Biomedical AI · global
GSK Expands Collaboration With Relation to Train Biology Foundation Models Using Human Cell Perturbation Data
The collaboration, worth up to $110 million, will connect automated cell experiments, multi-omics measurements, and AI models into a target discovery pipeline; however, guaranteed payments, disease scope, and model validation results have not been disclosed.
The AI race in drug development is gradually shifting from algorithms to how data are generated. UK biotechnology company Relation Therapeutics and GSK have expanded their collaboration and plan to conduct large-scale experiments using human cellular disease models to build proprietary data for training cellular biology foundation models. Relation is eligible to receive upfront and milestone payments totaling up to $110 million.
The collaboration is not primarily intended to deliver a drug candidate directly, but to observe how cells change over time after genetic alterations or drug treatment. Automated laboratories will repeat these “perturbation” experiments under more consistent conditions, then use multi-omics technologies to measure gene regulation, molecular pathways, and other cellular states, thereby mapping the relationships between interventions and biological responses.
The resulting data will be used to develop, validate, and train models including MORGAN. MORGAN stands for “Multi-Omic Regulatory Genomic Analysis with Artificial Neural Networks” and is designed to predict human cellular responses to genetic and pharmacological interventions across different cell types and disease contexts, helping researchers screen disease mechanisms and potential therapeutic targets.
This approach seeks to address common gaps in public biological data: samples, measurement methods, and time points from different studies are difficult to integrate and may not include the specific disease contexts required for drug development. Relation says its high-throughput experiments can generate petabyte-scale training data. However, the information currently available does not specify the volume of data, types of cellular models, or control designs, nor does it provide predictive accuracy from external testing or prospective experiments.
For GSK, the collaboration could deploy AI earlier in drug development, strengthening the basis for selecting biological targets before committing to expensive animal studies and clinical trials. However, reproducible responses in cellular models do not necessarily translate into efficacy in humans. Whether the models can maintain their performance on cells, diseases, and drugs not included in training will still need to be tested progressively through independent experiments.
Background
The two companies began collaborating in 2024 to explore targets for fibrotic diseases and osteoarthritis. The new agreement broadens the focus to the data infrastructure supporting the AI platform, but the diseases covered, allocation of research results, and subsequent development and commercial rights for targets have not been disclosed. The amount of the guaranteed upfront payment is also unclear. The $110 million should therefore be understood as the maximum potential payment if the relevant conditions are met, rather than confirmed revenue.
The deal indicates that competitive advantage in biomedical AI may come from both models and experimental systems: those able to continuously generate consistent-quality data with a temporal dimension that closely reflect disease physiology may have a better chance of improving their models. But until public benchmarks, wet-lab validation, and results showing progression from targets to drug candidates are available, this remains an investment in a research platform and cannot yet be regarded as validation of a new therapy.