GSK AI Drug Discovery Deal Puts Biological Data At Center
AI News reported that GSK and Relation Therapeutics expanded their AI-assisted drug discovery work with a deal worth up to $110 million, emphasizing disease-specific cellular data as model scaling shows limits.

GSK's new AI drug-discovery collaboration with Relation Therapeutics puts the commercial emphasis on disease-specific biological data rather than model scale alone.
AI News reported that the research agreement is worth up to $110 million and will have Relation generate large datasets on how cells react when genes are altered and drug candidates are introduced.
The agreement makes laboratory data production part of the model-building workflow.
Relation will use the generated cellular datasets to train models for identifying possible drug targets, including systems in its MORGAN platform, while the work extends earlier GSK and Relation projects in fibrotic diseases and osteoarthritis.
Data Quality Becomes The Operating Surface
Relation's Lab-in-the-Loop approach combines tissue profiling, single-cell and spatial transcriptomics, sequencing, target validation and computational analysis.
The same workflow uses perturbation experiments to measure how genetic changes affect disease-linked cellular characteristics, giving model builders evidence that is closer to a specific disease program than a broad public corpus.
Public datasets still anchor much of the field.
A 2025 Experimental & Molecular Medicine review identified three major single-cell repositories: CZ CELLxGENE, Human Cell Atlas and NCBI Gene Expression Omnibus.
The 2025 review found that CZ CELLxGENE alone gives access to more than 100 million standardised cells.
The review also warned that data produced under different sampling methods, sequencing protocols and processing pipelines can introduce noise, imbalance and leakage when the same or similar cells appear across public resources.
The practical issue for pharmaceutical AI teams is that a larger training pile does not automatically create a better biological representation.
Nature Methods research from June 2026 used a corpus of 22.2 million cells, trained 400 models and tested them across 6,400 experiments.
The study found performance plateaus in current single-cell foundation models and did not show the simple scaling pattern associated with large language models.
Biopharma Deals Follow The Dataset Constraint
The GSK-Relation deal fits a broader shift toward specialised biological inputs.
Relation has applied its approach to Osteomics, a functional single-cell bone atlas built from patient-derived samples and combining patient samples, imaging, genomics, proteomics, clinical phenotype records and both spatial and single-cell omics for osteoporosis research.
A 2025 Nature Biotechnology analysis treated specialised dataset providers as one trend in AI-focused biopharma partnerships.
The same analysis pointed to GSK's separate $37.5 million agreement with Ochre Bio for human liver single-cell and perfused-organ data, while AstraZeneca and Pathos AI entered a $200 million agreement with Tempus involving records for more than 150,000 patients, spanning de-identified clinical information, genomic material and imaging data.
For drug developers, those figures describe a procurement problem as much as a modelling problem.
Disease-relevant data, consent boundaries, experimental consistency and access rights can determine whether an AI target-discovery system produces useful candidates for a therapeutic program.
GSK and Relation have not published target names or model performance results from the new collaboration.
The work now turns on whether newly generated human cellular data can produce target evidence that is strong enough to move from computational prioritisation into drug-development decisions.




















