Proceedings · Session S-943 · filed October 10, 2026
AI & Emerging Tech in R&DSession paper
10-Pharma Consortium Pooled 2.6 Billion Data Points — and Won
A model trained on 2.6 billion pooled data points from ten pharma firms beat any single member's results, NIIMBL-led researchers report — and they argue biomanufacturing should follow.
By Rebecca Stone3 min read642 words
Summary
- The MELLODDY Consortium pooled 2.6 billion data points covering 21+ million molecules and 40,000+ assays from ten pharma companies.
- The shared model's median Relative Improvement of Proximity to Perfection for conformal efficiency exceeded 12%; one task exceeded 20%.
- Roger Hart (NIIMBL) and Kelvin Lee (University of Delaware) published the argument in Frontiers in Bioengineering and Biotechnology.
- Barriers include costly legacy infrastructure redesign, low vendor priority on sensor interoperability, unclear regulatory evaluation, and scarce hybrid-skilled staff.

A drug-activity model trained on 2.6 billion proprietary data points from ten pharma companies outperformed predictive models built by any single member of the consortium that shared the data. That result, from the MELLODDY Consortium — Amgen, Astellas, AstraZeneca, Bayer, Boehringer Ingelheim, GSK, Janssen, Merck KGaA, Novartis, and Servier — anchors a new argument that biomanufacturers should standardize and share big data industry-wide without exposing trade secrets.
The case appears in a paper in Frontiers in Bioengineering and Biotechnology by Roger Hart, PhD, senior fellow and Big Data Project lead at the National Institute for Innovation in Manufacturing Biopharmaceuticals (NIIMBL), and Kelvin Lee, PhD, professor at the University of Delaware, with colleagues. The consortium pooled decades of proprietary data spanning more than 21 million molecules and more than 40,000 assays in on-target and secondary pharmacodynamics and pharmacokinetics, using a secure computing platform to jointly train a shared quantitative structure-activity relationship (QSAR) model. The model, they write, used "the chemical structure of potential drug compounds to predict how they will function."
How big were the measured gains?
The consortium's results, reported separately by Wouter Heyndrickx, PhD, a machine learning scientist at consortium lead Janssen Pharmaceuticals, in the Journal of Chemical Information and Modeling, showed improvements in "most of the classification or regression tasks," with predictivity generally enhanced and in some cases substantially so. The hardest number: the median Relative Improvement of Proximity to Perfection for conformal efficiency exceeded 12%, and one task exceeded 20%.
Those are measured outcomes on pooled data — not projections. Hart and Lee's paper, by contrast, extrapolates the drug-discovery lesson to biomanufacturing: data from R&D, tech transfer, facilities operations, supply chain, and quality optimization could become predictive assets rather than storage overhead.
What would sharing require?
The authors are candid about the preparatory workload. Companies would need to:
- Standardize ontologies, schemas, and holistic data integration, and break down enterprise-wide data silos
- Enable predictive process-optimization strategies such as digital twins, in silico process design, and hybrid mechanistic/AI models
- Develop secure, shared datasets that allow pooling while protecting proprietary information
- Ensure interoperability and real-time connectivity among sensors from multiple vendors
- Build industry guidance and buy-in, including regulatory clarity, workforce development, and shared business cases
Each item carries budget weight. Redesigning legacy data infrastructure is expensive and time-consuming, Hart notes, and multivendor sensor operability ranks as a low priority for equipment developers. Two additional barriers sit outside any single company's control: regulators have not yet decided how such models will be evaluated in data submissions, and workers combining biomanufacturing and data science expertise remain scarce.
In-house or consortium?
The paper's core question for R&D managers is whether in-house analytics or ecosystem-wide collaboration delivers more value. "Many organizations are already pursuing in-house solutions," Lee told GEN. "The benefits are the ability to tailor the solution to one's specific situation. But by participating in larger-scale efforts, one can benefit from shared learnings and efforts to accelerate the digital transformation. Ecosystem-wide efforts can enable the entire industry to advance digital tools and benefits in ways that might not be apparent to individual organizations."
He adds that industry-wide solutions can accelerate individual organizations' learning, facilitate tech transfers, foster flexibility and interoperability, and support federated learning — a direct portability argument for companies running multiple sites or CDMO relationships.
The competitive framing is blunt. "Companies that do not adopt these capabilities are competitively disadvantaged from realizing the full benefits of their big data in the biopharmaceutical manufacturing marketplace," the team writes.
The MELLODDDY precedent — ten competitors, one secure platform, measurable gains over solo efforts — now stands as the reference case. Whether biomanufacturing data follows the same path depends on standards bodies, sensor vendors, and regulators converging on interoperability rules that today do not exist.
via frontiersin.org (Original)
Filed under
- machine-learning
- data-sharing
- pharma-consortium
- biomanufacturing
- drug-discovery
More from Rebecca Stone
Show full bio
Market editor covering marketplaces and e-commerce at Hypothesis Wire.
183 articles
References
- Structuring Tech Transfer: PharmTech Spotlights Collaboration as the New Default
- University of Chicago Team Swaps Single Atoms to Speed Drug Discovery
- AbbVie, Takeda Join Ginkgo-Apheris Antibody Developability Push
- AstraZeneca Puts $2 Billion Into Summit Therapeutics
- KRAS G12D Inhibitor Deal Values Cross-Border Pact at $2.13B