Proceedings · Session S-207 · filed October 10, 2026

AI & Emerging Tech in R&DSession paper

20,000 Simulated Data Points Plus 65 Bioreactor Runs: A Case for Multi-Fidelity ML in Bioprocessing

Melbourne researchers show multi-fidelity ML — 20,000 simulated points plus 65 bioreactor runs — beats experimental-only training for mAb titer prediction.

By Tom Whitfield4 min read747 words

Summary

  • Multi-fidelity models trained on 20,000 simulated CHO fed-batch data points plus 65 bioreactor runs predicted mAb titer more accurately than experimental-data-only training.
  • The review was led by Mohammad Golzarijalal, PhD, of the University of Melbourne's digital bioprocess hub, published in Biotechnology and Bioengineering.
  • Low-fidelity data sources include previous production runs, off-configuration bioreactor cultures, and low-cost mechanistic or empirical models.
  • High-fidelity bioprocess data is expensive to generate and scarce, limiting ML model accuracy and driving overfitting outside observed conditions.
  • Potential applications include predicting cell growth, viability, metabolite concentrations and titer, and guiding media and feeding optimization.

A machine learning model trained on 20,000 simulated CHO fed-batch data points plus just 65 experimental bioreactor runs predicted final monoclonal-antibody titer more accurately than a model trained on the experimental data alone. That result, from researchers at the University of Melbourne's digital bioprocess hub, sits at the center of a new review published in Biotechnology and Bioengineering arguing that multi-fidelity training sets could break the data bottleneck now limiting machine learning in biopharmaceutical process development.

The core problem the review addresses is cost. Machine learning models could, in principle, let process scientists predict the impact of parameter changes before running experiments. But the high-fidelity process data required to train such models is expensive to generate, and as a result it is scarce. Teams developing a cell culture process rarely have enough experimental runs to build a model that generalizes beyond the conditions already tested.

Lead author Mohammad Golzarijalal, PhD, a research fellow at the University of Melbourne's digital bioprocess hub, proposes a specific structural fix: pair the scarce high-fidelity data with abundant low-fidelity data that captures broad process trends.

What does multi-fidelity training actually do?

Golzarijalal described the mechanism directly: "The basic idea is that abundant low-fidelity data teach the model the broad global trend of the process, while a smaller amount of high-fidelity data corrects and refines these trends, aligning surrogate model predictions more closely with the available ground truth."

The approach counters a known failure mode. "A model trained only on limited high-fidelity data can overfit and perform poorly outside the conditions it has already observed," he said. "By learning useful trends from lower-cost data, a multi-fidelity model can improve predictive accuracy and explore a wider process space without requiring the same number of expensive experiments."

The headline evidence comes from the group's own study, cited in the review: multi-fidelity Gaussian-process models combining the 20,000 simulated data points with data from 65 bioreactor runs outperformed models trained only on experimental data for predicting final mAb titer. For R&D managers weighing simulation investment against wet-lab spend, that is the relevant comparison — measured accuracy gain, not projection.

Where does low-fidelity data come from?

The review identifies three practical sources already available to most process development groups:

  • Previous production runs
  • Cultures grown in bioreactor configurations that differ from the process under development
  • Low-cost mechanistic or empirical models that generate simulated data

Drug firms already use such data informally. Scientists routinely draw on previous experiments to set parameter ranges, design design-of-experiments studies, and steer process optimization. Golzarijalal contends that this usage, while common, is limited.

How systematic is the opportunity?

The gap the review targets is methodological rather than technological. "There is an opportunity to use these data more systematically," Golzarijalal said. "Multi-fidelity algorithms provide a structured way to determine how much information should be transferred from historical or simulated data to a current problem."

That framing carries a portfolio implication. Process groups typically treat each new program in isolation, discarding accumulated knowledge from prior campaigns. "This can make past experience more useful for prediction and decision-making, rather than treating each new program largely in isolation," Golzarijalal said.

The potential payoffs, if the accuracy gains hold at scale, span the standard cell culture workflow: predictions of cell growth, viability, metabolite concentrations and product titer, plus guidance on media optimization, feeding schedules, seeding density and operating conditions.

What are the caveats?

The strongest quantitative claim in the review rests on one in-house case study — 20,000 simulated points, 65 runs, one output variable (titer). The broader review synthesizes prior work rather than reporting new multi-program validation, and it does not quantify how the simulated-to-experimental ratio should scale across cell lines, scales of operation or product modalities. Teams evaluating the approach will need to weigh whether mechanistic models of their own processes are accurate enough to serve as the low-fidelity layer without biasing predictions.

What the review does establish is a concrete, measured proof point that historical and simulated data carry transferable signal, and a structured method — multi-fidelity Gaussian processes — for deciding how much of that signal to move into a new problem. Golzarijalal and colleagues suggest the next step is systematic deployment: folding legacy run data and cheap simulations into training pipelines so each new program starts from accumulated process knowledge rather than a blank slate.

via analyticalsciencejournals.onlinelibrary.wiley.com (Original)

Filed under

  • machine-learning
  • bioprocessing
  • multi-fidelity-modeling
  • monoclonal-antibodies
  • process-development
Share this article:

More from Tom Whitfield

Tom Whitfield

Show full bio

Senior reporter covering media and advertising at Hypothesis Wire.

190 articles

References

  1. Human-Guided AI Gains Ground in Translational Science Workflows
  2. Multiomics in 2026: Five Vendors Name Sample Prep and Coherence as Bottlenecks
  3. Sungkyunkwan Team at 80% on Hybrid Digital Twin for CHO Cell Bioreactors
  4. 10-Pharma Consortium Pooled 2.6 Billion Data Points — and Won
  5. AbbVie, Takeda Join Ginkgo-Apheris Antibody Developability Push

« Previous articleNext article »