Genetic Engineering and Biotechnology News

Research and development concepts in science laboratories.

Adding Low-Fidelity Data to ML Training Sets Could Yield Better Process Models

Credit: Lu ShaoJi/ Getty Images

Combining trend information with high-fidelity experimentally-derived data could improve the accuracy of machine learning-based process models, according to researchers.

Machine learning models have the potential to make biopharmaceutical process development faster and more efficient by, in theory, allowing scientists to predict the impact of changes to parameters ahead of time.

The problem, according to recent research, is that the high-fidelity process information required to train machine learning models is expensive to generate and, as a result, in short supply.

One potential solution would be to include low-fidelity data, which indicate general process trends, in training sets, says lead author Mohammad Golzarijalal, PhD, a research fellow at the University of Melbourne’s digital bioprocess hub.

“The basic idea is that abundant low-fidelity data teach the model the broad global trend of the process, while a smaller amount of high-fidelity data corrects and refines these trends, aligning surrogate model predictions more closely with the available ground truth.

“A model trained only on limited high-fidelity data can overfit and perform poorly outside the conditions it has already observed. By learning useful trends from lower-cost data, a multi-fidelity model can improve predictive accuracy and explore a wider process space without requiring the same number of expensive experiments,” he says.

More accurate models could improve predictions about cell growth, viability, metabolite concentrations, and product titer. They could also guide media optimization, feeding schedules, seeding density, and operating conditions.

“One example discussed in our review is a recent study, which was performed in our own group, that combined 20,000 simulated CHO fed-batch data points with experimental data from 65 bioreactor runs.

“The multi-fidelity Gaussian-process models predicted final monoclonal-antibody titer more accurately than a model trained only on the experimental data,” Golzarijalal says.

Low-fidelity data

Low-fidelity process data can be from previous production runs or from cultures grown in bioreactor configurations that differ from those in the process under development. Such data can also be generated using low-cost mechanistic or empirical models.

Drug firms already use low-fidelity data to an extent, although, as Golzarijalal points out, use is usually limited. For example, scientists routinely draw on previous experiments to define parameter ranges, develop design-of-experiments studies, and guide process optimization.

“There is an opportunity to use these data more systematically,” he says, adding, “Multi-fidelity algorithms provide a structured way to determine how much information should be transferred from historical or simulated data to a current problem.

“This can make past experience more useful for prediction and decision-making, rather than treating each new program largely in isolation,” Golzarijalal adds.