Genetic Engineering and Biotechnology News

Binary code

Big Data Opportunities in Biomanufacturing

Credit: Loops7/Getty Images

Is your data a predictive asset or something that merely takes up storage space? Odds are good the answer is the latter. Yet, data sets from R&D, tech transfer, facilities operations, supply chain operations, and quality optimization hold details that could be predictive assets for current and future biomanufacturing, product development, and supply chain operations.

Researchers led by Roger Hart, PhD, senior fellow, Big Data Project lead, National Institute for Innovation in Manufacturing Biopharmaceuticals, and senior researcher Kelvin Lee, PhD, professor, University of Delaware, make the case not only for leveraging big data in your own company, but standardizing and sharing key insights industry-wide…without compromising proprietary information.

That opinion is based on results from a big data model developed through the MELLODDY Consortium (Amgen, Astellas, AstraZeneca, Bayer, Boehringer Ingelheim, GSK, Janssen, Merck KGaA, Novartis, and Servier).

As Hart, Lee, and colleagues report in a recent paper, the consortium pooled decades of proprietary data—some 2.6 billion data points spanning more than 21 million molecules and more than 40 thousand assays in on-target and secondary pharmacodynamics and pharmacokinetics—using a secure computing platform to jointly train a predictive drug-activity model. That shared quantitative structure-activity relationship model used “the chemical structure of potential drug compounds to predict how they will function,” they note.

Results, they say, significantly outperformed those of any of the individual members. Detailing those results in a separate article, Wouter Heyndrickx, PhD, machine learning scientist from consortium lead Janssen Pharmaceuticals, reports benefits in “most of the classification or regression tasks,” writing that the model generally enhanced predictivity, sometimes substantially. Notably, the median Relative Improvement of Proximity to Perfection for conformal efficiency exceeded 12%, with one exceeding 20%.

Prepping to share

“Many organizations are already pursuing in-house solutions,” Lee tells GEN. “The benefits are the ability to tailor the solution to one’s specific situation. But by participating in larger-scale efforts, one can benefit from shared learnings and efforts to accelerate the digital transformation. Ecosystem-wide efforts can enable the entire industry to advance digital tools and benefits in ways that might not be apparent to individual organizations.”

Achieving such results takes a lot of data management preparation, the researchers admit. They must:

  • Standardize ontologies, schemas, and holistic data integration and break down data silo enterprise-wide
  • Enable predictive strategies to optimize processes, such as those supported by digital twins, in silico process design, and hybrid mechanistic/ AI models
  • Develop secure, shared datasets that enable data to be pooled while protecting sensitive and proprietary data
  • Ensure interoperability and real-time data connectivity among sensors from multiple vendors so they can communicate accurately and in real- or near-real time
  • Industry guidance and buy-in, including regulatory clarity, workforce development, and shared business cases that lower adoption barriers

It is, in a word, challenging.

Redesigning legacy data infrastructure is expensive and time-consuming, and multivendor operability is a low priority for equipment developers, Hart says. There are additional barriers, too. Regulators have not yet determined how such models will be evaluated in data submissions. And, finally, workers with combined expertise in biomanufacturing and data science are scarce.

Yet, there are clear advantages to standardizing big data and using it to populate shared models.

“Many organizations face similar challenges and opportunities related to development and manufacturing,” Lee says. “Supporting an industry-wide set of solutions can accelerate the learnings of individual organizations, facilitate tech transfers…foster greater flexibility and interoperability, and support federated learning.”

As the team emphasizes, “Companies that do not adopt these capabilities are competitively disadvantaged from realizing the full benefits of their big data in the biopharmaceutical manufacturing marketplace.”