ai robot in lab
Forget self-driving labs. For scientific AI to thrive, humans must always be in charge, setting objectives, constraints, risks, and interpreting results. [Maite Mueller/Getty Images]

AI is neither saint nor demon. Nor should it be a replacement for human scientists. As AI takes on greater roles in designing, executing, and analyzing experiments and processes, scientists understand that even the best AI needs human supervision.

The big question is how much oversight is needed and whether—or the extent to which—AI interactions should be documented and reported in regulatory filings.

Le Cong, PhD
Le Cong, PhD, associate professor, Stanford University, and co-founder of LabOS and MedOS

Although the lab of the future may be envisioned as a self-driving lab, that’s actually a bad idea, noted Le Cong, PhD, associate professor, Stanford University, and co-founder of LabOS and MedOS. Instead, he and leaders in the AI and biopharmaceutical industries see the future of scientific AI as agentic, with humans in charge.

“We think there is a positive trend toward using AI in a way that’s human, rather than as a self-driving lab,” he said. In that environment, AI is simply a tool—albeit a powerful, adaptive one—that can be managed as long as scientists use the right prompts.

Humans in the lead

“Today, much of the scientific research process remains inaccessible to machines,” Cong and colleagues wrote in a recent paper. Despite automation and some use of AI, “Scientific discovery remains fragmented.” Specifically, AIs lack the tacit knowledge, evolving experimental context, human observations, and adaptive decision-making inherent in human scientists.

Human involvement is needed, therefore, not just to oversee AI-based activities and check the output, but to ask the right questions and to ensure that analyses make sense in context. Specifically, he describes a scientific setting in which an AI would handle an experiment’s execution, and the scientists would be responsible for:

  • Framing objectives
  • Interpreting results
  • Setting constraints
  • Governing risks

“If AI can interpret everything, then it will start to generate fake stuff, right?” Cong asks. “We’ve seen this when AIs begin guessing in an effort to return results and supply citations that don’t exist. There are certain things that are useful for AI to do in the lab.”

But, as last summer’s sandbox breakouts illustrated, an AI needs firm guidelines as to what it can do, where it can access information, and the degree of autonomy it has in meeting a request.

For example, he recommends adding this phrase to instructions: “Any actions not explicitly stated in the protocol need human approval.” That default to human judgment also should apply to determining the risks associated with certain actions, such as editing a human gene, Cong said. With those guardrails, the paper points out, agentic AIs can freely handle “routine execution and coordination across models, instruments, protocols, and laboratory states.”

Cong equates the scientific use of AI to autonomous vehicles, which, according to the Insurance Institute for Highway Safety, have a 68% lower crash rate per mile traveled than human drivers in the same environment. Given those statistics, he added, “My thought is to elevate humans to setting destinations. We do not need humans to always execute the driving.”

In a university scientific lab, that equates to staffing a principal investigator and trainees, without much of the hierarchy that exists today. In a corporate environment, the hierarchy flattens to scientists who propose, design, execute, and interpret experiments, and a lab manager. The distinctions between senior and junior scientists blur because much of the hands-on work is automated.

The combination of AI and lab automation is expected to reduce human errors. Cong cited a 2016 Nature study of 1,500 scientists. When asked, “’Can you replicate other people’s experiments, and can you replicate your own after a few months?’ approximately 70% could not replicate others’ experiments, and half could not replicate their own!” Cong said. “AI can improve that.”

autonomous car
Cong equates the scientific use of AI to autonomous vehicles, which, according to the Insurance Institute for Highway Safety, have a 68% lower crash rate per mile traveled than human drivers in the same environment. Given those statistics, “My thought is, elevate humans to setting destinations,” said Cong. “We do not need humans to always execute the driving.” [Chesky_W/Getty Images]
Reasons for such poor reproducibility rest in the details, he elaborated. Was a step omitted? Was the protocol followed exactly? Is the protein being used identical to the one in the original experiment? Were the temperatures the same? Details like this—which AI can duplicate precisely—are behind many reproducibility challenges.

AI risks minimal

As yet, it’s unclear how AI involvement in experiments should be preserved and reported in regulatory submissions, Cong continued. “We’re still early in this journey.”

That said, the risk that AI will escape its constraints and cause physical harm—like designing and developing a physical virus—appears relatively low, according to Cong. That’s because a rogue AI still needs a human accomplice to allow a virus, for example, to be manufactured and released. “In areas where there is a physical execution step, I think AI is still incapable,” he said, “although we are seeing progress in connecting AI to biomedical labs and applications, and the physical execution layer.”

That’s due to the fact that there are multiple layers of human intervention needed to actually manufacture a product. Aside from logistics, he cites good manufacturing practices, safety and efficacy regulations, and digital safeguards like track and trace and the FDA’s 21 CFR Part 11, as well as real-time monitoring, periodic inspections, and quality control activities. Those regulations and checkpoints should also be sufficient to manage variations that occur during manufacturing as real-time conditions drift from specifications.

The catch, as last summer’s breakouts of frontier AIs underscore, is that sometimes AIs exceed their parameters. Whether there is sufficient appreciation of this among AI users remains to be seen, Cong said.

“A lot of people are connecting AI systems, which have access to more and more information and key decision-making systems,” Cong pointed out, without deeply understanding the risks and establishing appropriate guardrails. “People might be overly trusting of AI, perhaps.

“The more powerful the AI, the more capable it is of doing something. Are people keeping pace [with the technology and its risks]?”

Sometimes, small, highly specific AIs may be a better choice than always leveraging the large frontier models, Cong suggested. The reason, Cong, senior corresponding author Mengdi Wang, PhD, professor, Princeton University, and a dozen colleagues, noted in a 2025 paper in Nature Biomedical Engineering, is that “Large language models often lack domain-specific knowledge and struggle to accurately solve biological design problems.”

Whatever level of AI is used, however, “Humans need to be in the lead throughout the process,” Cong stressed.

Previous articleBuilt Without Bacteria: Bringing Cell-Free Synthesis to the Bench
Next articleScientists Analyze 267 Receptors That Control Protein Fate in Rare Diseases
Previous articleBuilt Without Bacteria: Bringing Cell-Free Synthesis to the Bench
Next articleScientists Analyze 267 Receptors That Control Protein Fate in Rare Diseases