MilikMilik

How AI Agents Are Moving Beyond Theory in Drug Discovery

How AI Agents Are Moving Beyond Theory in Drug Discovery
Interest|High-Quality Software

From talk to action: AI agents enter real drug discovery

AI agents in drug discovery are software systems that combine large language models with validated scientific tools so they can plan, run, and interpret real drug-discovery workflows instead of only describing them in natural language. That definition might sound academic, but the shift it signals is not. For years, AI agents in science have been better at writing protocols than carrying them out. Now, purpose-built platforms from NVIDIA and Boltz show that this era is ending: agents are starting to install tools, launch hit screens, and return ranked candidates without a human typing command-line scripts. The center of gravity is moving from chatty assistants to working systems, and it is reshaping how discovery teams think about automation, accountability, and trust.

NVIDIA’s BioNeMo Agent Toolkit: a scientific toolbox for agents

NVIDIA’s BioNeMo Agent Toolkit is the clearest sign that life sciences AI automation is being industrialized rather than improvised. The company has packaged its life-sciences software and models for protein-structure prediction, molecular docking, generative chemistry and genomic analysis as documented “skills” an AI agent can call on its own, so a general-purpose model can carry out real scientific work rather than merely describe it. The platform is open and harness-agnostic, giving AI agents and software platforms the building blocks to specialize for science. NVIDIA reports adoption from nearly 50 partners, including Eli Lilly, Thermo Fisher Scientific and Dassault Systèmes. That level of early buy-in from conservative pharmaceutical and tools companies is a clear vote of confidence that agent-driven pipelines are worth the integration cost and governance effort.

Kimberly Powell describes the “harness” that sits around these skills as the operating system for an AI agent, connecting models to tools, tracking workflow state, and enforcing rules about what the agent can and cannot do. Crucially, NVIDIA is not shipping an end-to-end agent, but a BioNeMo agent toolkit that is agent-agnostic. This is an opinionated choice: it pushes responsibility for orchestration to developers while insisting that agents must use validated domain-specific scientific tools that researchers already trust. In an ecosystem crowded with generic APIs, BioNeMo is betting that the path to safer AI agents drug discovery runs through documented tools, measurable task completion, and fewer failed tool calls rather than clever prompts.

How AI Agents Are Moving Beyond Theory in Drug Discovery

Boltz’s agent-first API and autonomous drug screening

If BioNeMo is the toolbox, Boltz is showing how an agent can wield it in practice. Boltz Bio has released a new API that includes BoltzProt-1, a protein design pipeline, and BoltzMol-1, its first small-molecule hit-discovery pipeline. The company is explicit that it built this infrastructure for agents as much as for people, with SDKs and integrations for coding agents such as Claude Code, Codex and Gemini CLI so scientists can run complex workflows through natural conversation interfaces. According to Boltz, its own scientists, chemists and protein engineers have already been interacting with its models through these coding agents for months. That internal pattern is now exposed as a productized interface for autonomous drug screening.

A concrete test shows what agentic life sciences AI automation looks like when it leaves the slide deck. Using Claude Code with a Boltz connector, a small hit-discovery screen was run in plain English against the EGFR kinase, a well-characterized cancer target. The agent installed and authenticated the command-line tool on its own, refused to fabricate the protein sequence or compound library when prompted, estimated the cost before spending anything, and returned ranked structures in about three and a half minutes. Experimental validation of BoltzMol-1 has already covered 10 targets spanning GPCRs, kinases, ion channels and protein-protein interactions. In other words, this is not a toy demo: it is a pipeline with domain-specific validation, designed from the ground up for AI agents drug discovery rather than making them scrape web forms.

Why this shift is happening now

This move from descriptive to executable AI agents did not appear out of nowhere. Since AlphaFold 2 cracked protein-structure prediction, a growing ecosystem of biomolecular models has emerged, including Boltz, Chai, OpenFold and Protenix. AlphaFold 3 broadened the field from protein folding to biomolecular interactions, but its public server is off-limits for commercial use, with the commercial engine kept inside Isomorphic Labs. Drug developers are left either running open models on their own GPU clusters and pipelines, or renting access through infrastructure and workflow providers such as Tamarind Bio, Rowan or NVIDIA BioNeMo. That combination of powerful models, uneven access and hungry developers almost guaranteed an intermediate layer would appear: agent-focused toolkits that expose high-value skills without forcing every lab to become an infrastructure company.

The broader landscape reinforces the idea that this is a maturation rather than a fad. Alongside Boltz’s launch, agent plugins tied to Anthropic’s Claude and OpenAI’s Codex have become the main way its chemists and protein engineers run its models. Meanwhile, a “near-autonomous AI chemist” pairing GPT-5.4 with Molecule.one’s Maria agent and an automated lab has already executed more than 10,000 reactions, improving a stubborn sulfonamide coupling that human chemists then reproduced at the bench. BioNeMo’s toolkit, in this context, is less a moonshot and more a realism check: it recognizes that general-purpose models cannot on their own decompose a request like “design me a binder” into the five to seven specialized tasks that real scientific work demands.

How AI Agents Are Moving Beyond Theory in Drug Discovery

Trust, adoption, and what changes for discovery teams

The uncomfortable truth is that many scientists still do not trust AI, even as they adopt it. One study of more than 25,000 agent runs across eight scientific domains found that LLM-based agents executed workflows competently but ignored evidence they had already gathered in 68% of cases and rarely revised conclusions when data contradicted them. A poll of 300 life-sciences leaders found only 1% saw AI delivering value in the wet lab and 13% in automating scientific workflows and experiments. Yet demand is rising: many researchers already use public generative-AI tools through personal accounts, a “shadow AI” trend that highlights both unmet need and serious IP and compliance risk.

In this climate, early adoption of BioNeMo’s agent toolkit by nearly 50 partners, including major players like Eli Lilly and Thermo Fisher Scientific, matters more than any keynote rhetoric. It signals that leading discovery organizations see agent-driven automation as worth piloting inside regulated pipelines, provided the tools come with clear governance and domain validation. On the other side, Boltz shows what it means when an API is “built for agents as much as for people”: autonomous drug screening workflows that run end-to-end under natural language control. The trajectory is clear. AI agents will not replace scientists, but they are on track to become “little experts” embedded into notebooks and platforms, taking routine tasks off human plates while still being judged by the same metrics as any lab instrument: accuracy, reliability, and time saved.

Milik earns a commission when you shop through our links, at no extra cost to you. This article was generated with AI from published sources and product data.

You May Also Like

Comments
Say something...
No comments yet. Be the first to share your thoughts!