From Talk to Action: What AI Agents in Drug Discovery Now Mean
AI agents in drug discovery are software systems that combine large language models with domain-specific scientific tools so they can plan, execute and revise multi-step research workflows, rather than only describing or summarizing experiments in natural language. This shift marks a move from single-task demos to context-aware systems able to carry out practical life-sciences work such as AI drug screening and protein design. The key change is not another headline model, but the appearance of infrastructure built explicitly for life sciences AI agents. NVIDIA’s BioNeMo Agent Toolkit gives agents documented “skills” for protein-structure prediction, molecular docking, generative chemistry and genomic analysis, so a general-purpose model can call those tools and perform real scientific tasks instead of staying in the realm of text commentary. In parallel, Boltz has released an API for its BoltzMol-1 and BoltzProt-1 pipelines that its team says is “built for agents as much as for people,” exposing the same workflows its own chemists use.

BioNeMo Agent Toolkit: Turning Frontier Models into Working Scientists
The most important signal that pharmaceutical AI automation is maturing is NVIDIA’s decision to ship tools, not yet another branded agent. The BioNeMo Agent Toolkit is an open, harness-agnostic platform that gives AI agents or software platforms the building blocks to specialize for science. Debuting with adoption from nearly 50 partners, including Eli Lilly, Thermo Fisher Scientific and Dassault Systèmes, it packages protein-structure prediction, docking, generative chemistry and genomics models as documented skills an agent can call on its own. NVIDIA is explicit about the division of labor: “Frontier models are the brains. BioNeMo is the scientific toolbox”. Instead of hoping a general model somehow “figures out” what it means to design a binder, BioNeMo lets developers wire in the five, six or seven domain-specialized tasks that process requires, with a harness that tracks workflows, enforces rules and remembers earlier steps. This is what moves life sciences AI agents from party tricks to reproducible research infrastructure.
Boltz’s Agent-First API: AI Drug Screening in Natural Language
If BioNeMo is a toolbox, Boltz shows what happens when the tools are built with AI agents as first-class users. Boltz’s new API wraps two pipelines: BoltzProt-1 for protein design and BoltzMol-1, its first small-molecule hit-discovery pipeline. These sit in a broader ecosystem that emerged after AlphaFold 2 cracked protein-structure prediction and AlphaFold 3 extended that success to biomolecular interactions, even as its public server remains off-limits for commercial use. Drug developers face a blunt choice: stand up open models on their own GPU clusters, or rent access through providers like Tamarind Bio, Rowan or NVIDIA BioNeMo. Boltz is betting that life sciences AI agents will tip that decision. Its API, SDKs and integrations with coding agents such as Claude Code, Codex and Gemini CLI expose the same workflows its internal scientists use. In the company’s words, “We built this for agents as much as for people”.
Multi-Turn, Context-Aware Agents Are Already Running Real Screens
The most telling evidence that AI agents drug discovery efforts are becoming practical is how these systems behave under real workloads. Boltz reports experimental validation of BoltzMol-1 across 10 targets covering GPCRs, kinases, ion channels and protein–protein interactions. In a test run through Claude Code and its Boltz connector, an AI agent was asked in plain English to carry out a small hit-discovery screen of commercially available compounds against EGFR kinase. The agent installed and authenticated the command-line tool on its own, refused to fabricate sequences or libraries, estimated cost before spending, and returned ranked structures in about three and a half minutes. The API priced the run at USD 0.025 (approx. RM0.12) per molecule, USD 0.20 (approx. RM0.92) for an eight-compound demo, but because Boltz only bills for compounds it scores, the run cost about USD 0.10 (approx. RM0.46). That is not a toy demo; it is a functioning AI drug screening service designed to be driven by agents.
From Single Tasks to Full Workflows—and a Cautious Path Forward
These developments matter because they mark a shift from single-step, brittle tools to multi-turn, context-aware life sciences AI agents. Kimberly Powell describes the agent harness as an operating system that connects models to tools, tracks where the workflow is, remembers what happened earlier and enforces what an agent can and cannot do. NVIDIA has already highlighted four AI agents built on BioNeMo across different R&D stages. At the same time, the industry has learned the hard way that clever workflows are not enough. A recent preprint covering more than 25,000 agent runs found that LLM-based agents executed scientific workflows competently but ignored gathered evidence in 68% of cases and rarely revised their conclusions when data contradicted them. In parallel, a poll of 300 life-sciences leaders found only 13% saw clear value in automating scientific workflows and experiments. The conclusion is uncomfortable but clear: real drug discovery work is finally within reach for AI agents, but only when they are grounded in validated tools and subjected to the same skepticism we apply to human scientists.






