From AI Commentators to AI Colleagues
AI agents in drug discovery are software systems that can call specialized scientific tools through natural-language instructions, break a drug-development goal into concrete steps, and then execute those steps across validated models and APIs without a human writing code for each action. For the past few years, pharma’s relationship with AI has been dominated by models that describe, summarize and predict, but stop short of doing the work. Now, a new generation of drug screening APIs and toolkits is pushing AI beyond analysis into execution, opening the door to more autonomous drug discovery workflows. This is not a cosmetic upgrade; it is a structural change in who—or what—can run a drug screen, design a protein or orchestrate a multi-step experiment.
This shift did not appear out of nowhere. Since AlphaFold 2 broke open protein-structure prediction and inspired an ecosystem of biomolecular models like Boltz, Chai, OpenFold and Protenix, the bottleneck has moved from model availability to usable access. Drug developers have been stuck choosing between running open models on their own GPU clusters or renting them via infrastructure platforms. That is a choice many teams are poorly equipped to make. The result: powerful models, limited adoption. Agent-focused platforms are a direct response to that gap, and they are changing who gets to participate in high-end computational biology.
Boltz’s Drug-Screening API: Turning Conversation into Chemistry
Boltz’s new drug-discovery API is the clearest example of this agent-first mindset. The company has released BoltzProt-1, a protein design pipeline, and BoltzMol-1, its first small-molecule hit-discovery pipeline, behind a usage-priced API. Crucially, they did not build this only for human programmers. Boltz’s own scientists, chemists and protein engineers have already been accessing these models through coding agents, and the same workflow is now exposed publicly via an API, SDKs and integrations with agent harnesses like Claude Code, Codex and Gemini CLI. The message is blunt: “We built this for agents as much as for people.”
What does that look like in practice? In one test, an AI agent connected to the Boltz API was asked in plain English to run a small hit-discovery screen: a handful of commercially available compounds against the EGFR kinase, a well-known cancer target. The agent installed and authenticated the command-line tool, refused to fabricate missing inputs, estimated the cost upfront, then delivered ranked structures in a few minutes. Because Boltz only bills for compounds it actually scores, the run came in below the initial estimate. This is not a chatbot explaining docking; it is an AI system performing a drug screen end-to-end. That removes the coding barrier and turns language into a control interface for complex chemistry workflows—exactly what AI agents in drug discovery were supposed to enable.
BioNeMo: A Toolkit So Agents Can Work, Not Perform
If Boltz is the new drug screening API, NVIDIA’s BioNeMo Agent Toolkit is the tool chest that many AI agents will reach for next. BioNeMo is an open, harness-agnostic platform that gives AI agents or software platforms the building blocks to specialize for science. The toolkit packages life-sciences software and models for protein-structure prediction, molecular docking, generative chemistry and genomic analysis as documented skills an agent can call on its own, so a general-purpose AI agent can carry out real scientific work rather than only describe it. One quotable framing from NVIDIA’s leadership is that “frontier models are the brains. BioNeMo is the scientific toolbox.”
The architecture matters. BioNeMo does not prescribe a single agent; it assumes multiple harnesses and models. Because it is agent-agnostic, any developer, regardless of which harness, model or platform they build on, can access the same accelerated scientific tools and skills with customizable governance. NVIDIA’s healthcare lead argues that general-purpose models cannot break a scientific request like “design me a binder” into the five to seven domain-specific tasks required without help, so agents must be given validated domain-specific scientific tools that researchers already use, with clear definitions of each tool’s purpose, inputs and outputs. That is a sharp rebuttal to the idea that bigger models alone will deliver autonomous drug discovery. Tooling and constraints, not raw model size, are what make agent behavior reliable enough for regulated science.

From Analytical Assistant to Autonomous Executor
Taken together, Boltz’s agent-ready drug screening API and NVIDIA’s BioNeMo toolkit mark a turning point: AI in pharma is moving from analytical assistant to autonomous executor. Agent plugins and integrations with models like Claude and Codex have already become the main way Boltz’s own chemists and protein engineers run its models. BioNeMo, meanwhile, launches with nearly 50 partners, including Eli Lilly, Thermo Fisher Scientific and Dassault Systèmes, adopting the toolkit. That scale matters. It signals that agent-based architectures are not fringe experiments but emerging defaults for pharmaceutical AI automation and life-sciences R&D.
At the same time, this shift is not fully trusted. Polls of life-sciences leaders show that only a tiny fraction see AI delivering value in the wet lab, and a small minority see value in automating scientific workflows and experiments. One expert describes the “agentic shift” as a future of “little experts” embedded into electronic lab notebooks to take routine tasks, but warns that earning science’s trust will take work. The near-autonomous AI chemist reported by OpenAI and Molecule.one—pairing a GPT-5.4 model, an agent (Maria) and an automated lab to run over 10,000 reactions and improve a stubborn sulfonamide coupling—shows what is already technically possible when agents, toolkits and automation align. Yet every such system will be judged less on novelty and more on whether its outputs reproduce at the bench.

Why Agent-Based APIs Standardize Drug Discovery Access
The deeper story here is standardization. Agent-based APIs are turning ad hoc, script-heavy drug discovery pipelines into services that broader research teams can use safely. Boltz’s release takes what had been an internal agent workflow and exposes it through an API, SDKs and integrations with mainstream coding agents. That means a biologist can ask an AI to run a hit-discovery campaign in plain language, while the agent quietly chains domain-specific skills behind the scenes. On the other side, BioNeMo’s agent-agnostic design gives developers a consistent set of documented skills and governance controls, regardless of the harness or model they favor. Together, these moves pull advanced computational chemistry out of the hands of a few scripting experts and into shared platforms.
This is why the move from AI commentary to AI execution matters politically inside organizations. Once drug screening APIs and BioNeMo-like toolkits become standard interfaces, the bottleneck shifts from writing code to defining questions, constraints and review criteria. Teams that once waited on scarce computational chemists can, in principle, trigger hit discovery or protein design from within their existing tools. Critics are right to point out that current agents still display gaps in reasoning and often fail to revise conclusions in light of new evidence, as recent large-scale evaluations show. But the direction of travel is clear: autonomous drug discovery will not arrive as a single, fully trusted AI scientist; it is emerging as an ecosystem of constrained, tool-using agents stitched into the everyday fabric of pharmaceutical workflows.






