From Simulation to Execution: Defining the Agentic Shift
AI agents in drug discovery are software systems that do not only model molecules, but can autonomously call specialized scientific tools, execute end‑to‑end screening workflows, and return ranked candidates, all through natural language or coding interfaces while humans supervise the outcomes and key decisions. This is more than a cosmetic upgrade to existing models; it is a change in who is “driving” the workflow. For a decade, AI in pharma mostly meant simulations, predictions and slide‑deck explanations of how discovery should work. Now, platforms like NVIDIA’s BioNeMo Agent Toolkit and Boltz’s drug‑discovery API are inviting agents to run the workflows themselves. The important question is no longer whether AI can predict a protein structure, but whether enterprises will trust agents to operate parts of their AI life sciences workflow in production.
NVIDIA’s BioNeMo Agent Toolkit: Turning Models into Scientific Tools
The BioNeMo Agent Toolkit is a clear statement that general‑purpose models are not enough for serious science. NVIDIA has packaged its life‑sciences software and models for protein‑structure prediction, molecular docking, generative chemistry and genomic analysis as documented “skills” an AI agent can call on its own, so a general‑purpose agent can carry out real scientific work rather than only describe it. Because the platform is harness‑agnostic, any developer, regardless of which agent framework or base model they use, can access the same validated domain‑specific scientific tools with customizable governance. Debuting with adoption from nearly 50 partners, including Eli Lilly and Thermo Fisher Scientific, the toolkit signals that large pharma and biotech are no longer content with toy demos; they want AI agents drug discovery workflows that plug into existing infrastructure and respect lab rules from the start.

Boltz’s Agent‑First API: Autonomous Drug Screening by Conversation
Boltz is taking an even more opinionated stance: its new drug‑discovery API was “built for agents as much as for people.” The service exposes BoltzMol‑1, a small‑molecule hit‑discovery pipeline experimentally validated across 10 targets spanning GPCRs, kinases, ion channels and protein‑protein interactions, and BoltzProt‑1 for protein design. Crucially, the interface is tuned for coding agents with SDKs and integrations into tools like Claude Code, Codex and Gemini CLI, making agents the default way chemists hit the models. In a test run through Claude Code, an AI agent in plain English carried out a small hit‑discovery screen against EGFR: it installed and authenticated the command‑line tool on its own, refused to fabricate sequences or compounds when prompted, estimated a cost of USD 0.025 (approx. RM0.12) per molecule and USD 0.20 (approx. RM0.92) for an eight‑compound demo, then returned ranked structures in about three and a half minutes. That is autonomous drug screening in practice, not theory.
Enterprise Reality: From Pilots to Production Agentic Workflows
The most telling shift is not in the technology but in who is willing to put it to work. BioNeMo’s nearly 50 partners spanning pharma, lab equipment makers and software platforms show that agentic AI life sciences workflow deployments are moving out of the innovation theater and into production planning. At the same time, Boltz reports that agent plugins, including integrations with Anthropic’s Claude and OpenAI’s Codex, have become the main way its own chemists and protein engineers run its models, treating agents as their front door to discovery tools. This sits uneasily beside survey data: one poll of 300 life‑sciences leaders found only 1% saw AI delivering value in the wet lab and 13% in automating workflows and experiments, even as 45% of scientists admitted using public generative‑AI tools through personal accounts. The contradiction is obvious: institutional trust is lagging, but individual demand is forcing enterprises to catch up.

Why This Shift Matters—and What Must Not Be Automated Away
These agent platforms matter not because they are more impressive demos, but because they redistribute work. When an AI agent can break a goal like “design me a binder” into five, six or seven specialized tasks and call the right tools for each, the human role changes from operator to supervisor. Done well, this should mean better task completion, fewer failed tool calls, less integration overhead and faster movement from data to prioritized candidates. Done poorly, it risks encoding bad habits at scale. A preprint already warns that LLM‑based agents executed workflows competently while ignoring evidence they had gathered in 68% of cases and rarely revising conclusions when data contradicted them. The path forward is not to reject agentic AI, but to insist that every autonomous drug screening run is paired with strong governance and critical human review. Agents can now perform drug discovery work; they must never be allowed to perform scientific judgment alone.






