From Descriptive Chat to Executable Science
AI agents in drug discovery are software systems that can translate natural language instructions into complete, autonomous research workflows, calling scientific tools through APIs to design proteins, screen small molecules, and analyze genomic data without constant human steering. This marks a clear shift from chatbots that only summarize papers or draft reports, toward agents that can run validated, domain-specific tasks end-to-end. Instead of scientists manually stitching together models for protein-structure prediction, molecular docking, and drug screening automation, enterprise AI agents now orchestrate these steps as callable skills. They install command-line tools, authenticate to services, track workflow state, and enforce rules about what they are allowed to do. For biotech companies, this change turns AI from a descriptive layer on top of existing pipelines into an executable platform that can sit inside R&D, where it starts to automate real experiments rather than commentary about them.
BioNeMo Agent Toolkit Signals Enterprise Readiness
NVIDIA’s BioNeMo Agent Toolkit shows how quickly enterprise AI agents are becoming standard infrastructure in life sciences. The toolkit packages NVIDIA’s life-sciences software and models as documented “skills” that any agent platform can call, covering tasks such as protein-structure prediction, molecular docking, generative chemistry, and genomic analysis. Because it is agent-agnostic, developers can connect their preferred large language model to the same scientific toolbox while keeping customizable governance over what workflows are allowed. In a press conference, NVIDIA Vice President of Healthcare Kimberly Powell described the toolkit as a flexible framework where the agent harness acts like an operating system, remembering earlier steps and enforcing rules. The early traction is notable: NVIDIA announced adoption from nearly 50 partners, including Eli Lilly, Thermo Fisher Scientific and Dassault Systèmes, which indicates that enterprise AI agents are moving beyond proof-of-concept and into production deployments across drug discovery and development teams.

Boltz’s Agent-First API and Natural Language Screening
Boltz Bio illustrates the agent-first design that is starting to define AI agents drug discovery platforms. Its new API exposes BoltzProt-1 for protein design and BoltzMol-1 for small-molecule hit discovery as usage-priced services that coding agents can call directly. Boltz’s own scientists have already been using agents like Claude Code, Codex and Gemini CLI as their main interface to these models, and the company now offers public SDKs and connectors tuned for autonomous operation. In a test, an AI agent using Claude Code and the Boltz connector was asked in plain English to run a small hit-discovery screen against the EGFR kinase, a well-known cancer target. The agent installed and authenticated the command-line tool, refused to fabricate sequences or compounds when prompted, estimated computational cost before running, and returned ranked structures within minutes. For drug screening automation, this shows how agents can now drive complete workflows by talking to APIs designed “for agents as much as for people.”
From Text Tools to Autonomous Research Workflows
Together, platforms like BioNeMo Agent Toolkit and the Boltz drug-discovery API show a broader transition from descriptive AI to executable agents. General-purpose models alone struggle to break a request like “design a binder” into the five to seven domain-specific tasks involved, so toolkits now expose those tasks as modular skills that agents can chain into full workflows. In parallel, frontier-model integrations are being tied to automated labs: OpenAI and Molecule.one recently reported a near-autonomous AI chemist that ran more than 10,000 reactions through an agent plus lab setup, with gains that human chemists later reproduced at the bench. These examples suggest a growing role for AI agents in end-to-end autonomous research workflows, where they do not merely propose hypotheses but also plan experiments, call validated tools, and generate results that can be checked and repeated in wet-lab settings.

Production Deployment Amid Ongoing Skepticism
Despite visible momentum, scientists’ trust in enterprise AI agents remains uneven. Recent surveys show that many researchers still see AI as most useful for text-heavy tasks like literature search and report writing, not for the wet lab. According to a Pistoia Alliance poll of 300 life-sciences leaders, only 1% said AI is delivering value in the wet lab and 13% in automating scientific workflows and experiments. Other large polls report rising worry about hallucinations and lingering doubts that AI outperforms humans on important tasks, even as adoption climbs and “shadow AI” use through personal accounts becomes common. At the same time, nearly 50 partners signing up for the BioNeMo Agent Toolkit and scientists at Boltz running their models mainly through coding agents suggest that validated tools and enterprise partnerships are pushing AI agents drug discovery platforms into daily production use, where skepticism will meet hands-on experience.






