From Chat Experiments to Production Infrastructure
Enterprise AI agents in production are software systems that use AI models, tools, and organizational data to execute ongoing tasks inside real business workflows, moving beyond isolated chat interfaces into governed environments that can persist, access internal systems, and meaningfully change how work gets done. The headline story is simple: AI agents are no longer a novelty, they are infrastructure. An analysis of companies running agents in production every month showed that the average number of artificial intelligence agents they activated nearly tripled over a 14‑month period as they moved beyond chatbots into business processes. "Companies were increasingly using agents to execute workflows instead of merely generating text". That shift forces a harder question than “what can it do?”—namely, “how do we operate this safely and reliably at scale?”

Blueberry Shows What Useful AI Incident Response Looks Like
If you want to see what truly useful AI incident response automation looks like, Instacart’s Blueberry is a better guide than any glossy demo. Blueberry is an AI-assisted incident response system designed to help on-call engineers investigate and troubleshoot production issues faster. It tackles the ugly, time‑consuming part of incidents: collecting context. In large operations, engineers burn precious minutes figuring out service ownership, recent deployments, logs, metrics, documentation, and whether this pattern has happened before. Instacart wired agents directly into that operational knowledge—incident history, ownership data, logs, and debugging signals—then integrated the output into Slack-based incident workflows so engineers never leave their existing channels.
The system launches about 10 subagents in parallel when an alert fires and delivers a grounded root cause hypothesis back into the Slack thread in roughly three minutes. Blueberry executed approximately 25,000 diagnostic passes in April across more than 270 Slack channels, and by grounding on 14 years of incident history, diagnostic accuracy improved from the mid‑60% range into the high‑90% range. This is not a toy assistant; it is AI incident response automation embedded in production, acting as a force multiplier that changes the starting point for on‑call engineers by giving them relevant information before deeper analysis begins.

Business Process Automation Agents Are Quietly Scaling
The noisy debate about chat interfaces hides the quieter reality: business process automation agents are already scaling across industries. The 2026 Agentic Enterprise Index examined usage data from organizations that consistently operated AI agents on one platform from February 2025 to April 2026 and found deployments nearly tripled in that window as companies moved them into business processes. These organizations had agents running in production every month, not sporadic pilots. Companies now create agents within an average of two days after provisioning them, and that time dropped by 53% over the measured period.
The work volume is growing, too: tasks completed by agents rose at a compound monthly rate of 15% as of April, measured as discrete “Agentic Work Units” per agent. Across industries, the average agent’s skill set expanded from two skills to six, including retrieving and summarizing information, drafting communications, updating records, and extracting structured data from user inputs. Retail agents tend to run narrow tasks most of the year, then ramp up to nine skills during peak shopping periods—a 350% jump—as companies assign them more complex customer-service workflows. Public-sector and healthcare organizations recorded 227‑fold and 19‑fold increases in agent work volume, respectively, and agents handled 170 times more customer‑service conversations over five quarters, resolving seven out of ten without human assistance. This is enterprise AI agents in production, quietly taking on real operational load.

The Agent Stack: Runtime, Access, and Risk Are the Real Story
The biggest mistake leadership can make now is treating agents like clever apps instead of infrastructure. The agent market has already started shifting from flashy demos to infrastructure decisions. For the last two years, executives watched multi-step task demos and asked, “what can it do?”—then signed pilot budgets. As agents begin to browse, click, spend money, and run unattended for hours, that question is not enough. Vendors are converging on the same hard problems: runtime environment, access boundaries, persistence, and cost control. One vendor announced Kitesurf, a browser built specifically for agents that runs inside V8 isolates on a worker platform with sandboxed outbound workers and durable objects. The point is not the brand name; it is the idea that agents need a purpose‑built runtime, not a full human browser that drags in unnecessary risk and cost.
Another provider added environment hooks, scheduled triggers, and an Environments API for managing sandboxes, plus budget controls to cap what an agent can spend. A third signalled agent‑specific “runtime instances” for persistent compute aimed at production agents. These moves recognize that an agent which persists across sessions, holds credentials, and keeps working carries a different risk profile than a disposable chat reply. If an agent can act outside the chat window—browse, call APIs, spend money, or persist—then AI infrastructure deployment questions must be answered before rollout: where does it run, what can it reach, how long does access last, how is spend capped, and what happens when it fails. Model quality still matters, but it no longer decides whether an agent is safe to deploy; infrastructure does.
What Actually Scales: Grounding, Governance, and Workflow Fit
The pattern across Instacart’s Blueberry and enterprise agent indexes is clear: what scales is not clever prompts, but grounded context, clear permissions, and tight workflow integration. Instacart’s experience shows that effective AI systems for operations depend on model capability plus the surrounding engineering framework—operational context, specialized workflows, tool integrations, and feedback loops. Blueberry works because it is wired directly into incident history, service data, and Slack workflows, and because it supports engineers rather than making unsupervised production changes. On the business side, companies that treat agents as workflow executors—retrieving, updating, and resolving across systems—see task volume and skill count grow, rather than stalling at single-purpose chat assistants.
The next phase will be less about “more agents” and more about better questions. Before adopting or expanding agent use, organizations should ask where the agent runs, whether the environment is sandboxed, what it can access and for how long, how spend and runtime are capped, and how failure is handled and logged. You are no longer buying a tool; you are choosing a piece of production infrastructure that will touch data and accounts and influence incident response and business processes. The companies that win with enterprise AI agents in production will be those that treat them as serious infrastructure decisions, connect them to their operational reality, and design them to complement human teams rather than replace them outright.







