From dashboards to decisions: what autonomous SRE agents really change
Autonomous SRE agents are AI-driven software components that continuously monitor observability data, correlate signals, and execute predefined or learned actions to investigate, triage, and remediate production incidents with minimal human intervention, while still preserving explicit controls and auditability for reliability teams.
The real shift in AI incident response is not about adding another chat assistant; it is about replacing probabilistic guesswork with deterministic, observable actions. Dynatrace is explicit that its autonomous SRE agents ground every step in a “deterministic, real-time system understanding,” instead of loosely correlated predictions. Grafana Labs is pushing in the same direction, turning its Assistant into an agentic operations layer that detects, investigates, and remediates issues based on live telemetry rather than static rules. This is the long-awaited move from dashboards that inform humans to agents that own a slice of production incident triage. The opinionated takeaway: teams that cling to pure human-on-call workflows will soon be competing with counterparts who have automated their “time to know” down to seconds.

Dynatrace and Grafana: determinism, not magic, in AI incident response
Dynatrace’s latest release is a clear statement: AI in observability has to act on facts, not guesses. The company is adding new autonomous agents for incident triage and remediation, plus no-code custom agent creation, all built on Dynatrace Intelligence, which “grounds every action in deterministic, real-time system understanding” and aims to automatically resolve incidents while keeping human oversight and governance in place. These autonomous SRE agents trigger on newly detected problems, decide whether they belong to an existing investigation, and enrich that investigation with fresh insights before orchestrating remediation.
Grafana Labs is taking a parallel, opinionated stance on observability automation. Its six newly available AI capabilities extend Grafana Assistant into an agentic operations layer that not only surfaces anomalies but “detects, investigates, and remediates production issues at the pace AI now creates them”. When an incident hits, Grafana Assistant Investigations forms hypotheses, swarms the telemetry, and proves or disproves each lead, while Grafana Assistant Automations turns those findings into executable workflows for production incident triage.

Expedia STAR: LLM-powered observability without giving up control
While platform vendors sprint toward agentic operations, Expedia’s Service Telemetry Analyzer (STAR) is a reminder that not every team is ready for full autonomy. STAR is an internal AI-assisted observability platform that analyzes service telemetry and produces structured root cause assessments to help engineers investigate production incidents. It intentionally stops short of autonomous action: telemetry is collected, passed through domain-specific prompts in a deterministic workflow, and consolidated into a report with likely causes and recommended next steps, with humans still responsible for validation and decisions.
Technically, STAR sits on FastAPI, pulls metrics from Datadog, and routes them through an internal generative AI gateway to LLM providers. It focuses on standardized telemetry from Kubernetes and JVM services, including throughput, latency, HTTP/gRPC/GraphQL error rates, CPU and memory, container restarts, probe failures, heap usage, and garbage collection. According to Expedia, STAR has already supported production incident investigations, post-incident analysis, Kubernetes troubleshooting, and JVM memory diagnostics. This is AI incident response as decision support, not decision maker—and that is a sensible intermediate step for many enterprises.

Why this shift is happening now: speed, complexity, and trust
The timing is not accidental. Engineers are shipping more code, faster, and traditional observability—instrument, dashboard, alert, hope you notice in time—cannot keep up. Most AI projects have promised automation but failed to deliver because they lacked real-time context and operational controls. Dynatrace’s response is to combine agentic AI with deterministic context, giving AI “real-time understanding of complex environments” so that it can act reliably in production.
On the Grafana side, the gap is stark: its 2026 Observability Survey reports that “92% of practitioners say they’d get real value from AI catching anomalies, yet only 57% say they’re currently implementing observability for their own AI systems in any capacity”. That mismatch is the business case for observability automation and AI incident response. Expedia’s STAR team talks openly about minimizing time to know and time to recover as their primary objective. The message across these efforts is blunt: in an environment of accelerating change, the only sustainable path is to automate more of production incident triage while designing for trust—deterministic workflows, auditable histories, and human-on-the-loop control.
What comes next: from human-in-the-loop to human-on-the-loop
The most important strategic question is not whether to adopt agentic operations, but how quickly to move from guided AI to autonomous SRE agents. Dynatrace openly describes a “crawl-walk-run” model: first use deterministic causal AI to pinpoint root causes, then add remediation automation with human-in-the-loop, and only then graduate to human-on-the-loop autonomous actions as confidence grows. Some components, like the Cloud SRE Agent and enhanced assistant, are already available to SaaS customers, while the Autonomous SRE Agent and Agent Builder are expected to follow in August.
Expedia’s roadmap for STAR is similarly pragmatic. Planned enhancements include adding service dependency data, more operational metadata, Model Context Protocol-based tools, and conversational interfaces, and even using STAR in chaos engineering to analyze controlled failure experiments. The lesson is clear: successful AI incident response will evolve in stages, from LLM-powered diagnostics to semi-automated remediation to truly autonomous observability automation. Teams that treat this as a linear upgrade project will fall behind; teams that treat it as a new operating model—where agents share accountability for uptime—will define the next generation of reliability practice.






