From reactive firefighting to deterministic, AI-led incident management
AI incident management is the emerging practice of using autonomous SRE agents and AI-assisted observability platforms to detect, analyze, and resolve production incidents in real time by grounding decisions in deterministic system context rather than probabilistic guesses, reducing manual triage while keeping humans in charge of remediation choices. Instead of waiting for dashboards to turn red and engineers to scramble, Dynatrace and Expedia are building systems that turn continuous telemetry into immediate, structured responses to problems. Dynatrace Intelligence moves beyond probabilistic outputs by grounding every action in deterministic, real-time understanding of complex environments, while Expedia’s Service Telemetry Analyzer (STAR) focuses on predefined diagnostic workflows that turn raw metrics into root cause assessments. The message is clear: the future of DevOps is not more alerts—it is fewer incidents that ever need a human’s attention.

Dynatrace: autonomous SRE agents that act on facts, not guesses
Dynatrace’s latest update to its observability platform automation is a direct challenge to probabilistic, LLM-only operations. The company has added new autonomous agents for incident triage and remediation, paired with no-code custom agent creation capabilities. These agents use deterministic, real-time system understanding to decide what is happening and what to do next, rather than improvising based on probabilities. The Autonomous SRE Agent triggers on newly detected problems, checks whether they belong to an existing investigation, and enriches that investigation with additional insights. The Cloud SRE Agent then coordinates remediation activities across AWS, Microsoft Azure, and Google Cloud environments, keeping a single auditable record for autonomous operations. This is AI incident management designed for real-time incident resolution: precise cause detection, clear remediation options, and human oversight and governance preserved throughout.
Expedia’s STAR: AI-assisted investigation, not runaway autonomy
Expedia’s STAR platform takes a more conservative—but no less important—path. Service Telemetry Analyzer is an internal AI-assisted observability platform that helps engineers investigate production incidents by analyzing service telemetry and generating structured root cause assessments. It combines operational metrics with large language models through predefined diagnostic workflows, aiming to cut the time engineers spend identifying the source of service degradation while keeping humans responsible for validation and decision-making. STAR focuses on standardized infrastructure telemetry from Kubernetes-based services and JVM applications—throughput, latency, error rates, CPU and memory use, container restarts, probe failures, heap utilization, and garbage collection activity. Crucially, Expedia has chosen not to use function calling, RAG, memory, or autonomous tool use, instead relying on prompt chaining in a deterministic workflow to produce consistent analyses. This is observability platform automation with guardrails: AI accelerates the "time to know", but engineers still own "what we do about it".
What this automation changes for MTTR, teams, and tools
Both approaches are a reaction to the same pain: AI systems often lack the real-time context and controls to make reliable decisions, so promises of automation fail in production. Dynatrace counters this by grounding autonomous agents in causal, deterministic context, then allowing customers to adopt a crawl‑walk‑run model where humans start in the loop and gradually move to human‑on‑the‑loop operation as confidence grows. Expedia defines STAR’s objective as minimizing time to know (TTK) and time to recover (TTR), using LLMs to automate root cause analysis so engineers can decide remediation faster. In practice, that means lower mean-time-to-resolution as incidents are triaged, correlated, and diagnosed in seconds, not hours. As one operations leader notes, Dynatrace’s automation grounded in real-time context reduces manual effort and frees teams to focus on higher‑value work while improving outcomes.
The road ahead: from AI-assisted SRE to unattended autonomy
The most important shift is philosophical: success is no longer measured by how quickly engineers react, but by how many incidents never need a human at all. Dynatrace Intelligence goes beyond providing answers to acting on them automatically, and its Agent Builder extends autonomous operations to custom workflows without code. Cloud SRE Agent, Enhanced Dynatrace Assist, and expanded integrations are already available to SaaS customers on its platform today, while Autonomous SRE Agent and Agent Builder are expected in August. Expedia, meanwhile, is planning to add service dependency information, more operational metadata, Model Context Protocol–based tools, and conversational interfaces, and is evaluating STAR for chaos engineering analysis. DevOps teams that cling to manual triage will fall behind. Those that treat autonomous SRE agents and AI-assisted observability as core infrastructure will gain a real advantage: predictable reliability in environments that only get more complex.






