What Is an AI SRE (and Which Parts of the SRE Job It Actually Does)
We define an AI SRE against the real SRE job, grade vendor definitions on Google's five autonomy levels, and list what to settle before one reads production.
Quick Answer
AI SRE tools mostly cover incident investigation, one slice of site reliability engineering. They gather logs, metrics, deploys and code, then propose a root cause and a fix for a human to approve. Google grades AI in operations on five autonomy levels, L0 to L4, and measured a 10% cut in time to mitigate from hypotheses alone. SLOs, error budgets and capacity decisions still sit with engineers.
Site reliability engineering is the practice of running production services with software engineering methods, against explicit reliability targets. When we searched "what is AI SRE" on 30 September 2026, every organic result on Google's first page came from a vendor. Several of them fold "resolves incidents" into the definition itself.
That makes it hard to tell what the tools do today. We think the cleaner test is the job itself. Google's SRE book lists what an SRE team owns, and a newer Google paper grades how much of that work an agent may touch.
In this post we'll map the AI SRE category against both, compare how vendors define it, and list the questions we'd settle before one reads production.
What an AI SRE Does Inside the Site Reliability Engineering Job
An AI SRE is an AI agent that investigates production problems across logs, metrics, deploys and code, then hands engineers a probable root cause with evidence. Measured against site reliability engineering as a whole, it covers the emergency-response and monitoring slices well, assists with a few more, and leaves target-setting alone.
We find it easiest to judge the category against a fixed list. The Google SRE book says an SRE team is responsible for availability, latency, performance, efficiency, change management, monitoring, emergency response and capacity planning. Here is how we read the AI SRE tools we reviewed against that list.
| SRE responsibility (Google SRE book) | What AI SRE tools do today | Who decides |
|---|---|---|
| Emergency response | Gather evidence, rank hypotheses, name a probable root cause, draft a fix or mitigation | An engineer approves any change |
| Monitoring | Enrich and group alerts before a human reads them | Engineers own alert design |
| Change management | Correlate an incident with recent deploys, flags and config changes | Engineers own rollout policy |
| Postmortems | Draft timelines and action items from chat and tickets | Engineers write the conclusions |
| Performance and efficiency | Surface slow queries, hot endpoints and cost spikes when asked | Engineers pick the work |
| Capacity planning | Summarize utilization trends | Engineers and finance decide |
| Availability and latency targets | Report SLO burn against targets someone else set | Engineers and product owners |
• Where It Does Real Work Today
Emergency response is where the evidence is strongest. Google's Incident Hypothesis system only suggests a likely cause and next checks. In an A/B test it cut mean time to mitigate (MTTM) by 10%, according to Google's paper on AI in SRE. The same paper reports a roughly 44% MTTM reduction from its Investigation Dashboards, for supported incidents.
Say your checkout API pages at 02:10 with p95 latency at three times its normal level. A human spends the first half hour opening dashboards, checking the last deploy and reading slow-query logs.
An AI SRE agent runs those checks in parallel and posts one lead, a migration dropped an index, and every checkout now scans the orders table. We'd still want an engineer to confirm that lead before anything changes.
• Where It Assists but Does Not Decide
Monitoring and change management get real help, but only as context. Google's AI Alert system intercepts alerts, queries monitoring, logs, change logs and dependency graphs within a budget of about two minutes, and appends what it finds. Google states that the system runs in read-only mode.
Postmortems fit the same pattern. A tool can reconstruct the timeline from chat, but the blameless postmortem still depends on people agreeing on contributing causes. We treat the AI draft as a starting point, never as the record.
• Where It Has No Role Yet
Choosing a service-level objective is a business decision, and so is spending an error budget. No AI SRE tool we reviewed claims to set those targets. The same goes for reliability architecture, such as running active-active, choosing which dependencies to accept, and deciding how teams own services.
This is also where most of the long-term value in site reliability engineering sits. We'd be wary of any pitch that blurs the line.
Pro tip: When a vendor demo shows an agent "handling" an incident, ask which row of the table above the demo covered. It is almost always emergency response.
Five Autonomy Levels Google Uses for AI in Site Reliability Engineering
Google grades AI in site reliability engineering on five levels, from L0 (humans do everything) to L4 (an agent runs multi-step resolution alone). The level is set per action, not per product, and an agent moves up only after it proves itself against human-verified data.
The table below reproduces the structure of Table 1 in Google's paper. "Mitigate" is the approval step, and "Actuate" is the change itself.
| Level | Monitor | Investigate | Mitigate (approve) | Actuate | Self-direct |
|---|---|---|---|---|---|
| L0 Manual | Automation | Human | Human | Human | Human |
| L1 Assisted | Automation | Automation | Human | Human | Human |
| L2 Partial | Automation | Automation | Human | Automation | Human |
| L3 High | Automation | Automation | Automation | Automation | Human |
| L4 Full | Automation | Automation | Automation | Automation | Automation |
1. L0 Manual, Humans Do Everything
Alerts fire automatically, and people do the rest. We'd guess plenty of teams with good dashboards still sit here, and that is a reasonable place to start.
2. L1 Assisted, the Agent Monitors and Investigates
The agent gathers data and proposes a hypothesis. A human approves and applies any fix. Google's 10% MTTM result came from this level alone, which we think is the most underrated number in the paper.
3. L2 Partial, the Agent Acts Only After Human Approval
The agent stages a mitigation, and a human must approve it before it runs. Google says its AI Operator agent works at this level for critical operations.
4. L3 High, the Agent Acts Alone on Bounded Scenarios
The agent can detect, decide and act without approval, but only for well-defined cases. Google limits AI Operator's L3 use to minor incidents, and routes every action through a separate safety gateway that can downgrade a request back to L2.
5. L4 Full, the Agent Runs Multi-Step Resolution
The agent plans, acts, watches the result and tries alternatives until the system is stable. Google describes L4 as the target for specific, well-bounded scenarios, not a general state.
Pro tip: Write your allowed level next to each action type, not next to the tool. Write "restart a pod is L2, roll back a deploy is L1, and schema change is never." That one page settles most security review debates.
How AI SRE Vendors Define the Category, Level by Level
Vendor definitions of an AI SRE disagree mainly on one point, how far up Google's autonomy ladder the tool goes. Read the wording closely and each definition implies a level, which tells you more about site reliability engineering risk than the feature list does.
When we build and evaluate AI SRE tooling, the level is the first thing we ask about. Here is what each vendor's own page said when we checked on 30 September 2026. The level column is our reading of that wording, not a claim the vendor makes.
| Vendor | What its own page says | Level the wording implies |
|---|---|---|
| Resolve | An autonomous agent that "resolves production incidents without human intervention" | L3 to L4 |
| Rootly | Graded autonomy, from read-only to approved actions and then bounded autonomous remediation | L1 rising to L3 |
| incident.io | "Autonomous remediation remains limited"; the engineer executes the remediation | L1 |
| Traversal | Triaging alerts, investigating infrastructure and diagnosing incidents | L1 |
| Datadog | An agent that investigates end to end to help on-call engineers pinpoint root causes | L1 |
| Splunk | Probable root cause delivered with a step-by-step remediation plan | L1 |
| Harness | "One-click runbooks to roll back, scale out, or toggle a feature flag" | L2 |
Three details stood out to us. Four of the seven SRE AI agents sit at L1 by their own wording. Only Resolve makes autonomous resolution the defining trait, while Rootly treats execution as a later stage. And Datadog's former Bits AI SRE documentation URL redirected to a page titled "Bits Investigation" when we checked it on 30 September 2026.
So the AI SRE meaning you get depends on who wrote the glossary. We'd read "AI SRE" on any page as "agent that investigates", then look for the words that describe actuation.
AI SRE vs AIOps, Runbook Automation and Chat Copilots
An AI SRE differs from AIOps, runbook automation and chat copilots because it produces a tested hypothesis about a specific incident, backed by evidence. The other three produce grouped alerts, executed steps or answers to questions someone already knew to ask.
The terms overlap in marketing, and much of what gets sold as SRE AI is one of these four things. We compare them by inputs and outputs, since each has a place in site reliability engineering.
| Tool type | Main input | Main output | Changes production? |
|---|---|---|---|
| AIOps | Metric and alert streams | Anomalies, grouped alerts, noise reduction | Rarely |
| Runbook automation | A known procedure | The procedure, executed | Yes, by design |
| Chat copilot | A human's question | An answer or a query | No |
| AI SRE agent | An alert or a question, plus logs, code, deploys and data | A root cause with evidence and a proposed fix | Depends on the level you allow |
Runbook automation speeds up the part of an incident that was already understood. We've argued before that execution was rarely the slow part. Diagnosis is. AIOps sits upstream, deciding which alerts belong together, and an AI SRE picks up from there.
The agentic AI SRE pattern is newer mainly because large language models can read code, logs and tickets in one loop. That capability is also the source of the failure modes below.
Where an AI SRE Agent Gets Site Reliability Engineering Wrong
An AI SRE agent fails in predictable ways because it confuses correlation with cause, trusts stale knowledge, inherits access it should not have, and loops without limits. None of these are exotic. Each has a known guardrail in site reliability engineering practice.
• It Mistakes a Correlation for a Cause
A deploy at 14:02 and an error spike at 14:07 look like cause and effect. Sometimes the deploy is innocent and a dependency failed at the same time. This is why we like a second, independent check on every proposed root cause analysis, the pattern we described in aviation's cross-check rule.
• It Trusts Stale Knowledge
Runbooks and wikis age faster than systems. An agent that retrieves an outdated page will state its conclusion just as confidently as a correct one. We want every claim linked to live evidence, such as a log line, a query result or a commit.
• It Inherits Credentials It Should Not Have
Google's paper sets a rule of "No Ambient Access" and says agents must not run with the standing credentials of the engineers who built them. It also requires a dry_run=true mode on any API an agent can call, so reviewers can see the blast radius first.
We think read-only access is the right default, and we explain why in our piece on automating root cause analysis without write access.
• It Loops Without a Budget
An agent that retries a failing mitigation can amplify the incident it was meant to contain. Google calls for agent-specific rate limits and circuit breakers, and lists prompt injection among the security risks. We'd add a hard cap on investigation time before the agent escalates to a person.
Pro tip: Before a pilot, give the agent one incident where the obvious suspect was innocent. How it handles that case tells us more than ten easy wins.
Questions We Would Settle Before an AI SRE Tool Reads Production
Before any AI SRE platform reads production, we'd settle five questions in writing. They cover autonomy per action, credentials, evidence, where data and the model run, and what happens when the agent gets stuck. These turn a site reliability engineering policy into something a security reviewer can check.
1. Allowed Autonomy Level for Each Action
List the actions, then assign L0 to L4 to each. We'd start every write action at L1 or L2, and let investigation run at L1 from day one.
2. Agent Credentials
The agent needs its own identity, separate from any human, with read-only scopes where possible. Every query should land in an audit log you control.
3. Evidence Required With Every Root Cause
A useful answer names the log lines, queries and commits it relied on, with links. A confident paragraph with no links is a guess, however good it sounds.
4. Data and Model Location
Logs and database rows often hold customer data. Decide up front where the agent runs, which model provider sees the prompts, and what leaves your network.
5. Escalation When the Agent Cannot Find the Cause
Google's AI Operator escalates to a human when it is outside safe bounds. It also posts its full investigation history, so the on-call engineer does not start from zero. We'd ask every vendor to show that handoff, not only the success path.
For a longer evaluation list, our buyer's guide to automated root cause analysis covers pricing, integrations and proof-of-value tests. Our note on the toil paradox in the survey data covers why adding AI does not automatically reduce toil.
Decide the Autonomy Level Before You Pick the Tool
An AI SRE is an investigator first. The evidence we found supports using it for emergency response and alert context, with measured gains even at L1. It also supports keeping SLOs, error budgets and architecture with the people who own site reliability engineering.
We'd write down the allowed level per action, require evidence with every answer, and ask to see the escalation path before signing anything.
Operate is built for that shape. It reads logs, the database and code inside your infrastructure, a second model verifies the root cause, and the fix arrives as a patch an engineer reviews.
Frequently Asked Questions
We don't read the evidence that way. Google's own paper on AI in site reliability engineering says SREs move up the abstraction ladder, from responding to incidents toward defining guardrails, curating evaluation data and governing agent behavior. The work changes shape, and accountability for reliability stays with people.
We don't think one tool wins for every team. The fairest test we know is a replay, where each candidate gets five resolved incidents from your own history and you compare its answer with what your postmortem found. The tool that names the right cause with linked evidence most often is the best one for you.
A common stack pairs Prometheus for metrics, Grafana for dashboards, OpenTelemetry for traces, PagerDuty or Opsgenie for paging, and Terraform for infrastructure as code. We'd get those basics, plus a status page and a postmortem template, working well before adding anything smarter on top.
Not exactly. The Google SRE book describes SRE as one specific implementation of DevOps with some idiosyncratic extensions, such as a cap on operational work. We treat the difference as emphasis because DevOps names a culture of shared ownership, and SRE adds measurable reliability targets and engineering time to meet them.
Partly. Wikipedia places site reliability engineering in both software engineering and IT infrastructure support. We see it as a software role pointed at operations because SREs write code to remove manual work instead of repeating that work by hand, so we'd hire for engineering skill first.
