Devops

AWS DevOps Agent goes multicloud: what changes for the AI SRE vendor map

Jorge de los Santos, CTO & Co-Founder · May 14, 2026 · 13 min read

Six weeks after GA, AWS DevOps Agent now investigates Azure workloads and discovers on-premises resources via MCP. The hyperscaler-native SRE agent isn't hyperscaler-only anymore.

AWS DevOps Agent goes multicloud: what changes for the AI SRE vendor map

The May 13 Announcement and Why It Reorganizes the 2026 Vendor Map

Six weeks after the March 31, 2026 general-availability launch, AWS announced on May 13, 2026 that AWS DevOps Agent now extends incident investigation beyond AWS environments — into Azure workloads, into on-premises applications, with cross-cloud topology correlation and unified incident response across AWS, Azure, and on-prem.

The mechanics matter:

  • Azure incident investigation. The agent investigates and correlates signals from Azure workloads alongside AWS workloads. The cross-cloud signal-correlation is the differentiation pitch.
  • On-premises resource discovery via MCP. The agent uses the Model Context Protocol to discover on-premises resources, analyze on-prem metrics, logs, and code, and build a comprehensive topology that spans cloud and on-prem. The MCP layer is the connective tissue between the cloud-native agent and the customer’s existing on-prem observability surface.
  • Custom skills. AWS shipped custom-skill support at GA, extending into the multicloud surface in this release.
  • Custom charts and reports. Output customization moves with the multicloud capability, which matters for the customer’s existing dashboards and reporting cycles.

The May 13 announcement reorganizes the 2026 AI-SRE vendor map that batch-twenty-four’s vendor-map post (published May 13) laid out. In that snapshot, hyperscaler-native agents were one of five categories — strong inside their home cloud, weak across clouds and on-prem. Six weeks of post-GA AWS DevOps Agent development has collapsed that category boundary. The hyperscaler-native agent is now a cross-cloud / hybrid-cloud surface.

The implication for platform-engineering leaders is structural, not tactical. The differentiation question for any non-hyperscaler AI-SRE surface — independent autonomous SRE agents (NeuBird, Causely, Aisera, Resolve.ai), observability incumbents extending into agentic surfaces (Dynatrace, Datadog, New Relic, PagerDuty, ServiceNow), skills-hub layers, and active-operational-layer agents like IAN — has shifted. “We run across clouds” is no longer a differentiation. The hyperscaler now runs across clouds too.

The Structural Re-Read: Where Differentiation Moves

With AWS DevOps Agent extending to Azure and on-prem via MCP, the differentiation surface for non-hyperscaler AI-SRE and active-operational-layer products moves to four places:

1. Cross-pillar coordination beyond incident response. Hyperscaler-native agents are still single-pillar in shape — AWS DevOps Agent is an incident-investigation agent, not a cost agent or a security agent or a deployment agent. The multicloud extension keeps it inside the SRE / incident-response pillar. Platform-engineering teams whose work crosses cost, security, compliance, deployment, and SRE need a coordination layer above the single-pillar surface. The active-operational-layer shape — a team of specialized agents on a shared fabric — is the structural answer.

2. Skills portability across vendors. AWS DevOps Agent supports custom skills today. NeuBird’s FalconClaw shipped a skills hub on April 6. ServiceNow’s AI Control Tower ships 300+ pre-built skills on the ServiceNow AI Platform. Red Hat Ansible Automation Platform 2.7 (announced May 12 at Red Hat Summit 2026) ships an MCP server with an automation orchestrator that combines deterministic, event-driven, and AI-driven automation in a single workflow canvas. Each vendor has its own skills surface. A platform team that codifies operational knowledge as skills inside one vendor’s surface accepts vendor lock-in by default. The active-operational-layer shape — where skills live in the customer’s repository and are portable across agent vendors that respect the format — keeps the customer’s operational knowledge as the customer’s asset.

3. BYOK economics on the model keys. Hyperscaler-native agents are priced with the model spend bundled into the orchestration price. Customer pays Microsoft / AWS / Google a single bill that includes the inference cost. The economic model is fine for customers who have not negotiated direct LLM commitments. The model is structurally expensive for customers who have. Enterprise customers in 2026 frequently have direct spend with Anthropic / OpenAI / their model provider; bundled inference is a hidden markup. BYOK on model keys is the structural answer, and it is the second differentiation surface for non-hyperscaler products.

4. Practitioner-grade interface inside the existing conversational client. AWS DevOps Agent ships with its own console surface and Slack / Microsoft Teams integration. The interface is fine. The pattern that produced NeuBird’s 74-versus-39 practitioner gap (74% of C-suite executives believe their organizations are actively using AI to manage incidents, only 39% of practitioners agree, per NeuBird’s 2026 State of Production Reliability and AI Adoption Report) was platform-side interface friction. Surfaces that show up natively in the conversational client the practitioner already uses (Claude Code, Cursor, the internal Mattermost on-call channel, the Slack pager bridge) are the third differentiation pillar.

The fourth is the active-operational-layer shape itself — a team of specialized agents on a shared fabric, doing the work, with capability-tier governance and an immutable audit trail on every action.

What This Means for the Hyperscaler-Native Agents

The May 13 AWS announcement is not bad news for AWS DevOps Agent. It is the natural product trajectory for a hyperscaler-native agent in a market where customers are multicloud by default. The InfoQ coverage of the AWS Frontier Agents launch in 2026 noted that “the DevOps Agent is your always-available operations teammate that resolves and proactively prevents incidents, optimizes application reliability and performance, and handles on-demand SRE tasks across AWS, multicloud, and on-prem environments.” The product positioning was already cross-cloud at GA; the May 13 announcement put the production-ready Azure and on-prem capability on the shipping calendar.

Microsoft will ship the matching capability for Azure SRE Agent (currently AWS / Azure-native, with on-prem on the published roadmap). Google Cloud will ship the matching capability for Gemini Cloud Assist’s incident-resolution surface. The three hyperscaler-native agents will converge on cross-cloud + on-prem in the next 12-18 months.

The structural read for platform-engineering teams is that the hyperscaler-native agents will be excellent at incident-response inside their own observability stacks, and serviceable at incident-response across other clouds and on-prem. The cross-pillar work, the skills-portability work, the BYOK work, and the practitioner-grade-interface work will live on a different surface.


See the IAN team run on your cloud. We connect to your AWS account via a scoped read-only role, run the Observe-tier agents, and leave you with a concrete audit report — cost waste, security exposure, compliance gaps, and a labor-offset estimate. You keep the findings regardless of next steps. Get a free infrastructure audit →


The 2026 SRE-Agent Posture for Mid-Market Platform Teams, Revisited

Re-read the five capabilities from batch-twenty-four’s vendor-map post against the May 13 announcement:

1. Cross-vendor skills compatibility. Reinforced by the announcement. AWS DevOps Agent custom skills, NeuBird FalconClaw skills, ServiceNow Control Tower skills, Ansible Automation Platform 2.7 skills, and the customer’s own skills should be portable across surfaces. The standard is the format, not the vendor.

2. Capability-tier classification on every action. Observe-tier (incident detection, anomaly correlation, runbook retrieval) runs automatically. Operate-tier (scoped remediation) is pre-authorized once. Administer-tier (IAM changes, billing-account actions) always requires explicit human approval. The tier model is the same across single-cloud, multicloud, and hybrid surfaces.

3. BYOK on model keys. More urgent now. With hyperscaler-native agents running cross-cloud, the inference-cost question becomes “whose model is doing the work, and who is paying for it.” Direct customer spend on Anthropic / OpenAI keys is the transparent answer.

4. Immutable audit trail in the customer’s database. The cross-cloud and hybrid surface multiplies the audit-trail-capture surface. Every incident across AWS / Azure / on-prem needs to land in the customer’s per-tenant audit-trail store, not the vendor’s. The audit trail is the reconciliation artifact for the post-incident review, the SRE blameless retro, the security audit, and the cyber-insurance claim.

5. Practitioner-grade interface inside the conversational client. Unchanged. The agent shows up in the same Slack channel, Claude Code session, Cursor IDE, or internal Mattermost where the practitioner already works.

How IAN Helps: The Incident / SRE Agent on the Active Operational Layer

IAN is the AI DevOps team for cloud infrastructure, delivered as a coordinated team of specialized agents on the active operational layer. The incident / SRE agent inside the IAN team is built for the post-May-13-2026 vendor-map shape:

  • Cross-cloud, hybrid-cloud, and on-premises by design. The SRE agent reads from CloudWatch, Stackdriver, Azure Monitor, the OpenTelemetry pipeline, Prometheus on-prem, and any MCP-exposed monitoring surface across every connected cloud account and on-prem location. There is no “primary cloud” — every cloud and every on-prem surface is a peer.
  • Cross-pillar coordination as the default. When the SRE agent investigates an incident, it reads from the cost-agent’s recent-rightsizing history, the resource-agent’s tagging-and-lifecycle changes, the security-agent’s drift detection, and the deployment-agent’s release log. Cross-pillar context is the default, not a separate integration.
  • OpenClaw-style skills layer. The customer’s operational runbooks are codified as skills inside the customer’s repository, executable by the SRE agent, and portable to other agents that respect the same skill format. AWS DevOps Agent custom skills, NeuBird FalconClaw skills, ServiceNow Control Tower skills, Ansible Automation Platform 2.7 skills should all be addressable from the same skills layer in the customer’s repo over time.
  • Capability-tier governance. Observe-tier (incident detection, runbook retrieval, investigation report generation) is automatic. Operate-tier (scoped pod restart, node-group scale, known-safe configuration apply, scoped credential rotation) is pre-authorized once. Administer-tier (IAM changes, billing-account actions) always requires explicit approval.
  • BYOK on model keys. Customer pays inference costs directly to Anthropic / OpenAI / their model provider. IAN charges for orchestration. Pricing is usage-based on agent actions, with a monthly minimum.
  • Immutable audit trail in the customer’s database. Every alert, every investigation, every remediation, every approval gate, every credential rotation across AWS / Azure / GCP / on-prem lands in the customer’s per-tenant audit-trail store.
  • Practitioner-grade interface inside the conversational client. The SRE agent shows up in the same Slack channel, Claude Code session, Cursor IDE, or internal Mattermost where the practitioner already works.

The Three-Phase Rollout

Phase 1 — Observe the multicloud + on-prem incident-response surface. Connect the SRE agent to the existing observability and alerting fabric across every cloud and every on-prem location. Run the Observe pass against the last 90 days of incidents. Surface the patterns — top recurring incidents, MTTR distribution, root-cause-category distribution, runbook-coverage gap across the entire surface. Two-to-four weeks.

Phase 2 — Codify the scoped remediation set per surface. Pre-authorize the Operate-tier scope per cloud and per on-prem location. The pre-authorization is the practitioner’s input, not the vendor’s. Two-to-three months.

Phase 3 — Cross the SRE / cost / security / compliance agent loop. SRE-agent investigations feed cost-agent rightsizing context. Cost-agent rightsizing events feed SRE-agent incident hypothesis space. Security-agent drift detection feeds SRE-agent investigation context. Compliance-agent reconciliation reads from every pillar. The agents become a coordinated team across the multicloud + on-prem surface.

The May 13 AWS announcement is one milestone on the path to a fully cross-cloud + on-prem AI-SRE category. The active-operational-layer shape is what stays distinct as the category matures: a coordinated team of specialized agents, skills portability, BYOK economics, capability-tier governance, an immutable audit trail, and a practitioner-grade interface — the structural answer to a 2026 vendor map that is no longer organized by cloud boundary.


Get a free infrastructure audit → | See pricing →

Next step: talk to the team

30 minutes. We'll look at your cloud together and scope what we'd take off your plate — see pricing.

Related Posts

');">
Devops

DevOpsCon London 2026: agentic developer platforms arrive

DevOpsCon London 2026 was the first major platform-engineering event to explicitly position agentic developer platforms as the successor to the DevOps tool stack — AI agents as first-class citizens with RBAC, quotas, and policies.

May 26, 2026 · 13 min