About the service
Three products you can buy together or separately: on-call with schedules and escalations, incident response inside Slack and Teams, and an AI SRE agent that hunts for the cause of a failure. The agent starts the moment an alert fires and tests several hypotheses in parallel — telemetry, recent commits, feature-flag and configuration changes, the history of similar incidents — then returns ranked theories with confidence scores. It also correlates a deployment with a metric dip, assesses blast radius, suppresses noisy pages, and drafts remediation steps together with a candidate fix, but never applies anything without a human. Alongside it sit an AI scribe on incident calls, matching against similar past outages and auto-generated retrospectives. You can bring your own model keys, and the agent plugs into the rest of your stack — Datadog, GitHub, Jira — and can be driven through the API and MCP.