Agent skill · vasilyu1983

ops-incident-response

Guides incident response from detection through postmortem. Use when designing on-call runbooks, triaging production incidents, writing status updates, or improving MTTD and MTTR.

What it needs

About 10k tokens when loaded.

What this skill does

Incident Response Production incident management: detect, triage, mitigate, resolve, learn. This skill covers the human and process side of incidents, not the infrastructure automation (use ops-devops-platform for that). Quick Reference Need Go to ------ ------- Classify the incident severity ## Severity Classification Run the end-to-end response loop ## Workflow Coordinate as incident commander ## Incident Commander Checklist Send status or stakeholder updates ## Communication Templates Use templates and runbooks ## Navigation Scope Use For Do NOT Use For --------- ---------------- Severity classification and escalation Infrastructure provisioning or CI/CD On-call runbook design and review Application resilience patterns (retries, circuit breakers) Incident commander workflows Monitoring/alerting tool setup Status page and stakeholder communication Root cause analysis of code bugs Blameless postmortem facilitation Performance benchmarking MTTD/MTTR measurement and improvement Security incident forensics (use software-security-appsec) Severity Classification Level Criteria Response time Who is paged ------- ---------- --------------- ------------- SEV1 Revenue-impacting, data loss, full outage Immediate On-call + incident commander + leadership SEV2 Degraded service, partial outage, SLO breach 15 min On-call + incident commander SEV3 Minor degradation, workaround available 1 hour On-call SEV4 Cosmetic, no user impact, internal tooling Next business day Ticket owner ASCII Flow Workflow Incident Commander Checklist [ ] Open dedicated incident channel [ ] Assign roles: IC, comms lead, subject matter experts [ ] Set a timer for status updates (every 15-30 min for SEV1/2) [ ] Keep a running timeline in the channel [ ] Delegate investigation — IC coordinates, does not debug [ ] Post status page updates at each phase transition [ ] Call "resolved" only when metrics confirm recovery [ ] Schedule postmortem within 48 hours Judgment Calls a Checklist Won't Make for You A chec …

How to use it

Reference it in AdaL, Claude Code, Cursor or any coding agent — nothing to install:

@skills vasilyu1983/ops-incident-response

View the source on GitHub

Browse the @skills marketplace