Agent skill · magnus919

site-reliability-engineering

Design, operate, and improve reliable production systems with SLOs, incident command, observability, error budgets, and operational practices.

What it needs

About 4k tokens when loaded.

What this skill does

Site Reliability Engineering A comprehensive methodology for designing, operating, and improving reliable production systems. Rooted in Google SRE principles and extended with modern practices for incident command, observability engineering, error budget governance, and operational excellence. When to Load This Skill Trigger What It Means --- --- "Design reliability into this system" SLO/SLI framework, error budget policy, resilience architecture "Run an incident postmortem" Blameless postmortem with timeline, 5 Whys, action tracking "Improve our on-call" Rotation design, alert tuning, toil reduction, escalation policy "Build observability" The Four Golden Signals, dashboard design, alert rule patterns "Do a reliability review" Architecture review against SRE principles, risk assessment "I need an incident commander" Incident command framework, role cards, communication templates "Automate this operational task" Toil assessment, automation decision tree, runbook pattern "Adopt SRE in this organization" Engagement boundaries, maturity, team model, and change adoption "Review this reliability design" User journeys, dependencies, overload, configuration, canary, durability "Our SRE team is overloaded" Operational-load diagnosis, protected engineering time, recovery plan "Improve incident learning or sustainable on-call" Cognitive load, psychological safety, documentation, exercises When not to use Use release-engineering to plan releases, compose promotion and rollback gates, or coordinate a release train. Use systematic-debugging to find the cause of a specific failure. Operating the telemetry stack itself — Prometheus scrape configs, OpenTelemetry Collector pipelines, Loki ingest and retention, Prometheus rules files — belongs to telemetry; this skill owns the SLI/SLO and alert design those rules implement. Grafana product work — dashboards, panels, Grafana-side alert rules, contact points, notification policies — belongs to grafana. …

How to use it

Reference it in AdaL, Claude Code, Cursor or any coding agent — nothing to install:

@skills magnus919/site-reliability-engineering

View the source on GitHub

Browse the @skills marketplace