Agent skill · magnus919
resilience-and-recovery
Design, exercise, and evidence graceful degradation, disaster recovery, and restoration behavior across systems and dependencies. Covers failure-mode analysis, RTO/RPO decision records, restore testing, game days, failover drills, data integrity verification, and recovery communication. Do not use for live incident command or incident response; route to site-reliability-engineering for those. Do not use for infrastructure implementation details; route to platform-engineering.
What it needs
About 8k tokens when loaded.
What this skill does
Resilience and Recovery Design, exercise, and evidence resilience and recovery behavior across systems and their dependencies. This skill joins failure modes, dependency behavior, degradation choices, restore testing, disaster recovery, game days, failover, data integrity, and recovery communication into a single method — producing an exercise-backed resilience plan, not only a design document. When to use Trigger What it covers --- --- "Design a resilience plan for this system" Failure-mode mapping, dependency analysis, degradation choices, RTO/RPO decision record, recovery plan template "Run a game day or restore test" Exercise design, scenario definition, evidence recording, follow-up work ledger "Assess our disaster recovery readiness" DR plan review against exercise evidence, gap analysis, data integrity verification "What happens if this dependency fails?" Dependency-loss scenarios, degradation paths, circuit-breaker and fallback strategy "Define our RTO and RPO" Context-specific decision record with tradeoff analysis, not universal prescription "Verify data integrity after a restore" Post-restore validation procedures, checksum and consistency checks, reconciliation protocol "Plan a failover drill" Failover exercise design, pre-conditions, success criteria, rollback/failback plan, evidence recording When not to use Live incident command or incident response: route to site-reliability-engineering for the incident command system, on-call operations, and real-time incident management. This skill owns the pre-incident resilience design and exercise-evidence method; SRE owns the live response. Infrastructure implementation details: route to platform-engineering for infrastructure-as-code, CI/CD pipeline implementation, container orchestration, and service networking. This skill owns the resilience requirements and exercise evidence; platform engineering owns the implementation. …
How to use it
Reference it in AdaL, Claude Code, Cursor or any coding agent — nothing to install:
@skills magnus919/resilience-and-recovery