Agent skill · elastic

observability-service-reliability

Design and operate service reliability targets in Elastic Observability: choose an SLI type and a defensible target, pick a time window and budgeting method, create and maintain SLOs through the Kibana API, attach burn-rate alert rules, and decide when an SLO is the wrong instrument and a threshold rule, anomaly job, or synthetics monitor is right. Use when defining or reviewing SLOs and error budgets, tuning burn-rate alerting, reducing alert noise, or setting up availability monitoring for a user-facing endpoint.

What it needs

About 11k tokens when loaded.

What this skill does

Service Reliability Design reliability targets that people will actually act on, then operate them. This skill covers the judgment before the API call — which service-level indicator fits the data you have, what target is achievable rather than aspirational, whether an SLO is even the right instrument — and then the mechanics of creating, alerting on, resetting, and retiring SLOs through the Kibana API. Reliability instruments are not interchangeable. An SLO measures a user-visible outcome against a spendable budget; a threshold rule fires on a raw condition; an anomaly job finds deviations where no fixed threshold exists; a synthetics monitor is the only one of the four that can see a service that has stopped emitting telemetry entirely. Choosing wrong produces alerts that are technically correct and operationally useless. For diagnosing a service that is already degraded, and for the incident workflow itself, use the observability-sre-triage skill; for general rule lifecycle mechanics use the kibana-alerting-rules skill. <!-- begin-partial: preamble --> Environment Configuration This skill executes Elasticsearch operations through the elastic CLI. If the elastic CLI is not installed, tell the user what it is needed for. Do not guess credentials, call the HTTP API directly, or attempt other workarounds. This skill references operations in HTTP-shorthand form (e.g., GET /, GET /cat/indices, GET /{index}/mapping, GET /{index}/settings/index.mode, POST /query). The Operations table at the end of this document maps each shorthand to the equivalent elastic CLI command — always use the CLI rather than calling the HTTP API directly. <!-- end-partial: preamble --> Analysis without cluster access The CLI check above gates querying the cluster — it does not gate analysis. When the user has already supplied the evidence in their question (metric values, counts, status reasons, log lines, alert payloads, configuration), reason from that evidence and deliver the conclusion. …

How to use it

Reference it in AdaL, Claude Code, Cursor or any coding agent — nothing to install:

@skills elastic/service-reliability--05eb7d

View the source on GitHub

Browse the @skills marketplace