Agent skill · vasilyu1983

ai-coding-agents-observability-evals

Designs coding-agent observability and evals. Use when measuring traces, replay, checkpoint lineage, quality trajectories, tool grading, regression, or cost.

What it needs

About 11k tokens when loaded.

What this skill does

AI Coding Agents Observability And Evals Use this skill to design or review the feedback loop around a coding-agent runtime: traces, replayable transcripts, eval packs, regression gates, tool-call grading, latency and cost accounting, and production failure triage. This skill covers how you operate a coding-agent product after the core runtime exists. It does not replace the runtime skills themselves. ASCII Flow Quick Reference Question Read Outcome ---------- ------ --------- What should the trace and telemetry model include? references/trace-and-telemetry-model.md Durable trace schema, session correlation, event stages, and replay boundaries How should evals, regressions, and cost controls work? references/evals-regression-and-cost-ops.md Golden tasks, iterative self-extension packs, trajectory scorecards, and cost-aware release gates How do I use the eval/trace substrate to improve the harness itself? references/harness-self-evolution.md Closed-loop harness evolution: three observability pillars, falsifiable-contract edits, attribution How does OpenAI Codex combine rollout replay, SQLite state, doctor reports, and telemetry? references/openai-codex-rollout-doctor-telemetry.md Replay artifacts, rebuildable state indexes, redacted diagnostics, W3C traces, token metrics How does Codex wire OTel exporters and what analytics events exist? references/openai-codex-otel-config.md OtelSettings TOML schema, exporter selection, W3C tracestate, contrast with proprietary analytics events When To Use Design tracing and replay for a coding-agent CLI Add regression evals for coding, review, or task-execution agents Evaluate whether a coding agent preserves correctness and structural quality while extending its own workspace across evolving specifications Grade tool calls, patch quality, verification behavior, or handoff quality Build latency, token, and cost accounting for agent sessions Review how incidents and bad runs should be debugged from stored traces Use Other Skills Nee …

How to use it

Reference it in AdaL, Claude Code, Cursor or any coding agent — nothing to install:

@skills vasilyu1983/ai-coding-agents-observability-evals

View the source on GitHub

Browse the @skills marketplace