@skills · Owner
Agent skills by hamelsmu
9 skills indexed from github.com/hamelsmu. Reference any of them in AdaL, Claude Code, Cursor or any coding agent — nothing to install.
- build-review-interface · Skill · 1,623 stars
Build a custom browser-based annotation interface tailored to your data for reviewing LLM traces and collecting structured feedback. Use when you need to b
- error-analysis · Skill · 1,623 stars
Help the user systematically identify and categorize failure modes in an LLM pipeline by reading traces. Use when starting a new eval project, after signif
- eval-audit · Skill · 1,623 stars
Audit an LLM eval pipeline and surface problems: missing error analysis, unvalidated judges, vanity metrics, etc. Use when inheriting an eval system, when
- evals-skills · Collection · 1,623 stars
- evaluate-rag · Skill · 1,623 stars
Guides evaluation of RAG pipeline retrieval and generation quality. Use when evaluating a retrieval-augmented generation system, measuring retrieval qualit
- generate-synthetic-data · Skill · 1,623 stars
Create diverse synthetic test inputs for LLM pipeline evaluation using dimension-based tuple generation. Use when bootstrapping an eval dataset, when real
- skills · Collection · 1,623 stars
- validate-evaluator · Skill · 1,623 stars
Calibrate an LLM judge against human labels using data splits, TPR/TNR, and bias correction. Use after writing a judge prompt (write-judge-prompt) when you
- write-judge-prompt · Skill · 1,623 stars
Design LLM-as-Judge evaluators for subjective criteria that code-based checks cannot handle. Use when a failure mode requires interpretation (tone, faithfu