Agent skill · refoundai

ai-evals

Help users build robust infrastructure for measuring, monitoring, and iterating on AI product performance using human, code-based, and LLM-as-a-judge metho

Reference it in any coding agent with:

@skills refoundai/ai-evals

Browse the @skills marketplace