Agent skill · zechenzhangagi

nemo-evaluator-sdk

Evaluates LLMs across 100+ benchmarks from 18+ harnesses (MMLU, HumanEval, GSM8K, safety, VLM) with multi-backend execution. Use when needing scalable eval

Reference it in any coding agent with:

@skills zechenzhangagi/nemo-evaluator-sdk

Browse the @skills marketplace