Agent Skills: Evals

Agent evaluation framework based on Anthropic's best practices. USE WHEN eval, evaluate, test agent, benchmark, verify behavior, regression test, capability test. Includes three grader types (code-based, model-based, human), transcript capture, pass@k/pass^k metrics, and ALGORITHM integration.

UncategorizedID: majiayu000/claude-skill-registry/Evals

Author

majiayu000

https://github.com/majiayu000 View all skills

Repository

majiayu000/claude-skill-registry

majiayu000License: MIT

11819

Install this agent skill to your local

pnpm dlx add-skill https://github.com/majiayu000/claude-skill-registry/Evals

Skill Files

Browse the full folder contents for Evals.

Download Skill

Loading file tree…

Select a file to preview its contents.