Agent Skills: Evals
Agent evaluation framework based on Anthropic's best practices. USE WHEN eval, evaluate, test agent, benchmark, verify behavior, regression test, capability test. Includes three grader types (code-based, model-based, human), transcript capture, pass@k/pass^k metrics, and ALGORITHM integration.
UncategorizedID: majiayu000/claude-skill-registry/Evals
11819
Install this agent skill to your local
Skill Files
Browse the full folder contents for Evals.
Loading file tree…
Select a file to preview its contents.