agent-evaluation
Design and implement comprehensive evaluation systems for AI agents. Use when building evals for coding agents, conversational agents, research agents, or computer-use agents. Covers grader types, benchmarks, 8-step roadmap, and production integration.
agent-evaluationevalsAI-agentsbenchmarks
autohandai
103
go-testing
>-
gogolangtestingtable-driven-tests
marsolab
0