Agent Skills: ai-evaluation-suite
Comprehensive AI/LLM evaluation toolkit for production AI systems. Covers LLM output quality, prompt engineering, RAG evaluation, agent performance, hallucination detection, bias assessment, cost/token optimization, latency metrics, model comparison, and fine-tuning evaluation. Includes BLEU/ROUGE metrics, perplexity, F1 scores, LLM-as-judge patterns, and benchmarks like MMLU and HumanEval.
UncategorizedID: majiayu000/claude-skill-registry/ai-evaluation-suite
57788
Install this agent skill to your local
Skill Files
Browse the full folder contents for ai-evaluation-suite.
Loading file tree…
Select a file to preview its contents.