Comprehensive evaluation patterns for AI agents including multi-turn conversation testing, LLM-as-judge frameworks, benchmark suites, regression detection, and systematic eval pipelines for measuring agent quality and safety.
$npx -y skills add pramoddutta/qaskills --skill ai-agent-evalInstalls into the current project.
Run `npx skills use "https://github.com/pramoddutta/qaskills" --skill "pramoddutta/qaskills/ai-agent-eval"` and follow the generated skill instructions now. Read its complete output, redirecting it to a temporary file first if necessary. Resolve relative paths from the supporting-files directory it provides.
Use the skills in "https://github.com/pramoddutta/qaskills" that are relevant to the current task. Run `npx skills add "https://github.com/pramoddutta/qaskills"` and select the relevant skills, then follow their instructions.