bymoltis-org· 59 skills
Evaluates LLMs across 60+ academic benchmarks (MMLU, HumanEval, GSM8K, TruthfulQA, HellaSwag). Use when benchmarking model quality, comparing models, reporting academic results, or tracking training progress. Industry standard used by EleutherAI, HuggingFace, and major labs. Supports HuggingFace, vLLM, APIs.
$npx -y skills add moltis-org/moltis --skill evaluating-llms-harnessInstalls into the current project.
Run `npx skills use "https://github.com/moltis-org/moltis" --skill "moltis-org/moltis/evaluating-llms-harness"` and follow the generated skill instructions now. Read its complete output, redirecting it to a temporary file first if necessary. Resolve relative paths from the supporting-files directory it provides.
Use the skills in "https://github.com/moltis-org/moltis" that are relevant to the current task. Run `npx skills add "https://github.com/moltis-org/moltis"` and select the relevant skills, then follow their instructions.