.fyi
SkillsMCPPluginsSubagents

Browse by category

DevOps & CI/CD SkillsProductivity & Workflow SkillsOther SkillsProduct & Project Management SkillsDocumentation & Knowledge SkillsCode Review & Refactor SkillsBackend & APIs SkillsAgent Meta & Communication SkillsResearch SkillsSecurity SkillsUX UI & Design SkillsTesting & QA SkillsSee all →

Every Claude Code skill, MCP server, plugin and subagent in one directory. Searchable, comparable, and one command from installed. Live stats from GitHub, npm and PyPI.

We're on Product HuntYour agent's app storeCheck it out →
Agent SkillsMCP ServersPluginsSubagentsCoding Agents
CollectionsOfficial publishersGlossaryFAQBlogSearchSavedFeedback
PrivacyTermsllms.txtSitemap

made with ♥ · © 2026 aaaa.fyi

Independent project · real data from public registries

…/agentscope-ai/openjudge
home/skills/agentscope-ai/openjudge
agentscope-ai avatar

agentscope-ai/openjudge

18 skills · 84 total installs

View on GitHub
$npx skills add agentscope-ai/openjudge
SkillInstalls
paper-reviewReview academic papers for correctness, quality, and novelty using OpenJudge's multi-stage pipeline.16find-skills-comboDiscover and recommend **combinations** of agent skills to complete complex, multi-faceted tasks.13claude-authenticityDetect whether an API endpoint is backed by genuine Claude (not a wrapper, proxy, or impersonator) using 9 weighted rule-based checks that mirror the…12auto-arenaAutomatically evaluate and compare multiple AI models or agents without pre-existing test data.11bib-verifyVerify a BibTeX file for hallucinated or fabricated references by cross-checking every entry against CrossRef, arXiv, and DBLP.11ref-hallucination-arenaBenchmark LLM reference recommendation capabilities by verifying every cited paper against Crossref, PubMed, arXiv, and DBLP.9openjudgeBuild custom LLM evaluation pipelines using the OpenJudge framework.8mmx-cliGenerate text, images, video, speech, and music via the MiniMax AI platform.2rl-rewardBuild RL reward signals using the OpenJudge framework. Covers choosing between pointwise and pairwise reward strategies based on RL algorithm, task type, and…200-meta-evalUse when the user wants to build an evaluation system for an LLM/agent application but doesn't know where to start — they have traces, prompts, RAG pipelines,…—01-eval-designUse when the user needs to design evaluation datasets, create test cases, stratify samples, generate adversarial examples, extract eval dimensions from…—02-metric-designUse when the user has evaluation principles or a dataset but needs help choosing the right graders, designing evaluation metrics, creating LLM-as-judge…—03-align-humanUse when the user has a judge/grader and human-labeled data, and wants to measure how well the judge agrees with humans, detect systematic biases, determine…—04-eval-reportUse when the user has run multiple evaluation skills and wants a comprehensive analysis — maturity assessment, cross-skill signals, trends, prioritized…—05-rag-evalUse when the user has a RAG (Retrieval-Augmented Generation) system and wants to evaluate its quality — separating retrieval issues from generation issues.—06-prompt-regressionUse when the user has changed a prompt (system prompt, RAG template, agent instruction, etc.) and wants to know whether the candidate is better or worse than…—07-redteamUse when the user wants to test their LLM/agent application for safety and security vulnerabilities — jailbreaks, prompt injection, PII extraction, harmful…—08-bootstrapUse when the user has nothing — no traces, no labels, no eval set — and needs to build a v0 evaluation from scratch.—