OpenJudge Skill
SkillAI & modelsLets your agent build pipelines that score and compare LLM outputs using configurable graders and batch evaluation.
Use OpenJudge Skill in Claude, ChatGPT or Ahel Desktop
Free. Sign in, add OpenJudge Skill and connect your AI. About a minute.
Also: Claude Code · Cursor · Codex
Then ask your AI: use the OpenJudge Skill
Details
Instructions available. Your AI can read the instructions. Execution depends on the setup they require.
Account requirements not reviewed. Check the skill instructions before use; Ahel provides instructions and does not run this skill.
No other account needed.
Add Ahel to your AI once: Claude, ChatGPT, Cursor, Claude Code or Codex. Then ask it to use this.
About this skill
Build custom LLM evaluation pipelines using the OpenJudge framework. Covers selecting and configuring graders (LLM-based, function-based, agentic), running batch evaluations with GradingRunner, combining scores with aggregators, applying evaluation strategies (voting, average), auto-generating grade
What this skill tells your AI
The instructions your AI receives, as published by agentscope-ai/openjudge in skills/openjudge-core/01-graders-and-pipeline/SKILL.md and read by Ahel’s review.
Build evaluation pipelines for LLM applications using the openjudge library.
When to Use This Skill
- User wants to evaluate LLM output quality (correctness, relevance, hallucination, etc.)
- User wants to compare two or more models and rank them
- User wants to design a scoring rubric and automate evaluation
- User wants to analyze evaluation results statistically
- User wants to build a reward model or quality filter
Sub-documents — Read When Relevant
| Topic | File | Read when… |
|---|---|---|
| Grader selection & configuration | graders.md | User needs to pick or configure an evaluator |
| Batch evaluation pipeline | pipeline.md | User needs to run evaluation over a dataset |
| Auto-generate graders from data | generator.md | No rubric yet; generate from labeled examples |
| Analyze & compare results | analyzer.md | User wants win rates, statistics, or metrics |
Read the relevant sub-document before writing any code.
Install
pip install py-openjudge
Architecture Overview
Dataset (List[dict])
│
▼
GradingRunner ← orchestrates everything
│
├─► Grader A ──► EvaluationStrategy ──► _aevaluate() ──► GraderScore / GraderRank
├─► Grader B ──► EvaluationStrategy ──► _aevaluate() ──► GraderScore / GraderRank
└─► Grader C ...
│
├─► Aggregator (optional) ← combine multiple grader scores into one
│
└─► RunnerResult ← {grader_name: [GraderScore, ...]}
│
▼
Analyzer ← statistics, win rates, validation metrics
5-Minute Quick Start
Evaluate responses for correctness using a built-in grader:
import asyncio
from openjudge.models.openai_chat_model import OpenAIChatModel
from openjudge.graders.common.correctness import CorrectnessGrader
from openjudge.runner.grading_runner import GradingRunner
# 1. Configure the judge model (OpenAI-compatible endpoint)
model = OpenAIChatModel(
model="qwen-plus",
api_key="sk-xxx",
base_url="https://dashscope.aliyuncs.com/compatible-mode/v1",
)
# 2. Instantiate a grader
grader = CorrectnessGrader(model=model)
# 3. Prepare dataset
dataset = [
{
"query": "What is the capital of France?",
"response": "Paris is the capital of France.",
"reference_response": "Paris.",
},
{
"query": "What is 2 + 2?",
"response": "The answer is five.",
"reference_response": "4.",
},
]
# 4. Run evaluation
async def main():
runner = GradingRunner(
grader_configs={"correctness": grader},
max_concurrency=8,
)
results = await runner.arun(dataset)
for i, result in enumerate(results["correctness"]):
print(f"[{i}] score={result.score} reason={result.reason}")
asyncio.run(main())
Expected output:
[0] score=5 reason=The response accurately states Paris as capital...
[1] score=1 reason=The response gives the wrong answer (five vs 4)...
Key Data Types
| Type | Description |
|---|---|
GraderScore | Pointwise result: .score (float), .reason (str), .metadata (dict) |
GraderRank | Listwise result: .rank (List[int]), .reason (str), .metadata (dict) |
GraderError | Error during evaluation: .error (str), .reason (str) |
RunnerResult | Dict[str, List[GraderResult]] — keyed by grader name |
Result Handling Pattern
from openjudge.graders.schema import GraderScore, GraderRank, GraderError
for grader_name, grader_results in results.items():
for i, result in enumerate(grader_results):
if isinstance(result, GraderScore):
print(f"{grader_name}[{i}]: score={result.score}")
elif isinstance(result, GraderRank):
print(f"{grader_name}[{i}]: rank={result.rank}")
elif isinstance(result, GraderError):
print(f"{grader_name}[{i}]: ERROR — {result.error}")
Model Configuration
All LLM-based graders accept either a BaseChatModel instance or a dict config:
# Option A: instance
from openjudge.models.openai_chat_model import OpenAIChatModel
model = OpenAIChatModel(model="gpt-4o", api_key="sk-...")
# Option B: dict (auto-creates OpenAIChatModel)
model_cfg = {"model": "gpt-4o", "api_key": "sk-..."}
grader = CorrectnessGrader(model=model_cfg)
# OpenAI-compatible endpoints (DashScope / local / etc.)
model = OpenAIChatModel(
model="qwen-plus",
api_key="sk-xxx",
base_url="https://dashscope.aliyuncs.com/compatible-mode/v1",
)
Signals
- GitHub stars
- 868
- Forks
- 73
- Last commit
- Sep 2026
Ahel review
K1binfo
installs-packages
Automated review, not a security audit. Ruleset v1+k2.
Advanced
- Item type
- skill
- Key
x-01-graders-and-pipeline- Source
- github.com/agentscope-ai/openjudge
github.com/agentscope-ai/openjudge
Related picks
Skill · wshobson
The pick for Pythonpython-pro
Skill · jeffallan
The pick for Pythonskill-creator
Skill · anthropics
More in AI & modelstriage
Skill · mattpocock
More in AI & modelswayfinder
Skill · mattpocock
More in AI & modelsalgorithmic-art
Skill · anthropics
More in AI & models