advanced-evaluation

SkillAI & models

Helps your agent build LLM-as-judge evaluations, scoring rubrics, and model output comparisons.

This skill should be used when the user asks to "implement LLM-as-judge", "compare model outputs", "create evaluation rubrics", "mitigate evaluation bias", or mentions direct scoring, pairwise comparison, position bias, evaluation pipelines, or automated quality assessment.

Serves today. ahel delivers it to your agents as a prompt through your gateway link.

Serve it through your gateway

One link, every agent. Your own credentials, stored once.

Signals

GitHub stars
18k
Forks
1k
Last commit
Aug 2026
Installs
17k stars