Evals Run Skill
SkillAI & modelsWorkflow skill for evals run.
Available today. Use it from your connected AI after setup.
No other account needed.
Connect ahel once, and every AI you use reads what you have installed.
Then ask your AI: use the Evals Run Skill skill
What this skill tells your AI
The instructions your AI receives, as published by hoangnguyen0403/agent-skills-standard in .codex/skills/evals-run/SKILL.md and read by ahel’s review.
[!IMPORTANT] Workflow skill for evals run.
Optional args: slug=, ticket=<id/url>, mode=interactive|autonomous|channel, channel=, auto_continue=true|false, profile=business|hybrid|technical.
Instructions
When the user asks to perform this workflow, execute the following steps:
description: Run blinded live skill evals and publish reproducible v2 results.
Goal
Measure whether a skill changes agent behavior with isolated, immutable, outcome-based eval evidence.
Steps
1. Choose or resume a run
-
For ordinary maintenance after a complete catalog baseline exists, run
pnpm evals:baselinefirst. It creates or resumes a selective manifest, reuses only compatible evidence, and prints the model, reasoning level, concurrency, and fresh-answer count without starting workers. -
Review that plan before spending quota. Start workers only with
pnpm evals:baseline -- --execute; the default isgpt-5.6-lunawithhighreasoning and one worker. Override intentionally withEVALS_MODEL,EVALS_REASONING_EFFORT, orEVALS_CONCURRENCY(maximum four workers). -
If usage is exhausted, keep the run directory and rerun the identical
--executecommand after access resumes; completed answers are reused automatically. -
Use
pnpm evals:manifest -- --category <category>for one category orpnpm evals:manifest -- --allfor the complete catalog. -
Use
pnpm evals:manifest -- --resume <runId>only when deliberately continuing an existing run; a new invocation always creates a collision-safe run ID. -
Record the printed run ID. The manifest records source hashes, the v2 schema, and the generation protocol.
2. Answer each blinded case
- Run each baseline and with-skill arm in a separate worker/context.
- Baseline receives only the prompt. With-skill receives the same prompt plus that skill's
SKILL.md. - Trigger cases receive only the skill name and one-line description; never open the full skill body or expose the expected label.
- Trigger prompt filenames use opaque case IDs; never infer the expected label from filenames or ordering.
- For
allruns, write answers underanswers/<category>/<skill>/<case>; category runs useanswers/<skill>/<case>. - Mark known compromised baselines in the manifest and do not use them for delta calculations until clean reruns replace them.
3. Complete and score
- Fill
metadata.agent,metadata.model, andmetadata.completedAtafter every required answer exists. - Run
pnpm evals:score -- --run <runId>. - Scoring refuses to write
results.jsonwhile any arm is pending, verifies source hashes, and writes one immutableinputs.jsonsnapshot before publishing v2 results.
4. Report and verify
- Run
pnpm evals:reportto project aggregate runs into the newest complete category partitions and update physical history/archive records. - Run
pnpm evals:verify -- --run <runId>and, before handoff,pnpm evals:verify -- --all. - Confirm case pass rate, assertion pass rate, trigger recall, trigger specificity, and balanced trigger accuracy. Treat baseline and delta as
n/afor compromised arms. - Never hand-edit
results.json, transcripts, history, or archives. Fix inputs or eval definitions and regenerate.
Output
Run Summary
Evidence
Known Risks
Outcome Report
feature_status: implemented | partially_implemented | blocked requirement_trace: manifest -> inputs -> results -> report -> verification completed_evidence: [] missing_evidence: [] decision_needed: [] recommended_next_workflow: verify-work
Signals
- GitHub stars
- 565
- Forks
- 164
- Last commit
- Sep 2026
Advanced
- Catalog kind
- skill
- Gateway key
evals-run- Source
- github.com/hoangnguyen0403/agent-skills-standard