decompose-evaluation-metric
SkillMonitoring & opsLets your agent break down an evaluation metric into its signals, aggregation rules, and gaming risks.
Use decompose-evaluation-metric in Claude, ChatGPT or Ahel Desktop
Free. Sign in, add decompose-evaluation-metric and connect your AI. About a minute.
Also: Claude Code · Cursor · Codex
Then ask your AI: use the decompose-evaluation-metric skill
Details
Instructions available. Your AI can read the instructions. Execution depends on the setup they require.
Account requirements not reviewed. Check the skill instructions before use; Ahel provides instructions and does not run this skill.
No other account needed.
Add Ahel to your AI once: Claude, ChatGPT, Cursor, Claude Code or Codex. Then ask it to use this.
About this skill
Decompose an evaluation metric into rewarded signals, aggregation choices, polarity, ceiling effects, and Goodhart vulnerabilities.
What this skill tells your AI
The instructions your AI receives, as published by yogsoth-ai/de-anthropocentric-research-engine in skills/decompose-evaluation-metric/SKILL.md and read by Ahel’s review.
Purpose
Decompose an evaluation metric into rewarded signals, aggregation choices, polarity, ceiling effects, and Goodhart vulnerabilities.
Input contract
required: [metric_definition, scored_outputs]
optional: [reference_standard, aggregation_rule, known_failure_cases]
constraints: [each component must have a declared direction and interpretation]
Procedure
- Split the metric into primitive signals and aggregation operations.
- Record polarity, scale, weighting, normalization, and ceiling/floor behavior.
- Map rewarded shortcuts and construct-irrelevant incentives.
- State interpretation limits and diagnostic needs.
If metric components are explicit but their link to the intended construct remains uncertain, consider assess-construct-validity as the next tactic.
Output contract
produces: [metric_components, aggregation_map, polarity_and_scale, ceiling_analysis, goodhart_risks]
delta_fields: [findings, evidence_updates, uncertainties, open_questions]
Quality gates
- Component contributions and aggregation are reconstructible.
- A high score is not treated as capability evidence without construct support.
Failure and counterexamples
Do not infer metric meaning from its name or ignore nonlinear aggregation and clipping.
Provenance map
resolved: metric-decomposition
Signals
- GitHub stars
- 503
- Forks
- 42
- Last commit
- Sep 2026
Advanced
- Item type
- skill
- Key
decompose-evaluation-metric- Source
- github.com/yogsoth-ai/de-anthropocentric-research-engine
github.com/yogsoth-ai/de-anthropocentric-research-engine
More in Monitoring & ops
Skill · anthropics
More in Monitoring & opsagent-eval
Skill · affaan-m
More in Monitoring & opspricing
Skill · coreyhaines31
More in Monitoring & opslark-okr
Skill · larksuite
More in Monitoring & opsdashboard-builder
Skill · affaan-m
More in Monitoring & opsbabysit
Skill · thedotmack
More in Monitoring & ops