decompose-evaluation-metric

SkillMonitoring & ops

Lets your agent break down an evaluation metric into its signals, aggregation rules, and gaming risks.

Use decompose-evaluation-metric in Claude, ChatGPT or Ahel Desktop

Free. Sign in, add decompose-evaluation-metric and connect your AI. About a minute.

Also: Claude Code · Cursor · Codex

Then ask your AI: use the decompose-evaluation-metric skill

Details

Instructions available. Your AI can read the instructions. Execution depends on the setup they require.

Add Ahel to your AI once: Claude, ChatGPT, Cursor, Claude Code or Codex. Then ask it to use this.

decompose-evaluation-metricStart free
About this skill

Decompose an evaluation metric into rewarded signals, aggregation choices, polarity, ceiling effects, and Goodhart vulnerabilities.

What this skill tells your AI

The instructions your AI receives, as published by yogsoth-ai/de-anthropocentric-research-engine in skills/decompose-evaluation-metric/SKILL.md and read by Ahel’s review.

Purpose

Decompose an evaluation metric into rewarded signals, aggregation choices, polarity, ceiling effects, and Goodhart vulnerabilities.

Input contract

required: [metric_definition, scored_outputs]
optional: [reference_standard, aggregation_rule, known_failure_cases]
constraints: [each component must have a declared direction and interpretation]

Procedure

  1. Split the metric into primitive signals and aggregation operations.
  2. Record polarity, scale, weighting, normalization, and ceiling/floor behavior.
  3. Map rewarded shortcuts and construct-irrelevant incentives.
  4. State interpretation limits and diagnostic needs.

If metric components are explicit but their link to the intended construct remains uncertain, consider assess-construct-validity as the next tactic.

Output contract

produces: [metric_components, aggregation_map, polarity_and_scale, ceiling_analysis, goodhart_risks]
delta_fields: [findings, evidence_updates, uncertainties, open_questions]

Quality gates

  • Component contributions and aggregation are reconstructible.
  • A high score is not treated as capability evidence without construct support.

Failure and counterexamples

Do not infer metric meaning from its name or ignore nonlinear aggregation and clipping.

Provenance map

  • resolved: metric-decomposition

Signals

GitHub stars
503
Forks
42
Last commit
Sep 2026
Advanced
Item type
skill
Key
decompose-evaluation-metric
Source
github.com/yogsoth-ai/de-anthropocentric-research-engine