detect-performance-discrepancy

SkillProductivity

Lets your agent compare benchmark scores for the same method across sources and explain why they differ.

Use detect-performance-discrepancy in Claude, ChatGPT or Ahel Desktop

Free. Sign in, add detect-performance-discrepancy and connect your AI. About a minute.

Also: Claude Code · Cursor · Codex

Then ask your AI: use the detect-performance-discrepancy skill

Details

Instructions available. Your AI can read the instructions. Execution depends on the setup they require.

Add ahel to your AI once: Claude, ChatGPT, Cursor, Claude Code or Codex. Then ask it to use this.

detect-performance-discrepancyStart free
About this skill

Detect material score discrepancies for the same method/task across sources and propose likely explanatory condition differences.

What this skill tells your AI

The instructions your AI receives, as published by yogsoth-ai/de-anthropocentric-research-engine in skills/detect-performance-discrepancy/SKILL.md and read by ahel’s review.

Purpose

Detect material score discrepancies for the same method or task across sources and identify plausible condition differences.

Input contract

required: [performance_records, method_key, task_key, metric_schema]
optional: [protocol_records, condition_schema, uncertainty_estimates]
constraints: [comparisons require aligned metric direction and declared conditions]

Procedure

  1. Align records by method, task, metric, and observation context.
  2. Quantify score differences with uncertainty and identify materially different pairs.
  3. Compare datasets, prompts, evaluators, budgets, and protocol conditions.
  4. Rank plausible explanations and retain unresolved alternatives.

Output contract

produces: [discrepancy_pairs, condition_difference_map, explanation_candidates, residual_uncertainties]
delta_fields: [findings, evidence_updates, uncertainties, open_questions]

Quality gates

  • Materiality uses a declared comparison basis.
  • Protocol mismatch is separated from method change.

Failure and counterexamples

Do not call rounding noise a discrepancy or infer a method improvement from non-equivalent evaluation conditions.

Provenance map

  • resolved: discrepancy-identification
  • resolved: discrepancy-analysis

Signals

GitHub stars
503
Forks
42
Last commit
Sep 2026
Advanced
Item type
skill
Key
detect-performance-discrepancy
Source
github.com/yogsoth-ai/de-anthropocentric-research-engine