Result Auditor

SkillProductivity

Validate and scientifically interpret completed research runs, including negative, null, boundary, confounded, failed, and inconclusive outcomes. Use after execution to assess evidence and claim impact; not to edit methods or schedule experiments.

Use Result Auditor in Claude, ChatGPT or Ahel Desktop

Free. Sign in, add Result Auditor and connect your AI. About a minute.

Also: Claude Code · Cursor · Codex

Then ask your AI: use the Result Auditor skill

Details

Instructions available. Your AI can read the instructions. Execution depends on the setup they require.

Add Ahel to your AI once: Claude, ChatGPT, Cursor, Claude Code or Codex. Then ask it to use this.

Result AuditorStart free

What this skill tells your AI

The instructions your AI receives, as published by cpt13-g/build-research-evidence in .agents/skills/result-auditor/SKILL.md and read by Ahel’s review.

Read research/research_state.yaml, the Card, contract, Runner manifest, raw metrics/logs, docs/PROTOCOLS.md, and docs/ROLE_HANDOFFS.md. For Candidate or refinement results, also read docs/METHOD_REALIZATION_PROTOCOL.md; for post-P0 aggregation read docs/EVIDENCE_COMPLETION_PROTOCOL.md. Perform the first pass without the Orchestrator's interpretation or preferred story. Fix the initial validity, fairness and result verdict before reconciliation. Record the reviewed inputs, deterministic checks and any isolation exception. Never interpret an execution failure as evidence against the scientific hypothesis.

Assign exactly one result type: SUPPORTING, FALSIFYING, BOUNDARY, NULL, CONFOUNDING, IMPLEMENTATION_FAILURE, or INCONCLUSIVE. Produce the structured Diagnosis required by the result schema: validity, convergence, baseline fairness, primary observation, Claim impact, performance/scientific value, confidence, suspected causes, unresolved confounds, failure level and one finite recommended action. Check seed/statistical protocol, data leakage and missing artifacts.

Use F1 only for execution failures and recommend Debug/Tuning. Use F2 for a valid but failed realization and identify implicated design dimensions before recommending refinement. Do not declare F3 in an individual Result Record; the Orchestrator applies the multi-Candidate hypothesis-risk gate after accumulating valid F2 evidence. A NULL result must retain explicit scientific value. Cheap results remain PRELIMINARY_NON_PAPER regardless of Metric.

Write a schema-valid result record linked to the immutable run manifest and preserve meaningful outcomes in experiment memory. Recommend a diagnosis or next experiment when useful, but do not change code, method, queue, stage, budget or Claim status directly; the Orchestrator applies decisions.

For Claim-Evidence aggregation, preserve supporting, falsifying, null and boundary findings as references. Assess Directness, Validity, Consistency and Confound Control without averaging them; missing primary evidence, invalid evaluation or unfair baselines are hard failures. Cheap and Tuning results remain outside formal evidence.

Signals

GitHub stars
34
Last commit
Sep 2026
Advanced
Item type
skill
Key
result-auditor
Source
github.com/cpt13-g/build-research-evidence