Experiment Runner

SkillMonitoring & ops

Implement and execute approved research Experiment, Tuning, or Design Cards with reproducible logs, metrics, environment, and artifacts. Use for bounded coding, debugging, and runs; not for changing claims, priorities, budgets, or the Research Spine.

Use Experiment Runner in Claude, ChatGPT or Ahel Desktop

Free. Sign in, add Experiment Runner and connect your AI. About a minute.

Also: Claude Code · Cursor · Codex

Then ask your AI: use the Experiment Runner skill

Details

Instructions available. Your AI can read the instructions. Execution depends on the setup they require.

Add Ahel to your AI once: Claude, ChatGPT, Cursor, Claude Code or Codex. Then ask it to use this.

Experiment RunnerStart free

What this skill tells your AI

The instructions your AI receives, as published by cpt13-g/build-research-evidence in .agents/skills/experiment-runner/SKILL.md and read by Ahel’s review.

Read research/research_state.yaml, the approved Card, research/experiment_contract.yaml, docs/PROTOCOLS.md, and docs/ROLE_HANDOFFS.md before execution. For Cheap Evaluation or refinement, also read docs/METHOD_REALIZATION_PROTOCOL.md. Refuse work before Gate 1 or a missing, draft, stale, budget-ineligible, dependency-blocked, phase-ineligible, priority-ineligible Card, and return the exact blocking fields.

Implement shared behavior in src/ and experiment differences in configs/. Preserve original data and checkpoints. Use tools/runner.py to emit the run manifest with command, Card/config identity, declared data and seed protocol, code/environment identity, timestamps, logs, raw metrics, exit status, artifact hashes, and explicit missing declarations. Do not describe a run as reproducible when its manifest is partial.

Engineering Debug and tuning are allowed within the Card. If a change affects a scientific mechanism, stop and request an approved Design Iteration. Never alter the Research Spine, Claim/EQ, priority, fairness protocol, budget or stop condition.

For cheap_evaluation, implement exactly the Card's proxy reductions and shared protocol across Candidates; label outputs preliminary and never promote them to paper evidence. For refinement_validation, modify only the approved Design Iteration dimensions and preserve the controlled comparison.

Stop once the Card completion contract is decided. Extra runs require an Orchestrator-approved named evidence gap; general completeness is insufficient. Return execution facts without a scientific verdict or the Orchestrator's desired interpretation. A failed run remains in experiment memory and goes to result-auditor when interpretable.

An Evidence Completion Queue entry is a recommendation, not execution authorization. Run it only after the Orchestrator creates and approves a normal Card and Registry activation succeeds. P3/SKIP entries are never default execution work.

Signals

GitHub stars
34
Last commit
Sep 2026
Advanced
Item type
skill
Key
experiment-runner
Source
github.com/cpt13-g/build-research-evidence