minimal-run-and-audit
SkillFiles & storageYour AI can run tasks and then audit the results, so finished work gets checked instead of being accepted as-is. minimal-run-and-audit is a skill in the dev-tools category. It is listed in the rigorpilot-skills repository on GitHub.
Available today. Use it from your connected AI after setup.
No other account needed.
After adding the skill, give your AI a task to run and ask it to audit the results once the task is done.
Then ask your AI: use the minimal-run-and-audit skill
What your AI can do with it
- Run tasks
- Audit the results of completed tasks
- Check its own work after finishing a task
What this skill tells your AI
The instructions your AI receives, as published by lllllllama/rigorpilot-skills in skills/minimal-run-and-audit/SKILL.md and read by ahel’s review.
Use this as the Rigor Run skill. The installed slug remains
minimal-run-and-audit for compatibility.
Use the shared operating principles in
../../references/agent-operating-principles.md; this skill should make run
evidence auditable without turning every command into a rigid protocol.
When to apply
- After a reproduction target and setup plan exist.
- When the main skill needs execution evidence and normalized outputs.
- When a smoke test, documented inference run, documented evaluation run, or other short non-training verification is appropriate.
- When the user already knows what command should be attempted and wants execution plus reporting only.
When not to apply
- During initial repo scanning.
- When environment or assets are still undefined enough to make execution meaningless.
- When the task is a literature lookup rather than repository execution.
- When the user is still deciding which reproduction target should count as the main run.
Clear boundaries
- This skill owns normalized reporting for an attempted command.
- It may receive execution evidence from the main skill or a thin helper.
- It does not choose the overall target on its own.
- It does not perform broad paper analysis.
- It does not own training startup, resume, or long-running training state.
- It should not normalize risky code edits into acceptable practice.
- It must not hide changes that alter evaluation, preprocessing, checkpoints, metrics, or other scientific meaning.
Input expectations
- selected reproduction goal
- runnable commands or smoke commands
- environment and asset assumptions
- optional patch metadata
Output expectations
- execution result summary
- standardized
repro_outputs/files SCIENTIFIC_CHANGELOG.mdfor changed scientific meaning and evidence statusCOMPARABILITY_REPORT.mdfor README/paper/baseline comparability- clear distinction between verified, partial, and blocked states
PATCHES.mdwhen repo files changed
Notes
Use references/reporting-policy.md, ../../references/research-rigor-principles.md, scripts/run_command.py, and scripts/write_outputs.py.
Signals
- GitHub stars
- 487
- Forks
- 17
- Last commit
- Sep 2026
- Installs
- 450k installs
Advanced
- Catalog kind
- skill
- Gateway key
minimal-run-and-audit- Source
- github.com/lllllllama/rigorpilot-skills