Debug
SkillDev toolsGuides your agent through debugging a stated code, JAX, TPU, or pipeline fault using a structured debug skill.
Use Debug in Claude, ChatGPT or Ahel Desktop
Free. Sign in, add Debug and connect your AI. About a minute.
Also: Claude Code · Cursor · Codex
Then ask your AI: use the Debug skill
Details
Instructions available. Your AI can read the instructions. Execution depends on the setup they require.
Account requirements not reviewed. Check the skill instructions before use; Ahel provides instructions and does not run this skill.
No other account needed.
Add Ahel to your AI once: Claude, ChatGPT, Cursor, Claude Code or Codex. Then ask it to use this.
About this skill
Diagnose a stated code, JAX, Marin, Iris, Zephyr, or TPU fault or startup/performance regression; do not activate for ordinary implementation or optimization without a symptom.
What this skill tells your AI
The instructions your AI receives, as published by marin-community/marin in .agents/skills/debug/SKILL.md and read by Ahel’s review.
Keep working notes in the active task. Do not add repository debug-log files.
At the start of a fresh investigation, use consult-echo when repository policy
requires prior-work search. During continued diagnosis or mitigation of the
same incident, reuse results already read instead of invoking consult-echo
again unless the scope or freshness requirement changes or the user asks for a
new search. After diagnosing a live infrastructure incident, use
write-ops-log to publish its standalone Echo record and link it from the
associated PR or issue. A code bug or local debugging session is not an
incident unless it caused a service, production run, or shared operational
system to fail or degrade.
Infrastructure faults
Read lib/iris/AGENTS.md or lib/zephyr/AGENTS.md for context, then follow
the matching OPS.md section:
| Symptom | Read |
|---|---|
| Stuck job, scheduling failure, resource leak, controller stalled | lib/iris/OPS.md → SQL Queries, Process Inspection & Profiling, Known Bugs, Troubleshooting |
| Iris task misbehaving, container inspection, profiling a running task | lib/iris/OPS.md → Task Operations, Process Inspection & Profiling |
| Zephyr pipeline slow / stragglers / data skew / worker failures | lib/zephyr/OPS.md → Diagnostic Patterns, Observability |
TPU bad node (No accelerator found, FAILED_PRECONDITION, Device or resource busy) | lib/iris/OPS.md → TPU Bad-Node Recovery |
Read the guardrails beside the commands. Never modify the controller database,
prefer iris process profile over SSH, and never run a full
iris cluster restart without approval. After a TPU recovery or Zephyr fix,
return to the active Iris job-monitoring or babysit-zephyr loop.
Code bugs
For code bugs, reproduce the failure, identify the smallest falsifiable hypothesis, change one cause at a time, and test the behavior that failed. Let exceptions propagate unless added context changes the diagnosis.
Signals
- GitHub stars
- 4k
- Forks
- 311
- Last commit
- Oct 2026
- Hacker News mentions
- 20
Advanced
- Item type
- skill
- Key
debug-marin-community- Source
- github.com/marin-community/marin
github.com/marin-community/marin