Debug

SkillDev tools

Guides your agent through debugging a stated code, JAX, TPU, or pipeline fault using a structured debug skill.

Use Debug in Claude, ChatGPT or Ahel Desktop

Free. Sign in, add Debug and connect your AI. About a minute.

Also: Claude Code · Cursor · Codex

Then ask your AI: use the Debug skill

Details

Instructions available. Your AI can read the instructions. Execution depends on the setup they require.

Add Ahel to your AI once: Claude, ChatGPT, Cursor, Claude Code or Codex. Then ask it to use this.

About this skill

Diagnose a stated code, JAX, Marin, Iris, Zephyr, or TPU fault or startup/performance regression; do not activate for ordinary implementation or optimization without a symptom.

What this skill tells your AI

The instructions your AI receives, as published by marin-community/marin in .agents/skills/debug/SKILL.md and read by Ahel’s review.

Keep working notes in the active task. Do not add repository debug-log files. At the start of a fresh investigation, use consult-echo when repository policy requires prior-work search. During continued diagnosis or mitigation of the same incident, reuse results already read instead of invoking consult-echo again unless the scope or freshness requirement changes or the user asks for a new search. After diagnosing a live infrastructure incident, use write-ops-log to publish its standalone Echo record and link it from the associated PR or issue. A code bug or local debugging session is not an incident unless it caused a service, production run, or shared operational system to fail or degrade.

Infrastructure faults

Read lib/iris/AGENTS.md or lib/zephyr/AGENTS.md for context, then follow the matching OPS.md section:

SymptomRead
Stuck job, scheduling failure, resource leak, controller stalledlib/iris/OPS.md → SQL Queries, Process Inspection & Profiling, Known Bugs, Troubleshooting
Iris task misbehaving, container inspection, profiling a running tasklib/iris/OPS.md → Task Operations, Process Inspection & Profiling
Zephyr pipeline slow / stragglers / data skew / worker failureslib/zephyr/OPS.md → Diagnostic Patterns, Observability
TPU bad node (No accelerator found, FAILED_PRECONDITION, Device or resource busy)lib/iris/OPS.md → TPU Bad-Node Recovery

Read the guardrails beside the commands. Never modify the controller database, prefer iris process profile over SSH, and never run a full iris cluster restart without approval. After a TPU recovery or Zephyr fix, return to the active Iris job-monitoring or babysit-zephyr loop.

Code bugs

For code bugs, reproduce the failure, identify the smallest falsifiable hypothesis, change one cause at a time, and test the behavior that failed. Let exceptions propagate unless added context changes the diagnosis.

Signals

GitHub stars
4k
Forks
311
Last commit
Oct 2026
Hacker News mentions
20
Advanced
Item type
skill
Key
debug-marin-community
Source
github.com/marin-community/marin