Parallel Debugging
SkillAI & modelsparallel-debugging is a skill that helps an AI agent debug complex issues by testing competing hypotheses in parallel. It guides the agent to generate multiple root-cause hypotheses across six failure categories, investigate them at the same time, and collect evidence with file and line citations. The agent then ranks hypotheses by evidence strength and confidence to declare a root cause and validates the fix against the original reproduction case.
Use Parallel Debugging in Claude, ChatGPT or Ahel Desktop
Free. Sign in, add Parallel Debugging and connect your AI. About a minute.
Also: Claude Code · Cursor · Codex
Then ask your AI: use the Parallel Debugging skill
Details
Instructions available. Your AI can read the instructions. Execution depends on the setup they require.
Account requirements not reviewed. Check the skill instructions before use; ahel provides instructions and does not run this skill.
No other account needed.
Have a bug with multiple plausible root causes or an issue that spans several modules.
What your AI can do with it
- Generate hypotheses across six failure mode categories
- Investigate competing hypotheses in parallel
- Collect evidence with file and line citations
- Rank hypotheses by evidence strength and confidence
- Declare a root cause and validate the fix
Getting started
- Have a bug with multiple plausible root causes or an issue that spans several modules.
- Add the parallel-debugging skill to your agent setup.
- Give the agent the bug report and any reproduction steps you have.
- Let the agent generate and investigate hypotheses, then review the ranked evidence and declared root cause.
What this skill tells your AI
The instructions your AI receives, as published by wshobson/agents in plugins/agent-teams/skills/parallel-debugging/SKILL.md and read by ahel’s review.
Framework for debugging complex issues using the Analysis of Competing Hypotheses (ACH) methodology with parallel agent investigation.
When to Use This Skill
- Bug has multiple plausible root causes
- Initial debugging attempts haven't identified the issue
- Issue spans multiple modules or components
- Need systematic root cause analysis with evidence
- Want to avoid confirmation bias in debugging
Hypothesis Generation Framework
Generate hypotheses across 6 failure mode categories:
1. Logic Error
- Incorrect conditional logic (wrong operator, missing case)
- Off-by-one errors in loops or array access
- Missing edge case handling
- Incorrect algorithm implementation
2. Data Issue
- Invalid or unexpected input data
- Type mismatch or coercion error
- Null/undefined/None where value expected
- Encoding or serialization problem
- Data truncation or overflow
3. State Problem
- Race condition between concurrent operations
- Stale cache returning outdated data
- Incorrect initialization or default values
- Unintended mutation of shared state
- State machine transition error
4. Integration Failure
- API contract violation (request/response mismatch)
- Version incompatibility between components
- Configuration mismatch between environments
- Missing or incorrect environment variables
- Network timeout or connection failure
5. Resource Issue
- Memory leak causing gradual degradation
- Connection pool exhaustion
- File descriptor or handle leak
- Disk space or quota exceeded
- CPU saturation from inefficient processing
6. Environment
- Missing runtime dependency
- Wrong library or framework version
- Platform-specific behavior difference
- Permission or access control issue
- Timezone or locale-related behavior
Evidence Collection Standards
What Constitutes Evidence
| Evidence Type | Strength | Example |
|---|---|---|
| Direct | Strong | Code at file.ts:42 shows if (x > 0) should be if (x >= 0) |
| Correlational | Medium | Error rate increased after commit abc123 |
| Testimonial | Weak | "It works on my machine" |
| Absence | Variable | No null check found in the code path |
Citation Format
Always cite evidence with file:line references:
**Evidence**: The validation function at `src/validators/user.ts:87`
does not check for empty strings, only null/undefined. This allows
empty email addresses to pass validation.
Confidence Levels
| Level | Criteria |
|---|---|
| High (>80%) | Multiple direct evidence pieces, clear causal chain, no contradicting evidence |
| Medium (50-80%) | Some direct evidence, plausible causal chain, minor ambiguities |
| Low (<50%) | Mostly correlational evidence, incomplete causal chain, some contradicting evidence |
Result Arbitration Protocol
After all investigators report:
Step 1: Categorize Results
- Confirmed: High confidence, strong evidence, clear causal chain
- Plausible: Medium confidence, some evidence, reasonable causal chain
- Falsified: Evidence contradicts the hypothesis
- Inconclusive: Insufficient evidence to confirm or falsify
Step 2: Compare Confirmed Hypotheses
If multiple hypotheses are confirmed, rank by:
- Confidence level
- Number of supporting evidence pieces
- Strength of causal chain
- Absence of contradicting evidence
Step 3: Determine Root Cause
- If one hypothesis clearly dominates: declare as root cause
- If multiple hypotheses are equally likely: may be compound issue (multiple contributing causes)
- If no hypotheses confirmed: generate new hypotheses based on evidence gathered
Step 4: Validate Fix
Before declaring the bug fixed:
- Fix addresses the identified root cause
- Fix doesn't introduce new issues
- Original reproduction case no longer fails
- Related edge cases are covered
- Relevant tests are added or updated
Signals
- GitHub stars
- 40k
- Forks
- 4k
- Last commit
- Sep 2026
Others that do the same job
Questions
- When should I use this skill?
- Use it when a bug has multiple plausible root causes, initial debugging attempts have not identified the issue, the issue spans multiple modules or components, or you need systematic root cause analysis with evidence.
- What failure mode categories does it cover?
- It generates hypotheses across six categories: logic error, data issue, state problem, integration failure, resource issue, and environment.
- How does it decide which hypothesis is the root cause?
- It ranks hypotheses by evidence strength and confidence, then declares a root cause and validates the fix against the original reproduction case.
Advanced
- Item type
- skill
- Key
parallel-debugging-wshobson- Source
- github.com/wshobson/agents
Related picks
Skill · mattpocock
Does the same job in other wordsce-debug
Skill · everyinc
Does the same job in other wordsdebugging
Skill · code-yeongyu
Does the same job in other wordssetup-ts-deep-modules
Skill · mattpocock
The pick for TypeScripttypescript-pro
Skill · jeffallan
The pick for TypeScriptskill-creator
Skill · anthropics
More in AI & models