Collab: Autonomous AI Team Collaboration

SkillProductivity

Start a collaborative AI team (Codex + Claude) to work on a task together. Use when the user says "werk samen met Codex", "collab", "team onderzoek", "laat Codex en Claude samenwerken", or wants multiple AI agents to analyze, research, or solve something together autonomously.

Available today. Use it from your connected AI after setup.

Connect ahel once, and every AI you use reads what you have installed.

Then ask your AI: use the Collab: Autonomous AI Team Collaboration skill

What this skill tells your AI

The instructions your AI receives, as published by michelhelsdingen/ensemble in skill/SKILL.md and read by ahel’s review.

Language rule: ALWAYS respond in the same language the user used to invoke /collab. If the user writes in English, all your output (status updates, summaries, everything) must be in English. If Dutch, respond in Dutch. Never mix languages.

Launch a Codex + Claude team. Runtime files are namespaced under /tmp/ensemble/<TEAM_ID>/.

Script Paths

Every script lives in __ENSEMBLE_DIR__/scripts/ and MUST be called with its full path and .sh extension. They are not on $PATH, and the permission allowlist installed by setup-claude-code.sh only covers the full paths. Calling a bare collab-launch fails.

Set this once per session and reuse it:

ES="__ENSEMBLE_DIR__/scripts"
ScriptPurpose
collab-launch.shStart a team (runs preflight, opens the monitor, arms postcheck)
collab-poll.shSingle-shot message poll with seen-state tracking
collab-status.shDashboard of active and recent collabs
collab-rescue.shRe-inject prompts when messages.jsonl stays empty
collab-replay.shReplay a finished session in the terminal
collab-livefeed.shLive colored feed (non-tmux)
collab-cleanup.shRemove old finished runtime dirs (dry-run by default)
open-herdr-monitor.shOpen the monitor in a herdr pane, called by launch
collab-preflight.shAuth/DNS/service checks, auto-run by launch
collab-postcheck.shAgent health check ~30s after spawn, auto-armed by launch

Path Convention

All collab artifacts live in /tmp/ensemble/<TEAM_ID>/:

  • messages.jsonl — agent + ensemble message log
  • summary.txt — written on disband by ensemble-service
  • bridge.pid, bridge.log — bridge process
  • poller.pid, feed.txt — background poller
  • postcheck.log: output of the automatic post-spawn health check
  • prompts/, delivery/ — agent prompt/delivery files
  • .finished — written by ensemble-service AFTER summary.txt
  • team-id — team ID marker

Workflow

Step 0: Detect environment

if [ "${HERDR_ENV:-}" = "1" ]; then
  echo "HERDR"
elif [ -n "$TMUX" ]; then
  echo "TMUX_YES"
elif [ "$(uname)" = "Darwin" ] && [ "${TERM_PROGRAM:-}" = "iTerm.app" ]; then
  echo "ITERM_NATIVE"
else
  echo "TMUX_NO"
fi

Four monitor modes:

  • HERDR — inside a herdr workspace; collab-launch.sh asks herdr for the split
  • TMUX_YES — already inside tmux; collab-launch.sh opens a split pane right
  • ITERM_NATIVE — macOS iTerm2 without tmux; collab-launch.sh uses osascript to open a native iTerm split pane (no tmux attach needed)
  • TMUX_NO — fallback: detached tmux session the user must attach to

Check HERDR_ENV first, and never conclude "iTerm" from TERM_PROGRAM alone. herdr owns the terminal and draws its own panes inside a host session, but passes TERM_PROGRAM=iTerm.app straight through. An AppleScript split then opens a real iTerm pane outside the layout the user is watching: it succeeds, reports success, and is never seen. HERDR_ENV=1 is the reliable signal (HERDR_PANE_ID, HERDR_TAB_ID and HERDR_WORKSPACE_ID name the exact pane).

Force a specific mode with COLLAB_MONITOR=herdr|tmux|iterm|none, change the herdr layout with COLLAB_HERDR_MODE=split|tab, or the iTerm layout with COLLAB_ITERM_MODE=split|tab|window (both default split).

Step 1: Launch the team

"$ES/collab-launch.sh" "$(pwd)" "$TASK_DESCRIPTION" [AGENTS] [TEMPLATE]

Agent selection (3rd argument, optional). Comma-separated keys from agents.json; the first one becomes lead. Default when omitted: codex (lead) + claude code (worker). Available keys: codex, claude, grok, gemini, glm, opencode. These are the keys actually present in agents.json. An unknown key is not an error: resolveAgentProgram() falls back to claude, so a typo silently spawns a second claude instead of failing. Only pass this when the user explicitly names agents in the task ("laat gemini en claude…"). Preflight checks only the CLIs you name here, so a codex quota wall does not block a grok,claude run. Naming agents explicitly also disables the auto-fallback: a dead agent becomes a hard failure instead of a silent swap.

Also settable via COLLAB_AGENTS, for a preferred line-up you do not want to retype (export COLLAB_AGENTS="codex,claude,grok"). Precedence: 3rd argument > COLLAB_AGENTS > the default pair. The env var counts as naming your agents, so it disables the auto-fallback too.

Template selection (4th argument, optional). A key from collab-templates.json that gives each agent an explicit role instead of the generic lead/worker prompt. Also settable via COLLAB_TEMPLATE. Pick one when the task clearly matches:

TemplateRolesUse when the task is
reviewREVIEWER + CRITICreviewing existing code
implementARCHITECT + DEVELOPERbuilding a feature
researchRESEARCHER-A + RESEARCHER-Bcomparing options or exploring a topic
debugREPRODUCER + ANALYSTchasing a bug

Leave it empty if the task does not fit cleanly. A wrong template is worse than none.

Extract TEAM_ID from the launch output (last line is TEAM_ID=<id>):

TEAM_ID=$(printf '%s\n' "$LAUNCH_OUTPUT" | sed -n 's/^TEAM_ID=//p' | tail -1)

Do not read /tmp/collab-team-id.txt unless you have no launch output: it is a single global file that a concurrent collab overwrites.

Step 1b: Health checks (automatic, do not re-run)

  • Preflight runs inside collab-launch.sh before the team is created. It checks the service, agent CLI auth, and DNS. Non-zero exit means launch aborted with the fix command printed. Do not paper over it with COLLAB_SKIP_PREFLIGHT=1 unless the user asks. Exit codes: 1 service down, 2 service started in an unauthenticated shell (restart it), 3 claude CLI broken, 4 codex CLI broken, 5 DNS/network.
  • Postcheck is armed automatically and fires ~25s after spawn. If an agent is stuck in an error state it kills the team and writes the diagnosis to /tmp/ensemble/<TEAM_ID>/postcheck.log. If a team dies within the first minute, read that file before guessing.

Step 2: Tell the user where the monitor is

  • HERDR: "Team is live in the new herdr pane on the right."
  • TMUX_YES: "Team is live in the right tmux pane."
  • ITERM_NATIVE: "Team is live in the new iTerm pane on the right."
  • TMUX_NO: "tmux attach -t ensemble-$TEAM_ID — live TUI monitor (steer, disband, scroll)"

Step 3: Monitoring — the user MUST see the conversation

CRITICAL RULE: The user wants to SEE the team's conversation as it happens. Every poll result must be presented clearly and formatted as a readable conversation. Do NOT just dump raw output — format it as a proper dialogue.

If TMUX_NO: poll and PRESENT messages inline

Use collab-poll.sh — a single-shot poller that tracks state automatically and gives clean output.

Poll command:

"$ES/collab-poll.sh" "$TEAM_ID" --sleep <seconds>

Output format: sender\tcontent lines, ending with one of:

  • ---STATUS:ACTIVE — new messages were found
  • ---STATUS:QUIET — no new messages (agents in deep work)
  • ---STATUS:DONE — team finished, followed by summary.txt content
  • ---STATUS:WAITING — messages file not yet created

Presentation rules — THIS IS THE KEY PART: After each poll, present the new messages to the user like this:

codex-1: [message content]

claude-2: [message content]

Use markdown bold for agent names. Show the FULL message content (up to 500 chars), not truncated summaries. Between polls, add a brief status line like "Team is working... next check in 15s."

Polling cadence:

  • First poll: --sleep 10
  • Normal: --sleep 15 to --sleep 20
  • If 3+ polls QUIET: --sleep 30 (agents in deep work)
  • On ---STATUS:DONE: stop polling, present final summary

When done, present structured summary + clean up:

TEAM_ID="<id>" && RD="/tmp/ensemble/$TEAM_ID" && kill "$(cat "$RD/poller.pid" 2>/dev/null)" 2>/dev/null || true; kill "$(cat "$RD/bridge.pid" 2>/dev/null)" 2>/dev/null || true; tmux kill-session -t "ensemble-$TEAM_ID" 2>/dev/null || true
If HERDR: background summary watcher

Same as TMUX_YES: the monitor pane is visible to the user, so don't inline-poll. Wait for completion in the background and present the final summary. The monitor closes its own herdr pane on exit (pane id in /tmp/ensemble/<TEAM_ID>/herdr-pane-id), so no tmux kill-session.

If ITERM_NATIVE: background summary watcher

Same as TMUX_YES: the monitor pane is visible to the user, so don't inline-poll. Wait for completion in the background (same snippet as below) and present the final summary when done. On cleanup, the iTerm pane lives on — the user closes it with q or Cmd+W. Do NOT try to tmux kill-session (no tmux session exists in this mode).

If TMUX_YES: background summary watcher

Monitor visible in right pane. Wait in background:

TEAM_ID="<id>" && RD="/tmp/ensemble/$TEAM_ID" && while [ ! -f "$RD/.finished" ] && [ ! -f "$RD/summary.txt" ]; do sleep 8; done && echo "COLLAB_COMPLETE" && cat "$RD/summary.txt" 2>/dev/null

Run with run_in_background: true, timeout: 600000.

When done: summarize + cleanup poller/bridge PIDs.

How a team ends

Auto-disband triggers on one of two paths:

  1. Explicit sentinel (preferred). EVERY active agent sends a team-say whose entire content is <<COLLAB_DONE>>. This bypasses the message-count minimum and the idle wait. The agent prompts already instruct this. The bar is every agent, not two, so a third agent is never cut off while its teammates close up.
  2. Idle + completion signals. A safety net for teams that go quiet without sending the sentinel: it needs a minimum number of exchanged messages AND at least 60s of silence.

Ordinary words like "done" or "klaar" in a status update do not kill a team while the conversation is running. Do not warn agents to avoid those words.

Troubleshooting

Work through these in order before reporting a failure to the user.

SymptomAction
Team died within a minutecat /tmp/ensemble/$TEAM_ID/postcheck.log, it names the broken agent
Agents visible but messages.jsonl stays empty"$ES/collab-rescue.sh" "$TEAM_ID", re-injects the prompts the service failed to deliver
Launch aborted before creating a teamPreflight printed the fix command; run it, then relaunch
Not sure what is still running"$ES/collab-status.sh"
/tmp filling up with old runs"$ES/collab-cleanup.sh" (dry-run), then --force
Want to review a finished session"$ES/collab-replay.sh" "$TEAM_ID", or open the auto-generated HTML replay

Escalation. Two failed rescue or relaunch attempts on the same task is the limit. Stop, report to the user what was tried, what the logs said (postcheck.log, bridge.log, /tmp/ensemble-server.log), and ask how to proceed. Never keep relaunching a team that fails the same way. Each retry spawns real CLI sessions and burns real tokens.

Important Notes

  • Agents run with auto-accept permissions (configured in agents.json: codex --dangerously-bypass-approvals-and-sandbox, claude --permission-mode auto). They should NEVER ask for file write approval.
  • Do not modify project code during a collab session unless the user explicitly asks
  • Do not truncate or remove messages.jsonl
  • Multiple collabs can run simultaneously — each has own /tmp/ensemble/<TEAM_ID>/ namespace
  • team-say.sh uses fcntl.flock for atomic JSONL writes
  • ensemble-bridge.sh has single-instance guard, health check, exponential backoff
  • .finished and summary.txt are written by ensemble-service, NOT by scripts
  • Bridge auto-stops when it sees .finished marker

Signals

GitHub stars
245
Forks
29
Last commit
Sep 2026
Advanced
Catalog kind
skill
Gateway key
collab
Source
github.com/michelhelsdingen/ensemble