Use Iris

SkillCloud & infra

Lets your agent use claude skill workflows to submit, debug, and monitor Iris jobs and reserve dev GPUs and TPUs.

Use Use Iris in Claude, ChatGPT or Ahel Desktop

Free. Sign in, add Use Iris and connect your AI. About a minute.

Also: Claude Code · Cursor · Codex

Then ask your AI: use the Use Iris skill

Details

Instructions available. Your AI can read the instructions. Execution depends on the setup they require.

Add ahel to your AI once: Claude, ChatGPT, Cursor, Claude Code or Codex. Then ask it to use this.

Use IrisStart free
About this skill

Use Iris to submit, inspect, debug, monitor, or recover jobs and tasks; diagnose scheduling and federation; deploy controllers; or reserve dev GPUs and TPUs. Use for ordinary Iris operations, including babysitting a job, controller rollouts, accelerator sessions, and stuck CoreWeave pods. Use a narr

What this skill tells your AI

The instructions your AI receives, as published by marin-community/marin in .agents/skills/use-iris/SKILL.md and read by ahel’s review.

Read only the material needed for the request:

Resolve cluster facts from lib/iris/config/<cluster>.yaml; do not copy live coordinates from memory.

Common reads

uv run iris --cluster=<cluster> job describe <job>
uv run iris --cluster=<cluster> task describe <task>
uv run iris --cluster=<cluster> task events <task>
uv run iris --cluster=<cluster> rpc controller list-backends

For a pending federated root, inspect all three parent-side views:

uv run iris --cluster=<parent> job list --prefix <root-job>
uv run iris --cluster=<parent> rpc controller list-peers
uv run iris --cluster=<parent> query \
  "SELECT job_id, peer_id, handoff_state FROM federated_jobs WHERE job_id='<root-job>'"

Only root jobs federate; their whole tree stays on the peer. Parent job describe is the liveness source, while forwarded logs may lag. CoreWeave tasks normally read regional S3 and GCP tasks read GCS.

Temporary outputs

Write bounded diagnostics to $IRIS_OUTPUT_DIR. Iris preserves that directory as one outputs.tar.zst archive per attempt without changing the command outcome when capture fails. Find the archive URI and its uploaded, empty, failed, or unavailable state with:

uv run iris --cluster=<cluster> attempt describe <task>:<attempt>

Use direct object-storage writes for large or durable outputs. See lib/iris/docs/task-outputs.md for retention, limits, and data-access boundaries.

Boundaries

  • Start read-only and name the evidence that distinguishes each cause.
  • Never run iris cluster restart without explicit approval for the named cluster; it kills all workers and jobs.
  • Treat a controller restart as a deployment and require an explicitly named target.
  • Cancel, complete, fail, preempt, resubmit, or change Kubernetes state only when the request or selected reference authorizes that exact action.
  • Avoid kubectl describe pod on task pods because it can print environment values.

Signals

GitHub stars
4k
Forks
311
Last commit
Oct 2026
Advanced
Item type
skill
Key
use-iris
Source
github.com/marin-community/marin