Use Iris
SkillCloud & infraLets your agent use claude skill workflows to submit, debug, and monitor Iris jobs and reserve dev GPUs and TPUs.
Use Use Iris in Claude, ChatGPT or Ahel Desktop
Free. Sign in, add Use Iris and connect your AI. About a minute.
Also: Claude Code · Cursor · Codex
Then ask your AI: use the Use Iris skill
Details
Instructions available. Your AI can read the instructions. Execution depends on the setup they require.
Account requirements not reviewed. Check the skill instructions before use; ahel provides instructions and does not run this skill.
No other account needed.
Add ahel to your AI once: Claude, ChatGPT, Cursor, Claude Code or Codex. Then ask it to use this.
About this skill
Use Iris to submit, inspect, debug, monitor, or recover jobs and tasks; diagnose scheduling and federation; deploy controllers; or reserve dev GPUs and TPUs. Use for ordinary Iris operations, including babysitting a job, controller rollouts, accelerator sessions, and stuck CoreWeave pods. Use a narr
What this skill tells your AI
The instructions your AI receives, as published by marin-community/marin in .agents/skills/use-iris/SKILL.md and read by ahel’s review.
Read only the material needed for the request:
- Normal jobs, tasks, scheduling, auth, or CoreWeave:
lib/iris/OPS.md. - Federation:
lib/iris/docs/federation.md. - Continuous job monitoring: references/monitor-job.md.
- Controller deploy or rollback: references/controller-rollout.md.
- Interactive GPU or TPU: references/dev-accelerators.md.
- Stuck terminating CoreWeave pod: references/stuck-pod.md.
- Temporary task outputs:
lib/iris/docs/task-outputs.md. - Logs or measurements: use
query-finelog.
Resolve cluster facts from lib/iris/config/<cluster>.yaml; do not copy live coordinates from memory.
Common reads
uv run iris --cluster=<cluster> job describe <job>
uv run iris --cluster=<cluster> task describe <task>
uv run iris --cluster=<cluster> task events <task>
uv run iris --cluster=<cluster> rpc controller list-backends
For a pending federated root, inspect all three parent-side views:
uv run iris --cluster=<parent> job list --prefix <root-job>
uv run iris --cluster=<parent> rpc controller list-peers
uv run iris --cluster=<parent> query \
"SELECT job_id, peer_id, handoff_state FROM federated_jobs WHERE job_id='<root-job>'"
Only root jobs federate; their whole tree stays on the peer. Parent job describe is the liveness source, while forwarded logs may lag. CoreWeave tasks normally read regional S3 and GCP tasks read GCS.
Temporary outputs
Write bounded diagnostics to $IRIS_OUTPUT_DIR. Iris preserves that directory as one outputs.tar.zst archive per attempt without changing the command outcome when capture fails. Find the archive URI and its uploaded, empty, failed, or unavailable state with:
uv run iris --cluster=<cluster> attempt describe <task>:<attempt>
Use direct object-storage writes for large or durable outputs. See lib/iris/docs/task-outputs.md for retention, limits, and data-access boundaries.
Boundaries
- Start read-only and name the evidence that distinguishes each cause.
- Never run
iris cluster restartwithout explicit approval for the named cluster; it kills all workers and jobs. - Treat a controller restart as a deployment and require an explicitly named target.
- Cancel, complete, fail, preempt, resubmit, or change Kubernetes state only when the request or selected reference authorizes that exact action.
- Avoid
kubectl describe podon task pods because it can print environment values.
Signals
- GitHub stars
- 4k
- Forks
- 311
- Last commit
- Oct 2026
Advanced
- Item type
- skill
- Key
use-iris- Source
- github.com/marin-community/marin
github.com/marin-community/marin
Related picks
Skill · docker
The pick for Dockerdocker-sandbox
Skill · joelhooks
The pick for Dockerazure-kubernetes
Skill · microsoft
The pick for Kubernetesinfra-containers-kubernetes
Skill · agents-inc
The pick for Kubernetesaws-architecture-diagram
Skill · awslabs
The pick for AWSaws-health-events
Skill · aws
The pick for AWS