start-experiment
SkillDocs & knowledgeStart the autoresearch optimization loop for a specific model + lane. Resolves the hierarchical program.md (root → model → lane), asks the user for hardware (local TPU VM or GKE cluster of a specified TPU type + topology), discovers available clusters from .env/, checks occupancy with USER_PREFIX-aware attribution, arms the launch-time process watcher (Step 9·0), and starts the prose loop, the session itself drives the iteration protocol. Invoke at the beginning of an autoresearch session.
Use start-experiment in Claude, ChatGPT or Ahel Desktop
Free. Sign in, add start-experiment and connect your AI. About a minute.
Also: Claude Code · Cursor · Codex
Then ask your AI: use the start-experiment skill
Details
Instructions available. Your AI can read the instructions. Execution depends on the setup they require.
Account requirements not reviewed. Check the skill instructions before use; Ahel provides instructions and does not run this skill.
No other account needed.
Add Ahel to your AI once: Claude, ChatGPT, Cursor, Claude Code or Codex. Then ask it to use this.
What this skill tells your AI
The instructions your AI receives, as published by vlasenkoalexey/tpu_performance_autoresearch_wiki in .claude/skills/start-experiment/SKILL.md and read by Ahel’s review.
You are starting one autoresearch session. Follow this sequence precisely. Do not skip steps.
Step 1 — Determine context (model, lane, parallelism)
KERNEL FAST-PATH — check this FIRST. If the invocation names a kernel (a kernel/benchmark path, a wiki/kernels/ family slug, or "tpu chip N"), this is a KERNEL session. Steps 2–8 are model-lane machinery — skip them ALL (no XPK cluster discovery, no xprof probe, no model-lane hardware question). The kernel lane asks its own, much smaller run-target question at Step 9·K below. Do exactly:
- Read
wiki/kernel_experiments/program.mdend-to-end; derive the family slug per its Quick-start rule 1. - Step 9·K — choose the run target (ASK ONCE, then it is fixed for the whole run). See below.
- Step 9·0 — arm the watcher (kernel mode:
family+home_repofrom the family binding; bootstrap the binding first if new, per Quick-start rule 2). - Step 9b — start marker, into the FAMILY's log:
wiki/kernel_experiments/<slug>/pallas/log.md. Record the run target chosen at 9·K. - Hand off: run the kernel Quick start (load
/author-kernel, then K0–K9 viakexec.sh run).
Step 9·K — Run target: local chip or cluster pod (KERNEL ONLY)
The chip named in the prompt is the local default. Before arming anything, ask the user with AskUserQuestion (--yes skips the question and takes the local default — same convention as Step 8):
Question: "Run this kernel family on the local chip, or on a cluster pod?" Options:
Local chip <N>(recommended) — fastest edit→measure cycle; right for authoring.Cluster pod— persistent GKE pod,kubectl exec. For a different target generation or more chips than local. Costs contended capacity for the whole run.
(When to pick which — the full rationale — is canonical in wiki/kernel_experiments/program.md's run-target section; don't restate it here.)
If Cluster pod:
tools/kernel_exec/kexec.sh discover— prints the TPU capacity actually present (nodepool, accelerator, topology, machine type, chips). Never assume a generation; the GKE accelerator label value is generation-specific (tpu7x,tpu-v6e-slice,tpu-v5p-slice,tpu-v5-lite-podslice, …).- Ask which row to use, and for how long (
--hours, default 8 — the pod's hard TTL). - Bring it up — it stays up for the whole run:
tools/kernel_exec/kexec.sh up --family <slug> \ --accelerator <A> --topology <T> [--chips N] --hours <H> --image <IMG>--imagemust havejax[tpu]+ libtpu + kgate.--chipsdefaults to the node's allocatable count. kexec.sh sync --family <slug>after each K4 authoring pass, so the pod sees the current.repo.
Two properties to state to the user when the pod comes up, because both are load-bearing:
- Teardown is mandatory —
/stop-experimentdoes it, but the pod also carriesactiveDeadlineSecondsso the cluster reaps it even if every software path fails.sleep infinitynever exits on its own; a leaked pod holds scarce chips indefinitely. - Artifacts must be
gs://—kexecrefuses to run otherwise. All four classes (HLO / LLO / Mosaic / xprof trace) write straight to GCS, which keeps the pod stateless and immune to theephemeral-storageeviction that has destroyed captures before.
Record the choice in the Step 9b start marker (**Run target**: local chip N or **Run target**: cluster pod <pod> — <accel>/<topo>, <chips> chips, TTL <H>h). Do not re-ask per iteration.
Everything below this line is the MODEL-lane path.
Model + lane, try in order:
- Infer from CWD: if the current working directory contains
wiki/experiments/<model>_autoresearch_optimization/<lane>/, use<model>and<lane>. - Read user invocation: the user may have typed
/start-experiment <model> <lane>(e.g./start-experiment <model> tpu). Parse if present. - Ask: if neither resolves both, use
AskUserQuestionto ask. Offer the model folders underwiki/experiments/(currently:<model>_autoresearch_optimization,qwen3_autoresearch_optimization,gemma4_autoresearch_optimization,llama3_8B_autoresearch_optimization); then ask for lane (typicallytpu,jax,torchax— list whatever subdirectories of the chosen model folder actually contain aprogram.md).
Parallelism (how many clusters to run in parallel as independent tracks):
- Parse
--parallelism N(or--parallelism all) from the invocation if present. - If not given, default to 1 (single-cluster, current behavior — safest for new sessions).
- If user asked for
all, treat as "up to the number of free clusters matching the requested TPU type + topology" (resolved in step 6). - If user asked for
N > free clusters available, use however many ARE free, surface to user: "requested N, only K free, proceeding with K".
Hold the parallelism value through step 6 (cluster selection) and step 9 (loop start).
Step 2 — Resolve hierarchical program.md
Read in order, gracefully skipping any that don't exist:
wiki/experiments/program.md
wiki/experiments/<model>_autoresearch_optimization/program.md
wiki/experiments/<model>_autoresearch_optimization/<lane>/program.md
(Kernel families never reach this step — the Step 1 fast-path handled them. Their spec is self-contained: wiki/kernel_experiments/program.md → <family>/pallas/program.md. Both lanes share one supervision pattern: the Step 9·0 watcher, armed once at launch, plus the prose loop the runner drives at Step 9c. /loop and the Stop-hook/marker machinery are retired — see SCHEMA "Launch-armed process watcher"; rationale in the 2026-07-21 enforcement design record.)
Apply replace-per-section resolution (per the inheritance rule in root program.md): later files completely replace earlier files' definitions of the same H2 section; new sections in later files are taken as-is.
Print a one-line audit summary to the user showing which level each section came from. Example:
Effective program for <model> / tpu:
Inheritance model ← root
Concurrency model ← root
Setup ← lane (overrides model + root)
Workload naming ← root (no model/lane override)
Branching model ← model
Operational env vars ← lane
Kernels available ← lane
The goal ← model
The experiment loop ← root
... etc.
Step 2b — Derive USER_PREFIX and MODEL_NAME
USER_PREFIX resolution chain (first hit wins):
- If model-level
program.mdhas an explicitUSER_PREFIX = <value>line, use it. $USERbefore first underscore, lowercased:echo "${USER%%_*}" | tr '[:upper:]' '[:lower:]'. Example:alekseyv_google_com→alekseyv.- Git first-name fallback:
git config user.name | awk '{print tolower($1)}'. Example:Aleksey Vlasenko→aleksey. - If still empty, refuse and ask the user to set it explicitly.
MODEL_NAME is auto-derived from the model folder name:
echo "<model>_autoresearch_optimization" | sed 's/_autoresearch_optimization$//' | tr '_' '-' | tr '[:upper:]' '[:lower:]'
Examples: <model> → <model>, gemma4 → gemma4, llama3_8B → llama3-8b.
Print both to the user before continuing:
USER_PREFIX = alekseyv (from $USER segment)
MODEL_NAME = <model> (auto-derived from folder)
LANE = tpu (from context)
Step 3 — Determine hardware target
Ask the user via AskUserQuestion:
Question: "Where should experiments run?" Options:
Local (this TPU VM)— runs in the master session, no XPK. Use when the local machine itself is a TPU VM with enough chips for the requested experiment.GKE (XPK)— submits workloads to a GKE cluster. Will require TPU type + topology selection in the next step.
If the user picks Local:
- Confirm the local TPU is appropriate for the model size:
python -c "import jax; print(jax.devices())"to see what's available locally. - Skip the cluster discovery (steps 4–6) and proceed to step 7.
If the user picks GKE:
- Continue to step 4.
Step 4 — Ask for TPU type + topology (GKE only)
Ask the user via AskUserQuestion:
Question 1: "Which TPU generation?"
Options: v5p, v6e (list from what .env/ actually contains — scan filenames for v5p, v6e, etc.)
Question 2: "Which topology / chip count?"
Options: scan .env/*.md for the chosen generation, extract topology from the "Topology" section of each cluster file (look for lines like **TPU**: v5p, 2x2x2 topology = 8 chips per slice or **TPU**: v6e, 2x4 = 8 chips). Group cluster files by topology and present each topology as an option. Example for v5p: 2x2x1 (4 chips), 2x2x2 (8 chips), 4x2x2 (16 chips), etc.
Step 5 — Discover candidate clusters (GKE only)
From .env/, list every cluster file matching the chosen (TPU generation, topology). Parse each cluster file to extract:
cluster_name(from filename patterngke-<tpu>-<topo>-<owner>.mdor from the "Cluster" line in the Topology section)region,zone,project(from the Connection or Topology section)context_name(kubectl context — typically printed as a comment aftergcloud get-credentials)tpu_type(the xpk-style--tpu-typevalue — note the v5p TC-vs-chip distinction: xpk'sv5p-16= 16 TC = 8 chips;v5p-8= 8 TC = 4 chips)chip_count
Build a candidate list. Example:
v5p, 2x2x2 (8 chips per slice) candidates from .env/:
- atwigg-v5p-16 (europe-west4-b, cloud-tpu-multipod-dev)
- tsbao-v5p-16 (europe-west4-b, cloud-tpu-multipod-dev)
- wenxindong-pw-v5p-16-2 (europe-west4-b, cloud-tpu-multipod-dev)
- niting-v5p-16 (europe-west4-b, cloud-tpu-multipod-dev)
Step 6 — Occupancy check + cluster selection (GKE only)
For each candidate cluster in turn, run the occupancy check:
# Fetch credentials (skip if already in kubeconfig)
gcloud container clusters get-credentials <cluster_name> --location=<region> --project=<project> 2>/dev/null
# List active workloads
xpk workload list --project=<project> --zone=<zone> --cluster=<cluster_name> 2>&1
# Or:
kubectl --context=<context_name> get jobset -A --no-headers 2>&1
Classify each active workload from the name alone (the format <USER_PREFIX>-<MODEL_NAME>-<LANE>-v<NNN>-<SLUG>[-<retry>] makes this one-step):
| Workload name pattern | Classification |
|---|---|
Starts with <USER_PREFIX>-<MODEL_NAME>-<LANE>- (requested lane) | mine, same model+lane → cluster busy, skip |
Starts with <USER_PREFIX>-<MODEL_NAME>-<other-lane>- | mine, same model other lane → conflict, skip |
Starts with <USER_PREFIX>-<other-model>- | mine, other model → conflict, skip |
Starts with <other-prefix>- | foreign → cluster occupied by another user, skip |
Image-tag inspection is the backstop only — use it if a workload name doesn't follow the convention (legacy workloads or manually-submitted ones), to verify lane/model. Pattern:
kubectl --context=<ctx> get jobset <workload> -o jsonpath='{.spec.replicatedJobs[0].template.spec.template.spec.containers[0].image}'
# Returns: <base>:<branch> where branch encodes <model>-<lane>-<date>-v<NNN>-<slug>
Selection rule (parallelism-aware):
- Walk candidates, classify each as
freeoroccupiedper the attribution table above. - Select up to N free clusters (where N = parallelism from step 1;
all= every free candidate). - If parallelism = 1: pick the first free cluster (current single-cluster behavior).
- If parallelism > 1: pick the first N free clusters. Pool will operate as N independent tracks.
- If fewer than N are free, use however many are free and surface: "requested N, only K free, proceeding with K".
- If zero are free, report each candidate's occupancy with attribution and STOP. Do NOT pick an occupied cluster. Do NOT start the run without targets.
Example output for parallelism=3:
v5p (2x2x2) cluster pool (4 candidates, 3 selected):
✓ atwigg-v5p-16 → selected (track 0)
✓ tsbao-v5p-16 → selected (track 1)
✗ wenxindong-pw-v5p-16-2 → occupied by my own jax-lane workload `alekseyv-<model>-jax-v204-...`
✓ niting-v5p-16 → selected (track 2)
Proceeding with 3 parallel tracks.
Step 7 — Re-ground from the wiki
Before launching the loop, read the current state of the project:
- Last 50 lines of the lane's log:
wiki/experiments/<model>_autoresearch_optimization/<lane>/log.md(this is the per-lane log per SCHEMA's two-tier convention; if it doesn't exist yet, this is the first session on this lane — create it empty at Step 9's loop-start marker) - Last 30 lines of the global
wiki/log.md(cross-cutting events — schema changes, ingests, lane scaffolding — that may affect this lane) - The active model page:
wiki/models/<model>-<lane>.md(variant matrix, Current best, Open hyps, Frontier exp) - The last 2–3 experiment pages in
wiki/experiments/<model>_autoresearch_optimization/<lane>/ - Any open hypotheses in
wiki/hypotheses/tagged for this model + lane
Summarize to the user in 5–10 lines: which variant is the frontier, what was just learned, what's open, what hypothesis you'd run first.
Step 7.5 — Probe xprof-cli (serverless — no MCP server, no :8791)
The profile-analyzer agent (dispatched per experiment) runs all xprof/HLO/LLO reads through the xprof-cli CLI in serverless local mode (XPROF_MODE=local, in-process; --logdir takes GCS paths and local dirs identically). There is no server to start. If the CLI is missing, every dispatch fails Phase 1 → no ## Profile / ## HLO Dump sections → LINT failures.
Probe once:
XPROF_MODE=local xprof-cli list_runs --logdir=<shared-profiles-tree>
- Returns a (possibly empty) run list → proceed.
- Command not found / import error → surface to the user: install with
pip install -e raw/code/xprof-cli(the xprof-cli checkout; see.claude/agents/profile-analyzer.mdfor the tool inventory). Ifraw/code/xprof-cliis empty: the submodule is a private repo markedupdate = none— the user needs access (reach out to @vlasenkoalexey), thengit submodule update --init --checkout raw/code/xprof-cli; see README "Configuring xprof-cli". Re-probe before continuing.
Do NOT proceed past this step if xprof-cli is not functional. (The legacy mcp__xprof__* server transport is retired from this flow — never ask the user to start xprof --port=8791.)
Step 8 — Confirm with user
Use AskUserQuestion:
Question: "Start the experiment run with this plan?" Options:
Yes, start— proceed to step 9. (Description shown to the user: "Runs autonomously in this session; a background audit subagent supervises and revives it. Stop anytime with /stop-experiment.")Different first hypothesis— let user redirect.Cancel— exit without starting.
(Do not mention internal step numbers like "Step 9·0" in the question or option text — the user hasn't read this skill; describe mechanics in plain terms only.)
CLI flag override: --yes — skip the question (assume "Yes, start"; safe default for unattended scripted invocation).
(There is no never-stop question anymore. The former opt-in Stop-hook/marker mechanism is retired — premature-stop protection is now unconditional, provided by the Step 9·0 watcher's braking + reviving services and by reader-side validation: a stop without the required retrospective/evidence is flagged by the watcher and voided+reopened by LINT.)
Step 9 — Start the run (arm watcher, mark log, drive the prose loop)
Step 9·0 — Arm the process watcher (ONE-TIME, both lanes)
ARM IT NOW, EXACTLY ONCE. The process-auditor runs as a persistent, self-rescheduling watcher — your harness's scheduler drives the repetition, never your per-iteration memory. Never re-arm per iteration. Inputs: kernel mode family + home_repo; model mode model + lane (the auditor picks its battery from the mode).
How to arm — native scheduler only, no scripts or scaffolding:
| harness | arming action (do this now) |
|---|---|
| claude | launch a background Agent (Task) that re-runs every ~5 min or on new commits under the lane/family dir; findings arrive as task notifications |
| agy | define_subagent the auditor + arm a self-rescheduling Schedule/ManageTask loop (timer or new-commit wake) that invokes it as a background subagent and returns findings to your context; the reschedule is part of this one arming |
| codex | no scheduled watcher: dispatch the process-auditor subagent (.codex/agents/process-auditor.toml) after filing each experiment; its result auto-returns to your context |
What the watcher does (it replaces the retired /loop + Stop hook):
- Check — delta audit since
.audit-cursor; findings + paste-ready corrections land in your context. Your rule: apply them before your next K3 (kernel) / next iteration (model). - Brake — a stop/at-ceiling claim without its artifacts ⇒ "stop blocked"; LINT voids+reopens unearned closes.
- Revive — its scheduled firing wakes an idle session (proven on agy, 2026-07-21), restarting a runner that stopped early.
Kernel families — clear any stale stop authorization as part of this arming: rm -f wiki/kernel_experiments/<family>/pallas/.stop-authorized. That file is the auditor-written close authorization (/stop-experiment Step 1·0); one left over from a PREVIOUS close must never satisfy a new run's test -e gate — a fresh run voids all prior authorization.
(Why launch-time arming: the per-iteration dispatch buried in program.md fired for 1 of 6 agy families — see the 2026-07-21 enforcement design record, "Reversal".)
Step 9b — Write start marker to the lane's log
Write a start marker to the lane's log so the lane's log starts with the session boundary. Path: wiki/experiments/<model>_autoresearch_optimization/<lane>/log.md (kernel fast-path: the FAMILY's log, wiki/kernel_experiments/<slug>/pallas/log.md, with Cluster pool = chip <N> and Parallelism = 1). Create the file if it doesn't exist; insert at the top (newest-first):
## [YYYY-MM-DD] start | /start-experiment session begin
**Op**: start
**Cluster pool**: <comma-separated cluster names>
**Parallelism**: <N>
**First-pick hypothesis**: <one-line from Step 7's summary>
**Notes**: session opened via /start-experiment.
Step 9c — Run the prose loop (the runner is the driver)
There is no /loop skill invocation and no external loop machinery: this session IS the loop driver. Adopt the iteration protocol below as your standing operating instructions and execute it repeatedly — kernel lanes run K0→K9 synchronously per wiki/kernel_experiments/program.md; model lanes run the iteration protocol below. The Step 9·0 watcher supervises (checks, brakes, revives); you drive. Substitute <model>, <MODEL_NAME>, <lane>, <USER_PREFIX>, and <CLUSTER_POOL> (a list of {name, context} for the N selected clusters from step 6).
You are running the <model> / <lane> autoresearch loop, one iteration at a time.
Session constants (derived once by /start-experiment, do not re-derive):
USER_PREFIX = <USER_PREFIX>
MODEL_NAME = <MODEL_NAME>
LANE = <lane>
CLUSTER_POOL = [
{name: "<cluster_1>", context: "<context_1>"},
{name: "<cluster_2>", context: "<context_2>"},
...
] # N independent tracks, one per cluster
ARCHITECTURE: parallel-tracks-via-background-subagents.
- Each cluster is an INDEPENDENT TRACK with its own experiment lifecycle.
- Cluster-runner subagents are dispatched with run_in_background=true.
- Master does NOT block on subagents — it walks the pool, dispatches idle clusters,
processes completed background notifications, and exits the iteration.
- When a background subagent completes, the master is auto-notified — process on
next iteration's step 2(a).
Iteration steps:
0. BACKFILL missing wiki pages (catches subagent silent-fail / iteration-race /
direct-kubectl-bypass failure modes):
For each cluster in CLUSTER_POOL, list Completed workloads matching
`<USER_PREFIX>-<MODEL_NAME>-<LANE>-v<NNN>-*` via `kubectl get jobset` (or
`xpk workload list`). For each Completed workload:
- Extract `v<NNN>` from the workload name.
- Check if `wiki/experiments/<model>_autoresearch_optimization/<lane>/`
contains a `*-v<NNN>-*.md` page.
- If NO page exists: this is a dispatch that completed without filing.
File a page from `kubectl logs <pod> --tail=200` (extract MFU, loss,
exit code, headline metrics).
VERDICT POLICY for backfilled pages (NEVER assigns supported/refuted —
those require profile-analyzer's hypothesis-firing audit, which did
not run):
- If logs indicate crash / non-zero exit → `verdict: invalid`,
reason: "crashed; logs: <one-line summary>"
- If logs show clean completion but no analyzer ran →
`verdict: inconclusive`, reason: "backfilled — profile-analyzer
not dispatched"
FRONTMATTER add: `backfilled: true` — this is the LINT exception
marker. SCHEMA's LINT check for missing `## Profile` / `## HLO Dump`
skips pages with `backfilled: true`. The frontmatter persists; the
page documents the gap rather than failing LINT.
Page body: `## Hypothesis under test` is unknown (no stub was filed),
so leave it as: "**Hypothesis not recovered** — page filed by
BACKFILL after the run completed without a stub. The original
dispatch context was lost; treat this experiment as
observation-only."
Surface to user: "Backfilled N missing pages (all marked invalid or
inconclusive — no supported/refuted verdicts assigned without analyzer)."
If N=0, no mention.
This step is cheap (1 kubectl call + 1 dir listing) and prevents the
wiki from drifting out of sync with cluster reality.
1. RE-GROUND from disk (ORDER MATTERS):
(a) PROGRAM (methodology — the drift-prevention anchor; do NOT skip):
Read wiki/experiments/program.md (root).
Read wiki/experiments/<model>_autoresearch_optimization/program.md (model-level).
Read wiki/experiments/<model>_autoresearch_optimization/<lane>/program.md
(lane-level, if exists; gracefully skip if not).
Apply replace-per-section resolution. Use additive-section convention for
sections like "<Model>-specific CAN additions".
(b) STATE (what's happened):
Read last 50 lines of the LANE'S log:
wiki/experiments/<model>_autoresearch_optimization/<lane>/log.md
(per SCHEMA's two-tier log convention — loop-iteration entries
live here, not in the global wiki/log.md). If the file doesn't
exist, this is the lane's first iteration — proceed; the loop
creates it at first append.
Read last 30 lines of global wiki/log.md (cross-cutting events
that may affect this lane — schema changes, ingests, etc.).
Read the active model page variant matrix (wiki/models/<model>-<lane>.md):
current best, open hyps, frontier exp.
Read the last 2-3 experiment pages in your lane.
(c) LIVE (what's running):
For each cluster in CLUSTER_POOL, xpk workload list to enumerate in-flight
workloads matching <USER_PREFIX>-<MODEL_NAME>-<LANE>-* (yours).
Shortened here. Read the whole file on GitHub.
Signals
- GitHub stars
- 56
- Forks
- 5
- Last commit
- Sep 2026
Ahel review
K1binfo
installs-packages
Automated review, not a security audit. Ruleset v1+k2.
Advanced
- Item type
- skill
- Key
start-experiment- Source
- github.com/vlasenkoalexey/tpu_performance_autoresearch_wiki
github.com/vlasenkoalexey/tpu_performance_autoresearch_wiki
Related picks
Skill · jeremylongshore
The pick for GCPgcp-security-scanner
Skill · a5c-ai
The pick for GCPdocker-agent-run
Skill · docker
The pick for Dockerdocker-sandbox
Skill · joelhooks
The pick for Dockerazure-kubernetes
Skill · microsoft
The pick for Kubernetesinfra-containers-kubernetes
Skill · agents-inc
The pick for Kubernetes