Testgrid failure analysis
SkillMonitoring & opsFetches logs and support bundles from a failed Testgrid kURL run so failures can be analyzed.
Use Testgrid failure analysis in Claude, ChatGPT or Ahel Desktop
Free. Sign in, add Testgrid failure analysis and connect your AI. About a minute.
Also: Claude Code · Cursor · Codex
Then ask your AI: use the Testgrid failure analysis skill
Details
Instructions available. Your AI can read the instructions. Execution depends on the setup they require.
Account requirements not reviewed. Check the skill instructions before use; Ahel provides instructions and does not run this skill.
No other account needed.
Add Ahel to your AI once: Claude, ChatGPT, Cursor, Claude Code or Codex. Then ask it to use this.
About this skill
Use when analyzing a failed Testgrid kURL run to fetch the run results, failure logs, and encrypted support bundles from the Testgrid API and write them into a directory for offline analysis; trigger with "testgrid failure analysis", "fetch testgrid logs", "get support bundle from Testgrid", or "ana
What this skill tells your AI
The instructions your AI receives, as published by replicatedhq/kurl in .opencode/skills/testgrid-failure-analysis/SKILL.md and read by Ahel’s review.
This skill helps an agent collect the artifacts of a failed Testgrid run so they can be analyzed locally.
What it does
- Queries the Testgrid API for a run by
refId. - Identifies every failed instance (
isSuccess == false, not unsupported, not skipped, and finished). - For each failure, fetches:
- The instance metadata (
instance.json) - The main instance logs (populated when the VM fails to start)
- Sonobuoy results, if any
- The per-node logs from the actual test VMs (
{nodeId}.log.txt) - Any encrypted support bundles whose S3 URLs are printed in the node logs
- The instance metadata (
- Writes everything into a structured output directory ready for an agent to inspect.
Important details from the codebase
- Public API base path is
/api/v1. The endpoints used are:POST /api/v1/run/{refId}— returns the run with itsinstancesarray, plussuccess_countandfailure_count.GET /api/v1/instance/{instanceId}/logs— returns{"logs": "..."}from thetestinstance.outputcolumn.GET /api/v1/instance/{nodeId}/node-logs— returns{"logs": "..."}from theclusternode.outputcolumn.GET /api/v1/instance/{instanceId}/sonobuoy— returns{"results": "..."}.
- The open-source
/api/v1endpoints are not authenticated by default (theapi-tokenauth middleware only protects the runner endpoints under/v1). However, an optional--api-tokenis accepted and sent as HTTP Basic Auth with usernametokenand the provided password, for deployments that add authentication.--api-keyis kept as a deprecated alias for backward compatibility. - Support bundles are collected by the test script (
tgrun/pkg/runner/vmi/embed/runcmd.sh→collect_support_bundle) and uploaded to S3 with the handler atPOST /v1/instance/{instanceId}/bundle. The S3 URL is printed in the node log output, which is why this skill scans the logs for it. - The bundle is encrypted with the
agefile format using a scrypt passphrase. The API stores it with key pattern{instanceId}-{unix}/bundle.tgz.age. The downloaded file keeps the.ageextension. - If you provide the age passphrase, the helper script will try to decrypt each bundle in place with
age -d -p.
Node IDs used by the runner
Testgrid creates one initial-primary node plus optional additional nodes. The node IDs are predictable from the instance ID and the numPrimaryNodes / numSecondaryNodes fields, so the skill tries:
{instanceId}-initialprimary{instanceId}-primary-1...{instanceId}-primary-{numPrimaryNodes-1}{instanceId}-secondary-0...{instanceId}-secondary-{numSecondaryNodes-1}
Only nodes that actually produced logs will be saved.
How to use
Run the helper script shipped with this skill:
python3 .opencode/skills/testgrid-failure-analysis/fetch.py \
--api-endpoint https://api.testgrid.kurl.sh \
--ref-id <RUN_REF_ID> \
--output-dir ./testgrid-analysis/<RUN_REF_ID> \
[--api-token <TOKEN>] \
[--age-passphrase <PASSPHRASE>]
Environment variables are also supported:
TESTGRID_API_TOKEN→--api-token(TESTGRID_API_KEYis still read as a fallback)TESTGRID_AGE_PASSPHRASE→--age-passphrase
Output layout
<output-dir>/
run.json # full run response
<instanceId>/
instance.json # instance metadata
logs.txt # main instance output, if any
sonobuoy.txt # sonobuoy results, if any
<instanceId>-initialprimary.log.txt
bundle-<nodeId>-0.tgz.age # encrypted support bundle
bundle-<nodeId>-0.tgz # decrypted support bundle (if passphrase supplied)
What to do next
After fetching, read the run.json summary, open the per-instance logs, and inspect any decrypted support bundles. If a bundle could not be downloaded, grep the corresponding node log for bundle.tgz.age to find the raw S3 URL.
Signals
- GitHub stars
- 809
- Forks
- 81
- Last commit
- Oct 2026
Advanced
- Item type
- skill
- Key
testgrid-failure-analysis- Source
- github.com/replicatedhq/kurl
More in Monitoring & ops
Skill · anthropics
More in Monitoring & opsagent-eval
Skill · affaan-m
More in Monitoring & opspricing
Skill · coreyhaines31
More in Monitoring & opslark-okr
Skill · larksuite
More in Monitoring & opsdashboard-builder
Skill · affaan-m
More in Monitoring & opsbabysit
Skill · thedotmack
More in Monitoring & ops