Fix Integration workflow failures

SkillProductivity

Diagnose and fix a failing scheduled (or standalone) TMDb Integration workflow run — re-run transients, or fix real drift on a branch off main, then PR and (optionally) merge. Use for the weekly Sunday Integration cron failure, or any red Integration run not tied to an open PR.

Available today. Use it from your connected AI after setup.

Connect ahel once, and every AI you use reads what you have installed.

Then ask your AI: use the Fix Integration workflow failures skill

What this skill tells your AI

The instructions your AI receives, as published by adamayoung/tmdb in .claude/skills/fix-integration-failures/SKILL.md and read by ahel’s review.

The Integration workflow (.github/workflows/integration.yml) runs the live-API integration suite on a weekly schedule (cron: '0 0 * * 0' — Sunday 00:00 UTC), as well as on PRs/pushes. The scheduled run is a live-API canary: nothing in the code changed, so a failure means the live TMDb API drifted, a test's assumed data went stale, or a transient error/rate-limit hit. This skill takes such a failure from red run → green main: diagnose it, then either re-run a transient or fix the real cause on its own branch and open (and optionally merge) a PR.

Scope. This is for an Integration failure not attached to an open feature PR — the scheduled canary, a workflow_dispatch, or a failure on main. When a failing Integration check belongs to an open PR, /watch-pr and /fix-pr-checks own that; they delegate the pre-existing/unrelated case back here (fix on a branch off main), which is exactly this skill's job.

Mode — check the arguments passed to this skill (shown at the end). If they include merge (e.g. /fix-integration-failures merge), auto-merge the fix PR once green. Otherwise stop at ready-to-merge and hand off (default).

Agent behaviour contract

  1. Never edit main directly (CLAUDE.md). Any fix lands on a fix/<slug> branch off main, via a PR.
  2. Diagnose before fixing. Always run /diagnose-integration-failure first; let its ranked cause drive the fix. Don't guess. Check the diagnosis carries its evidence. A shape-change (cause 1) or drifted-data (cause 2) conclusion must arrive with an observed: line — the live call that saw today's response. Missing one, re-run the diagnosis once asking for it. Still missing → treat that cause as unverified: proceed, but say so in the PR body and in claude-analysis.md, and never describe it as confirmed drift. Re-run once, not in a loop — a headless run has no MCP and legitimately cannot observe (it reports observed: unavailable (headless)), so an unbounded retry would spin forever. Transient (cause 3) and in-diff regression (cause 0) carry no observed: line by design; do not demand one.
  3. Distinguish transient from real. A re-run is the cheap test. Only open a PR for a cause that survives a re-run (or is clearly a data/shape drift).
  4. Test-first for real fixes. A model/decoder fix follows canon-tdd (failing unit test + fixture, then the fix). A drifted-assertion fix updates the integration test to assert behaviour, not a brittle exact value.
  5. The gate is make ci. Never open the PR until it passes locally.
  6. Don't paper over a real regression. If the cause is a genuine library bug (not drift/transient), fix it properly or stop and report — never just relax a test to hide it.

0. Find the failing run

Interactive vs headless. The steps below use the GitHub MCP (mcp__github__*, owner/repo from the origin remote). When this skill runs headless from integration-failure.yml — a CI runner, where the user-scoped MCP is not mounted — use the gh equivalents instead (given inline at each step and in Running headless below); that path stays 100% gh/git.

Find the failing run with mcp__github__actions_list method list_workflow_runs (owner/repo from origin, resource_id: integration.yml, workflow_runs_filter: { event: schedule, status: completed }). The status enum has no failure value, so filter the results to conclusion == "failure" yourself.

  • Prefer the most recent schedule (or workflow_dispatch) run on main.
  • Note its run id and event (sets the cause ranking).
  • No failing run → report "Integration is green, nothing to fix" and stop.
  • Headless: gh run list --workflow Integration --status failure --limit 5 --json databaseId,event,headBranch,conclusion,createdAt,displayTitle.

1. Diagnose

Invoke /diagnose-integration-failure, handing it the run id (it fetches the failed-job logs — via mcp__github__get_job_logs, or gh run view --log-failed when headless). It returns the three-section analysis — Summary, Likely cause (ranked; for a scheduled run it leads with backend/data drift, not a code regression), and Suggested fix. Use that ranking to choose the path below.

2. Transient? Re-run first

If the diagnosis points to case 3 (HTTP 429 / rate-limit, a timeout near the 30-min cap, or a truncated log with no assertion failure):

Re-run the failed jobs with mcp__github__actions_run_trigger method rerun_failed_jobs (owner/repo from origin, run_id: <id>), then block on the re-run with gh run watch <run-id> (the MCP has no blocking-wait equivalent). After it returns, re-read the conclusion with mcp__github__actions_get method get_workflow_run (resource_id: <id>) — don't trust the rerun call to surface it.

gh run watch <run-id>   # blocking wait — kept on gh
  • Green on re-run → it was transient. Report and stop; no PR needed.
  • Fails the same way again → treat as deterministic; go to §3.
  • Headless: gh run rerun <run-id> --failed then gh run watch <run-id>.

(The integration client already retries 429/5xx with backoff, so a true transient that survives a re-run is uncommon — a repeat failure is usually real drift.)

3. Fix the real cause on a branch off main

Reproduce locally first to confirm and to get a fast edit loop. Branch off origin/main directly — never git checkout main first. This is worktree-safe: when this skill is invoked from a /deliver worktree (via /watch-pr §1c), main is usually checked out in the main working copy, so git checkout main fails with fatal: 'main' is already used by worktree … — the same trap /pr documents for its rebase step:

git fetch origin
git checkout -b fix/<slug> origin/main   # e.g. fix/<service>-integration-drift

Invoked mid-/deliver (from /watch-pr)? Don't create the fix branch in the deliverable's worktree — that switches its checkout away from the feature branch being watched. Give the fix its own worktree instead and work there, removing it once the fix PR merges:

git worktree add .claude/worktrees/fix-<slug> -b fix/<slug> origin/main

Run the failing suite locally to reproduce — /integration-test (or swift test --filter <Suite>/<test>). Then fix per the diagnosis:

  • TMDb backend / response-shape change (a field added, removed, renamed, or now nullable) → fix the Swift model / CodingKeys / fixture test-first (canon-tdd): add a failing unit test + a JSON fixture matching the current live shape. The diagnosis's observed: line tells you which endpoint drifted and how, but it is a one-line summary (and unavailable headless) — not a response body, so it cannot source a fixture. Fetch the real response with mcp__tmdb__* and build the fixture from that, cross-checking the OpenAPI spec for the documented shape — then the model fix.
  • Stale assumed data (a test asserts a specific live title, count, id, date, or ordering that drifted) → relax the integration assertion to verify behaviour, not a brittle exact value (e.g. assert non-empty / a stable property / >= 1 rather than an exact count). Match the robust pattern already used by sibling tests in the same suite.
  • A genuine library regression → fix it properly, test-first. If it is not a quick, confident fix, stop and report — don't merge a workaround.

Verify: /integration-test green for the touched suite, then make ci (mandatory full gate). If make ci is red, fix and re-run — never PR on red.

4. Open the PR (and optionally merge)

Run /pr to commit (gitmoji), push, and open the PR (🐛//♻️ as fits; make ci runs again inside /pr). Then run /watch-pr to drive it to ready:

  • Default (no merge)/watch-pr watch-only: report the ready PR URL and stop for the user to merge.
  • merge/watch-pr merge: squash-merge once green, then report.

/watch-pr handles any further transient flakes on the PR's own checks (re-run), so a live-API hiccup during CI won't strand the fix.

5. Report

Close with: the run id that failed, the diagnosis verdict (transient vs real), what you changed (file + one line) or that a re-run cleared it, the PR URL and whether it merged, and anything left for the user (e.g. a real regression you chose not to auto-fix).

Running headless (from integration-failure.yml)

When the Integration Failure Alert workflow invokes this skill on a scheduled failure, it runs non-interactively, so adapt:

  • A failing-step log is already at failure-log.txt in the workspace — diagnose from it (pass it to /diagnose-integration-failure); don't re-download logs.
  • Verify with the targeted suite + a buildswift build --build-tests and swift test --filter <Suite>/<test>not the full make ci. The opened PR's own CI (ci.yml + integration.yml) is the authoritative gate; the alert job need not replicate the whole pinned-lint/xcsift toolchain.
  • Open the PR with git/gh directly (git checkout -b → commit → push → gh pr create), not /pr/pr runs the full make ci, which the lightweight alert job can't satisfy. The format/lint hooks still reshape files on edit, so the diff stays clean.
  • Then STOP — do not run /watch-pr and do not merge (a human reviews it). Write the diagnosis to claude-analysis.md and, if you open a PR, its URL on one line to pr-url.txt.
  • If the cause is transient (re-run territory) or a genuine regression you should not auto-fix, open no PR — explain in claude-analysis.md so the alert issue carries it.

Guardrails

  • Never edit .github/workflows/* to force a check green, and never force-push, without surfacing to the user first.
  • One failing run can surface several drifted assertions — fix them together in one PR, but keep each change minimal and behaviour-preserving.
  • If the live API is broadly down/throttled (many unrelated suites failing fast), that's an outage, not a fix target — report and wait it out.
  • Capture anything durable you learned (a new live-API shape, a recurring drift) with /capture-knowledge before the PR, so the fixture/notes land with it.

Arguments: $ARGUMENTS

Signals

GitHub stars
176
Forks
47
Last commit
Aug 2026
Advanced
Catalog kind
skill
Gateway key
fix-integration-failures
Source
github.com/adamayoung/tmdb