Triaging visual review runs

SkillWeb & browsing

This skill lets an agent inspect PostHog Visual Review runs, the screenshot regression checks that gate PR merges. It gathers evidence from the run, judges whether a snapshot diff is real or flaky, and helps resolve the check by quarantining flakes, tolerating noise, or approving intended changes. Finalizing baselines always requires explicit human confirmation.

Use Triaging visual review runs in Claude, ChatGPT or Ahel Desktop

Free. Sign in, add Triaging visual review runs and connect your AI. About a minute.

Also: Claude Code · Cursor · Codex

Then ask your AI: use the Triaging visual review runs skill

Details

Instructions available. Your AI can read the instructions. Execution depends on the setup they require.

Have access to the PostHog project that runs Visual Review for the repository.

Triaging visual review runsStart free

What your AI can do with it

  • Inspect a Visual Review run's status, scope, and history
  • Judge whether a snapshot diff is real or flaky
  • Quarantine a flaky story or lift an existing quarantine
  • Tolerate noisy snapshot changes
  • Check whether a story has been changing across runs
  • Approve intended visual changes, pending human confirmation

Getting started

  1. Have access to the PostHog project that runs Visual Review for the repository.
  2. Make sure the agent can reach the Visual Review MCP tools.
  3. Point the agent at a PR with a failing visual-review status check or an open VR run on the current branch.
  4. Ask the agent to triage the run and explain what changed visually.
  5. Confirm explicitly before any baseline is finalized.

What this skill tells your AI

The instructions your AI receives, as published by posthog/posthog in products/visual_review/skills/triaging-visual-review-runs/SKILL.md and read by ahel’s review.

Visual Review is PostHog's screenshot-regression product: CI captures storybook + playwright screenshots, diffs them against committed baseline hashes, and gates the PR until every changed snapshot is resolved. A PR with visual changes carries a visual-review GitHub status check and a required "Visual regression tests pass" job check. Both stay red until each diffed snapshot is approved and finalized, tolerated, or quarantined, and the job re-runs: finalize does that for approvals, recompute-create for tolerations and quarantines. The VR UI offers the same actions. A PR from a fork is the exception: it gets no Visual Review run at all. See Fork PRs have no Visual Review run.

This skill teaches an agent how to answer the questions a human reviewer would actually ask, by chaining the VR MCP tools — instead of reaching for gh pr view and tab-hopping to the VR web UI. The read tools cover status / scope / history / triage. An agent may resolve flakes on its own, with a quarantine or, rarely, a toleration. It may not ship a visual change on its own: finalize-create commits the baseline and needs explicit per-run human confirmation.

Decide first

Gather the evidence with Is the diff real or unrelated? and the flake check, then take the first row that matches each changed snapshot.

EvidenceActionHuman yes needed
PR comes from a fork (isCrossRepository: true)Report only. See Fork PRsNo VR writes possible
The diff comes from your change and is intendedapprove-create, then ask for finalizeYes, for each run, before finalize
The diff comes from your change and is not intendedFix the code and push. No VR writeNo
Your change renders a quarantined story changed or new, and the change is intendedapprove-create for that identifier, then ask for finalize. If the change also fixes the flake, add lift-on-merge-create after the approval. See Quarantined storiesYes, for each run, before finalize
Your change fixes a quarantined story's flake, and the story renders unchangedlift-on-merge-create for that identifier. See Quarantined storiesNo
Story outside your change, flakiness entry unstable, hard_count ≥ 5, last_flaked_at in the last 7 daysquarantine-create, then recompute-create. Report itNo
Story outside your change, flakiness entry brokenDo not quarantine. Its baseline on the default branch is wrong. Report it and recommend a re-baselineYes
Story outside your change, quiet history, same width and height, change_kind: pixel, a noise source you can nametolerate-create, then recompute-createNo
Anything else: unstable with fewer failures, a real-looking change you did not make, unsureStop and report what you sawYes

Each theme is its own identifier. Judge the --light and --dark snapshots of a story separately, and quarantine only the ones that match.

Expect tolerations to be rare. The diff already absorbs most real render noise below the threshold, and a small diff percentage is often a real structural change. A toleration accepts one exact hash forever and cannot be undone through the API, so it is never a way past a gate. A story that renders differently from run to run and meets the quarantine row above gets a quarantine, which also protects every other developer. With less evidence, report it instead.

Quarantined stories in your run

A quarantine hides a story's diff from the gate and from the PR comment, so a green check does not show that your change left a quarantined story alone. The story still renders and is diffed on every run that selects it. A PR run renders only the stories its diff affects. Check every run of a change that touches UI. List the changed quarantined snapshots with posthog:visual-review-runs-snapshots-list { id: <run_id>, include_quarantined: true, exclude_unchanged: true }. With exclude_unchanged, quarantined_count counts only the changed ones. The list is paginated and does not put quarantined rows first, so follow next until every row is read.

  • A quarantined story that your change renders differently needs its new picture approved by identifier, then finalized. "Approve all" and approve_all skip quarantined snapshots. Without the approval, the default branch keeps the old entry, and every run fails on the day the quarantine is lifted or expires.
  • A quarantined story that your change does not touch can still show changed, because it is flaky. Leave it.
  • A fix for the flake changes nothing VR can see in one run, so the story renders unchanged and the list above leaves it out. Record the fix with posthog:visual-review-runs-lift-on-merge-create { id: <run_id>, identifier: <identifier> } for each identifier the fix should release, and name the identifiers in the PR description. The quarantine lifts only after the PR merges and a default-branch run that contains the merge renders the same picture against a matching entry. Until then it stays, unless its expiry date passes first, and posthog:visual-review-runs-quarantine-lifts-list { id: <run_id> } shows each request's state and detail.
  • A change that deletes a quarantined story leaves its baseline entry behind. Only a full run classifies the story removed, and only finalize prunes the entry. The run-ci-frontend label takes effect on the next push or ready-for-review, not when it is added, so push after labeling and check that the new run is full. Finalize that run before the merge, or every full run reports the story removed once the quarantine ends.
  • Requesting a lift never approves a picture. For a changed or new quarantined snapshot, approve it by identifier first, or the request returns 400. Finalize the run too, so the baseline entry the lift checks lands with the merge.
  • One clean render does not prove a rare flake is gone, and neither does variant_count: 0, which counts only absorbed variants.

When this skill applies

Trigger this skill on any of:

  • A PR number, branch name, or commit SHA paired with words like visual review, VR, snapshot, screenshot, storybook diff, playwright snapshot, baseline, approve, tolerated, quarantine.
  • Questions about why a PR is blocked, what visually changed, or whether a diff is real.
  • "Is my run done?" / "What's left to review?" / "Has this story flaked recently?"
  • A failing visual-review GitHub check or a PR comment from the posthog-bot mentioning visual review.
  • A failing Visual regression tests pass check on a PR from a fork, which is the offline fallback and not a VR run.

When the user asks for the rendered diff image itself, the VR web UI is faster — direct them there. This skill is for everything around the diff: status, scope, history, triage.

First, check whether the PR comes from a fork. A fork PR has no Visual Review run, so every run-scoped VR tool below returns nothing for it. Only the repo-scoped flakiness tool still answers. Read the flag before you query the tools:

gh pr view <n> --json isCrossRepository

If isCrossRepository is true, stop here. Go to Fork PRs have no Visual Review run.

Fork PRs have no Visual Review run

Visual Review needs a secret, and CI does not give a secret to a fork. The Visual Review upload is therefore skipped. No run, no snapshot row and no visual-review check exists for the PR. That is the designed behavior, not a fault.

Each Storybook shard instead compares its own screenshots with the committed baseline file frontend/snapshots.yml, offline, in the Verify snapshots against the baseline offline step. A mismatch fails the Visual regression tests pass check.

The offline fallback is weaker than Visual Review in ways that change the triage:

Visual ReviewOffline fallback on a fork
Diffs with a noise thresholdExact pixel hash match only
Knows tolerated alternate hashesKnows none — a tolerated variant still fails
Applies quarantineApplies none — a quarantined story still fails
Rendered diff images in the VR UINo images; the job log names the snapshots that differ
Triage and finalize through the MCP toolsNo tools apply; a maintainer updates the baseline file instead

So a flaky story can fail a fork PR that changes nothing visible.

How to triage a fork PR failure:

  1. Read the failing Visual regression tests pass job. The step summary and the Baseline mismatch error name each snapshot that differs. gh run view <run_id> --log-failed gets the log.
  2. Run the scope check from Is the diff real or unrelated? against the named identifiers. It needs only git diff, so it works without a run.
  3. Judge flakiness from the default branch, not from the fork PR. Use posthog:visual-review-repos-flakiness-retrieve { id: <repo_id> }, with the repo id from posthog:visual-review-repos-list. This is the one VR tool that still helps, because it reports repo-level history and does not need a run.
  4. Report the verdict and stop. You cannot approve, tolerate or finalize anything, because there is no run to act on. A snapshot that must change needs a maintainer to update frontend/snapshots.yml on the PR branch, with a Visual Review run on an in-repo branch.

Never push a fork's head to an in-repo branch to get it a Visual Review run. See Pull requests from forks.

Tools

Read tools (safe to call freely):

ToolPurpose
posthog:visual-review-runs-listList runs, filter by pr_number / commit_sha / branch / review_state. Start here.
posthog:visual-review-runs-retrieveFull detail for a single run (status, summary counts, supersession). Carries the repo_id the repo tools need.
posthog:visual-review-runs-snapshots-listPer-snapshot results inside a run: identifier, result, diff %, classification, baseline + current artifact URLs. Quarantined snapshots are excluded by default (see quarantined_count); pass include_quarantined=true to see them. Pass exclude_unchanged=true to skip the thousands of unchanged rows a large run holds.
posthog:visual-review-repos-flakiness-retrieveRepo snapshots whose rendering is untrusted, with a flakiness_state and flake rates for each. The flake check starts here.
posthog:visual-review-runs-snapshot-history-listOne story's baseline timeline on the default branch: one row per baseline change. Takes { id: <run_id>, identifier: <identifier> }.
posthog:visual-review-runs-counts-retrieveAggregate counts for queue triage (how many runs in needs_review, etc.).
posthog:visual-review-runs-tolerated-hashes-listHashes the team has explicitly accepted as "known flake / acceptable variation". Takes the same two parameters as the history tool.
posthog:visual-review-repos-listRepos (one per GitHub repo) — usually only one matters; useful for filtering.
posthog:visual-review-repos-retrieveRepo metadata: baseline file paths, PR-comment configuration.
posthog:visual-review-repos-quarantine-listActive quarantines with reason, author, expiry and source run. Pass identifier for its full history.
posthog:visual-review-repos-toleration-pileups-retrieveStories that keep getting tolerated: candidates for a fix in the story.
posthog:visual-review-runs-quarantine-lifts-listRequests to lift a quarantine when the run's PR merges, with state and the latest check's detail. Takes { id: <run_id> }.

Triage tools (they do NOT change the baseline; the gate changes only after recompute-create, except a lift on merge, which a default-branch run applies on its own and needs no recompute):

ToolPurpose
posthog:visual-review-runs-approve-createMark changed / new snapshots reviewed (approved) in the DB. Does NOT commit or green the gate — ship via finalize.
posthog:visual-review-runs-tolerate-createAccept one changed snapshot's current hash as an alternate in every future run. Render noise only, see Decide first. Cannot be undone through the API.
posthog:visual-review-repos-quarantine-createRemove one identifier of one run type from pass or fail on every PR until it expires (30 days if expires_at is omitted). Undo with quarantine-expire-create.
posthog:visual-review-repos-quarantine-expire-createLift a quarantine, so the story gates runs again.
posthog:visual-review-runs-lift-on-merge-createLift a quarantine once the run's PR merges and a default-branch run renders the snapshot's picture against a matching entry. Never approves a picture. Takes { id: <run_id>, identifier }.
posthog:visual-review-runs-quarantine-lifts-cancel-createWithdraw a pending lift on merge. The quarantine stays. Takes { id: <run_id>, request_id }.
posthog:visual-review-runs-recompute-createRecount a completed, unfinalized run, post the visual-review status, and re-run the CI job recorded on the run, so the required check reads the new verdict.

Branch protection requires the Visual regression tests pass and Playwright tests pass job checks, not the visual-review status. So a quarantine or toleration unblocks the PR only after recompute-create re-runs the CI job recorded on the run. Approved changes keep the gate red until finalize commits them, so recompute never ships an approval.

Ship tool (irreversible, outward-facing — requires explicit per-run human confirmation; see the gate):

ToolPurpose
posthog:visual-review-runs-finalize-createCommit the approved baseline to the PR branch and green the GitHub visual-review check. This ships the change.

Mark-reviewed call shape (approve-create):

  • id (required) — the run UUID. It's the route parameter, so the call fails without it.
  • snapshots: [{identifier, new_hash}] — new_hash is the content_hash of each snapshot's current_artifact. This only records the review in the DB; nothing is committed and the gate stays red until you finalize.

Quarantine call shape (quarantine-create):

  • id — the repo UUID (the run's repo_id), and run_type — the failing run's run_type. Both are route parameters.
  • identifier, and a reason with the evidence, for example "unstable on master: 7 failed default-branch runs in 7 days, chart animation timing; unrelated to PR 1234".
  • source_run_id — the run that failed. expires_at — omit it for 30 days, or set an earlier date, never a later one.
  • notify_owners: true — posts the quarantine to the owning team's Slack channel, so the people who must fix the story hear about it.
  • Then call recompute-create { id: <run_id> } on the PR's newest non-stale run of the same run_type, and check ci_rerun_triggered.

Toleration call shape — both fields are required:

  • id (required) — the run UUID. It's the route parameter, so the call fails without it.
  • snapshot_id (required) — the UUID of the individual snapshot to tolerate (from visual-review-runs-snapshots-list). This identifies which snapshot inside the run; it does not replace the run id.

Finalize call shape (finalize-create) — the all-or-nothing ship action:

Shortened here. Read the whole file on GitHub.

Signals

GitHub stars
40k
Forks
3k
Last commit
Sep 2026

Questions

Why is my PR blocked?
A PR with visual changes carries a visual-review status check and a required Visual regression tests pass job check. Both stay red until every diffed snapshot is approved and finalized, tolerated, or quarantined, and the job re-runs.
What changed visually?
The skill inspects the Visual Review run to show which snapshots differ from their committed baselines, so the agent can report the changed stories and screenshots.
Can the agent approve a visual change on its own?
No. The agent may resolve flakes with a quarantine or, rarely, a toleration, but finalizing a baseline with finalize-create needs explicit per-run human confirmation.
Advanced
Item type
skill
Key
triaging-visual-review-runs
Source
github.com/posthog/posthog