Triaging visual review runs
SkillWeb & browsingLets your agent inspect PostHog screenshot checks, figure out why a PR is blocked, and flag flaky snapshots.
Use Triaging visual review runs in Claude, ChatGPT or Ahel Desktop
Free. Sign in, add Triaging visual review runs and connect your AI. About a minute.
Also: Claude Code · Cursor · Codex
Then ask your AI: use the Triaging visual review runs skill
Details
Instructions available. Your AI can read the instructions. Execution depends on the setup they require.
Account requirements not reviewed. Check the skill instructions before use; ahel provides instructions and does not run this skill.
No other account needed.
Add ahel to your AI once: Claude, ChatGPT, Cursor, Claude Code or Codex. Then ask it to use this.
About this skill
Inspects PostHog Visual Review (VR) runs that gate PR merges with screenshot regression checks. Use when the user mentions "visual review", "VR", "snapshot diff", "screenshot test", "storybook regression", "playwright snapshot", asks why a PR is blocked or what changed visually, wants to triage the
What this skill tells your AI
The instructions your AI receives, as published by posthog/skills in skills/omnibus/triaging-visual-review-runs/SKILL.md and read by ahel’s review.
Visual Review is PostHog's screenshot-regression product: CI captures storybook + playwright screenshots,
diffs them against committed baseline hashes, and gates the PR until every changed snapshot is resolved.
A PR with visual changes carries a visual-review GitHub status check and a required "Visual regression tests pass" job check.
Both stay red until each diffed snapshot is approved and finalized, tolerated, or quarantined, and the job re-runs:
finalize does that for approvals, recompute-create for tolerations and quarantines. The VR UI offers the same actions.
A PR from a fork is the exception: it gets no Visual Review run at all.
See Fork PRs have no Visual Review run.
This skill teaches an agent how to answer the questions a human reviewer would actually ask, by chaining
the VR MCP tools — instead of reaching for gh pr view and tab-hopping to the VR web UI. The read tools
cover status / scope / history / triage. An agent may resolve flakes on its own, with a quarantine or, rarely, a toleration.
It may not ship a visual change on its own: finalize-create commits the baseline and needs explicit per-run human confirmation.
Decide first
Gather the evidence with Is the diff real or unrelated? and the flake check, then take the first row that matches each changed snapshot.
| Evidence | Action | Human yes needed |
|---|---|---|
PR comes from a fork (isCrossRepository: true) | Report only. See Fork PRs | No VR writes possible |
| The diff comes from your change and is intended | approve-create, then ask for finalize | Yes, for each run, before finalize |
| The diff comes from your change and is not intended | Fix the code and push. No VR write | No |
Your change renders a quarantined story changed or new, and the change is intended | approve-create for that identifier, then ask for finalize. If the change also fixes the flake, add lift-on-merge-create after the approval. See Quarantined stories | Yes, for each run, before finalize |
Your change fixes a quarantined story's flake, and the story renders unchanged | lift-on-merge-create for that identifier. See Quarantined stories | No |
Story outside your change, flakiness entry unstable, hard_count ≥ 5, last_flaked_at in the last 7 days | quarantine-create, then recompute-create. Report it | No |
Story outside your change, flakiness entry broken | Do not quarantine. Its baseline on the default branch is wrong. Report it and recommend a re-baseline | Yes |
Story outside your change, quiet history, same width and height, change_kind: pixel, a noise source you can name | tolerate-create, then recompute-create | No |
Anything else: unstable with fewer failures, a real-looking change you did not make, unsure | Stop and report what you saw | Yes |
Each theme is its own identifier.
Judge the --light and --dark snapshots of a story separately, and quarantine only the ones that match.
Expect tolerations to be rare. The diff already absorbs most real render noise below the threshold, and a small diff percentage is often a real structural change. A toleration accepts one exact hash forever and cannot be undone through the API, so it is never a way past a gate. A story that renders differently from run to run and meets the quarantine row above gets a quarantine, which also protects every other developer. With less evidence, report it instead.
Quarantined stories in your run
A quarantine hides a story's diff from the gate and from the PR comment, so a green check does not show that your change left a quarantined story alone.
The story still renders and is diffed on every run that selects it. A PR run renders only the stories its diff affects.
Check every run of a change that touches UI.
List the changed quarantined snapshots with
posthog:visual-review-runs-snapshots-list { id: <run_id>, include_quarantined: true, exclude_unchanged: true }.
With exclude_unchanged, quarantined_count counts only the changed ones.
The list is paginated and does not put quarantined rows first, so follow next until every row is read.
- A quarantined story that your change renders differently needs its new picture approved by identifier, then finalized.
"Approve all" and
approve_allskip quarantined snapshots. Without the approval, the default branch keeps the old entry, and every run fails on the day the quarantine is lifted or expires. - A quarantined story that your change does not touch can still show
changed, because it is flaky. Leave it. - A fix for the flake changes nothing VR can see in one run, so the story renders
unchangedand the list above leaves it out. Record the fix withposthog:visual-review-runs-lift-on-merge-create { id: <run_id>, identifier: <identifier> }for each identifier the fix should release, and name the identifiers in the PR description. The quarantine lifts only after the PR merges and a default-branch run that contains the merge renders the same picture against a matching entry. Until then it stays, unless its expiry date passes first, andposthog:visual-review-runs-quarantine-lifts-list { id: <run_id> }shows each request'sstateanddetail. - A change that deletes a quarantined story leaves its baseline entry behind.
Only a full run classifies the story
removed, and only finalize prunes the entry. Therun-ci-frontendlabel takes effect on the next push or ready-for-review, not when it is added, so push after labeling and check that the new run is full. Finalize that run before the merge, or every full run reports the storyremovedonce the quarantine ends. - Requesting a lift never approves a picture. For a
changedornewquarantined snapshot, approve it by identifier first, or the request returns 400. Finalize the run too, so the baseline entry the lift checks lands with the merge. - One clean render does not prove a rare flake is gone, and neither does
variant_count: 0, which counts only absorbed variants.
When this skill applies
Trigger this skill on any of:
- A PR number, branch name, or commit SHA paired with words like visual review, VR, snapshot, screenshot, storybook diff, playwright snapshot, baseline, approve, tolerated, quarantine.
- Questions about why a PR is blocked, what visually changed, or whether a diff is real.
- "Is my run done?" / "What's left to review?" / "Has this story flaked recently?"
- A failing
visual-reviewGitHub check or a PR comment from theposthog-botmentioning visual review. - A failing
Visual regression tests passcheck on a PR from a fork, which is the offline fallback and not a VR run.
When the user asks for the rendered diff image itself, the VR web UI is faster — direct them there. This skill is for everything around the diff: status, scope, history, triage.
First, check whether the PR comes from a fork. A fork PR has no Visual Review run, so every run-scoped VR tool below returns nothing for it. Only the repo-scoped flakiness tool still answers. Read the flag before you query the tools:
gh pr view <n> --json isCrossRepository
If isCrossRepository is true, stop here.
Go to Fork PRs have no Visual Review run.
Fork PRs have no Visual Review run
Visual Review needs a secret, and CI does not give a secret to a fork.
The Visual Review upload is therefore skipped.
No run, no snapshot row and no visual-review check exists for the PR.
That is the designed behavior, not a fault.
Each Storybook shard instead compares its own screenshots with the committed baseline file frontend/snapshots.yml, offline, in the Verify snapshots against the baseline offline step.
A mismatch fails the Visual regression tests pass check.
The offline fallback is weaker than Visual Review in ways that change the triage:
| Visual Review | Offline fallback on a fork |
|---|---|
| Diffs with a noise threshold | Exact pixel hash match only |
| Knows tolerated alternate hashes | Knows none — a tolerated variant still fails |
| Applies quarantine | Applies none — a quarantined story still fails |
| Rendered diff images in the VR UI | No images; the job log names the snapshots that differ |
| Triage and finalize through the MCP tools | No tools apply; a maintainer updates the baseline file instead |
So a flaky story can fail a fork PR that changes nothing visible.
How to triage a fork PR failure:
- Read the failing
Visual regression tests passjob. The step summary and theBaseline mismatcherror name each snapshot that differs.gh run view <run_id> --log-failedgets the log. - Run the scope check from Is the diff real or unrelated? against the named identifiers.
It needs only
git diff, so it works without a run. - Judge flakiness from the default branch, not from the fork PR.
Use
posthog:visual-review-repos-flakiness-retrieve { id: <repo_id> }, with the repo id fromposthog:visual-review-repos-list. This is the one VR tool that still helps, because it reports repo-level history and does not need a run. - Report the verdict and stop.
You cannot approve, tolerate or finalize anything, because there is no run to act on.
A snapshot that must change needs a maintainer to update
frontend/snapshots.ymlon the PR branch, with a Visual Review run on an in-repo branch.
Never push a fork's head to an in-repo branch to get it a Visual Review run. See Pull requests from forks.
Tools
Read tools (safe to call freely):
| Tool | Purpose |
|---|---|
posthog:visual-review-runs-list | List runs, filter by pr_number / commit_sha / branch / review_state. Start here. |
posthog:visual-review-runs-retrieve | Full detail for a single run (status, summary counts, supersession). Carries the repo_id the repo tools need. |
posthog:visual-review-runs-snapshots-list | Per-snapshot results inside a run: identifier, result, diff %, classification, baseline + current artifact URLs. Quarantined snapshots are excluded by default (see quarantined_count); pass include_quarantined=true to see them. Pass exclude_unchanged=true to skip the thousands of unchanged rows a large run holds. |
posthog:visual-review-repos-flakiness-retrieve | Repo snapshots whose rendering is untrusted, with a flakiness_state and flake rates for each. The flake check starts here. |
posthog:visual-review-runs-snapshot-history-list | One story's baseline timeline on the default branch: one row per baseline change. Takes { id: <run_id>, identifier: <identifier> }. |
posthog:visual-review-runs-counts-retrieve | Aggregate counts for queue triage (how many runs in needs_review, etc.). |
posthog:visual-review-runs-tolerated-hashes-list | Hashes the team has explicitly accepted as "known flake / acceptable variation". Takes the same two parameters as the history tool. |
posthog:visual-review-repos-list | Repos (one per GitHub repo) — usually only one matters; useful for filtering. |
posthog:visual-review-repos-retrieve | Repo metadata: baseline file paths, PR-comment configuration. |
posthog:visual-review-repos-quarantine-list | Active quarantines with reason, author, expiry and source run. Pass identifier for its full history. |
posthog:visual-review-repos-toleration-pileups-retrieve | Stories that keep getting tolerated: candidates for a fix in the story. |
posthog:visual-review-runs-quarantine-lifts-list | Requests to lift a quarantine when the run's PR merges, with state and the latest check's detail. Takes { id: <run_id> }. |
Triage tools (they do NOT change the baseline; the gate changes only after recompute-create, except a lift on merge, which a default-branch run applies on its own and needs no recompute):
| Tool | Purpose |
|---|---|
posthog:visual-review-runs-approve-create | Mark changed / new snapshots reviewed (approved) in the DB. Does NOT commit or green the gate — ship via finalize. |
posthog:visual-review-runs-tolerate-create | Accept one changed snapshot's current hash as an alternate in every future run. Render noise only, see Decide first. Cannot be undone through the API. |
posthog:visual-review-repos-quarantine-create | Remove one identifier of one run type from pass or fail on every PR until it expires (30 days if expires_at is omitted). Undo with quarantine-expire-create. |
posthog:visual-review-repos-quarantine-expire-create | Lift a quarantine, so the story gates runs again. |
posthog:visual-review-runs-lift-on-merge-create | Lift a quarantine once the run's PR merges and a default-branch run renders the snapshot's picture against a matching entry. Never approves a picture. Takes { id: <run_id>, identifier }. |
posthog:visual-review-runs-quarantine-lifts-cancel-create | Withdraw a pending lift on merge. The quarantine stays. Takes { id: <run_id>, request_id }. |
posthog:visual-review-runs-recompute-create | Recount a completed, unfinalized run, post the visual-review status, and re-run the CI job recorded on the run, so the required check reads the new verdict. |
Branch protection requires the Visual regression tests pass and Playwright tests pass job checks, not the visual-review status.
So a quarantine or toleration unblocks the PR only after recompute-create re-runs the CI job recorded on the run.
Approved changes keep the gate red until finalize commits them, so recompute never ships an approval.
Ship tool (irreversible, outward-facing — requires explicit per-run human confirmation; see the gate):
| Tool | Purpose |
|---|---|
posthog:visual-review-runs-finalize-create | Commit the approved baseline to the PR branch and green the GitHub visual-review check. This ships the change. |
Mark-reviewed call shape (approve-create):
id(required) — the run UUID. It's the route parameter, so the call fails without it.snapshots: [{identifier, new_hash}]—new_hashis thecontent_hashof each snapshot'scurrent_artifact. This only records the review in the DB; nothing is committed and the gate stays red until you finalize.
Quarantine call shape (quarantine-create):
id— the repo UUID (the run'srepo_id), andrun_type— the failing run'srun_type. Both are route parameters.identifier, and areasonwith the evidence, for example "unstable on master: 7 failed default-branch runs in 7 days, chart animation timing; unrelated to PR 1234".source_run_id— the run that failed.expires_at— omit it for 30 days, or set an earlier date, never a later one.notify_owners: true— posts the quarantine to the owning team's Slack channel, so the people who must fix the story hear about it.- Then call
recompute-create { id: <run_id> }on the PR's newest non-stale run of the samerun_type, and checkci_rerun_triggered.
Toleration call shape — both fields are required:
id(required) — the run UUID. It's the route parameter, so the call fails without it.snapshot_id(required) — the UUID of the individual snapshot to tolerate (fromvisual-review-runs-snapshots-list). This identifies which snapshot inside the run; it does not replace the runid.
Finalize call shape (finalize-create) — the all-or-nothing ship action:
Shortened here. Read the whole file on GitHub.
Signals
- GitHub stars
- 72
- Forks
- 7
- Last commit
- Oct 2026
ahel recommends instead
Advanced
- Item type
- skill
- Key
triaging-visual-review-runs-2- Source
- github.com/posthog/skills
More in Web & browsing
Skill · browser-use
More in Web & browsingwebapp-testing
Skill · anthropics
More in Web & browsingplaywright-cli
Skill · microsoft
More in Web & browsingbenchmark
Skill · affaan-m
More in Web & browsingopen-source
Skill · browser-use
More in Web & browsingimpeccable
Skill · pbakaus
More in Web & browsing