ab-test-analysis

SkillProductivity

Lets your agent analyze A/B test results for statistical significance and recommend whether to ship, extend, or stop a variant.

Available today. Use it from your connected AI after setup.

Connect ahel once, and every AI you use reads what you have installed.

Then ask your AI: use the ab-test-analysis skill

About this capability

Use when a task needs analysis of A/B test results, interpretation of p-values and confidence intervals, statistical significance checks, or a principled ship/no-ship decision.

What this skill tells your AI

The instructions your AI receives, as published by jshsakura/awesome-opencode-skills in skills/ab-test-analysis/SKILL.md and read by ahel’s review.

Instructions

Own A/B test analysis as a principled ship/no-ship decision, not a p-value lookup.

Separate statistical significance from practical significance, and catch common analysis traps before recommending action.

Working mode:

  1. Establish the experiment design: primary metric, guardrails, pre-specified minimum detectable effect, and duration.
  2. Validate test integrity (sample ratio match, sufficient power, no peeking) before trusting any result.
  3. Quantify the effect: observed lift, confidence interval, p-value, and whether the result is practically meaningful.
  4. Issue a ship / no-ship / iterate / inconclusive verdict with explicit rationale.

Focus on:

  • p-value meaning: probability of data this extreme under the null, not the probability the variant is better
  • effect size and confidence interval as the decision drivers, not significance alone
  • guardrail metrics: never ship a primary win that significantly harms a guardrail
  • pre-specified segments only; flag post-hoc segmentation as p-hacking
  • common errors: peeking, multiple comparisons, Simpson's paradox, survivorship bias
  • frequentist vs bayesian framing and which the tooling actually used
  • power and sample adequacy when results are trending but underpowered

Quality checks:

  • confirm the verdict aligns with primary metric, effect size, and guardrail status together
  • verify no sample ratio mismatch invalidates the test
  • check the test ran to its predetermined sample size and duration
  • ensure each segment claim was pre-planned, not mined after the fact
  • distinguish "no effect detected" from "underpowered, extend the test"

Return:

  • results summary table (control vs treatment: n, conversion rate, lift, CI, p-value)
  • statistical significance verdict and effect-size interpretation
  • guardrail metric status
  • ship / no-ship / iterate / inconclusive recommendation with rationale
  • next step (ship plan, extension, or learning to capture)

Do not declare a win on statistical significance alone when practical significance or guardrails are unmet unless explicitly requested by the parent agent.

Signals

GitHub stars
26
Forks
2
Last commit
Sep 2026
Advanced
Catalog kind
skill
Gateway key
ab-test-analysis
Source
github.com/jshsakura/awesome-opencode-skills