A/B 实验分析
SkillMonitoring & opsLets your agent analyze existing A/B test results for significance, uplift, and launch recommendations.
Use A/B 实验分析 in Claude, ChatGPT or Ahel Desktop
Free. Sign in, add A/B 实验分析 and connect your AI. About a minute.
Also: Claude Code · Cursor · Codex
Then ask your AI: use the A/B 实验分析 skill
Details
Instructions available. Your AI can read the instructions. Execution depends on the setup they require.
Account requirements not reviewed. Check the skill instructions before use; ahel provides instructions and does not run this skill.
No other account needed.
Add ahel to your AI once: Claude, ChatGPT, Cursor, Claude Code or Codex. Then ask it to use this.
About this skill
Analyzes the group quality, conversion or continuous metric uplift, statistical significance, and confidence intervals of existing A/B experiment data, and outputs cautious launch recommendations. Use when the user asks about A/B testing, experiment results, significance testing, conversion rate com
What this skill tells your AI
The instructions your AI receives, as published by zafer-liu/data-analysis-agent in skills/ab-test-analysis/SKILL.md and read by ahel’s review.
分析已采集的实验数据;不要配置线上分流、修改用户分组或宣称因果关系超出随机实验设计所支持的范围。
必要口径
先确认或从数据中识别以下信息。缺少会改变结论的定义时,只问一个最关键的问题。
- 对照组和实验组的分组字段及取值;默认以对照组为基准。
- 随机化单位(用户、设备、订单等)及其唯一标识字段。
- 主指标与指标类型:二元转化、用户级连续值,或需先按用户聚合的收入/次数。
- 实验起止时间、过滤条件,以及主指标的分母和成功条件。
不得把曝光数、订单行数或事件行数直接当作独立用户样本;先按随机化单位聚合。只在数据明确存在恰好两组且分组定义清晰时输出 A/B 结论;否则说明这是多组或非实验比较。
工作流
- 用
get_schema确认候选表、字段类型和时间覆盖范围;不臆造字段或业务事件。 - 用
query_data检查每个实验单位是否只有一个分组、各组样本量、缺失值,以及实验期间是否一致。发现重复分组、空分组或明显的样本比例失衡时,先报告风险。 - 用
run_analysis执行AB_Test_Analysis。SQL 必须返回每个随机化单位恰好一行、一个数值指标列和一个分组列;不要传原始事件行。 - 在
analysis_options中传入control_group;可选传入metric_type(binary/continuous/auto)及expected_allocation。二元指标必须为 0/1。 - 读取结果中的比例/Welch 检验、95% 置信区间、提升幅度与 SRM。不要自行计算 p 值或置信区间。
- 同时检查并报告样本比例是否与预期分流一致、实验周期是否完整,以及多个指标/分组比较带来的假阳性风险。未预先声明为主指标的结果仅标为探索性。
- 先给出一张组间对比表;数据支持时,用
generate_chart绘制主指标及 95% 置信区间,或按日期的组别趋势图。
结论规则
- 仅在双侧
p < 0.05、95% 置信区间不跨 0、方向与业务目标一致,且样本质量没有重大问题时,建议“实验组在该主指标上有统计证据优于对照组”。 p >= 0.05或置信区间跨 0 时,写“尚无足够证据证明差异”,不要写“两个版本相同”。- 显著但效果很小、区间很宽、样本比例异常、实验过早停止或存在多重比较时,建议继续实验或复核,不直接建议全量上线。
- 不把统计显著性当作业务显著性;将预先约定的最小可接受提升、成本和护栏指标纳入建议。数据中没有这些信息时明确列为未验证项。
输出格式
按以下顺序输出:实验口径与数据范围、样本质量检查、主指标结果表、统计结论、业务建议、局限性与下一步。每个关键数字标注来源表和过滤条件;区分观察结果、统计推断和业务假设。
工具路由
- 使用
get_schema识别实验表、随机化单位、分组、指标与时间字段。 - 使用
query_data完成单位级聚合与数据质量预检;SQL 中显式写出过滤条件与分组字段。 - 使用
run_analysis并设置analysis_name="AB_Test_Analysis"、target_column、groupby_column和analysis_options。 - 只在主指标结果已验证后使用
generate_chart,图表必须标明组别、指标、时间范围与置信区间口径。
Implementation reference
- 查询工具:
agent/tools/business/data.py::_tool_query_data - 图表工具:
agent/tools/business/data.py::_tool_generate_chart - 图表实现:
Function/Charts_generation/chart_generate.py
Signals
- GitHub stars
- 3k
- Forks
- 226
- Last commit
- Sep 2026
Advanced
- Item type
- skill
- Key
ab-test-analysis-zafer-liu- Source
- github.com/zafer-liu/data-analysis-agent
github.com/zafer-liu/data-analysis-agent
More in Monitoring & ops
Skill · anthropics
More in Monitoring & opsagent-eval
Skill · affaan-m
More in Monitoring & opsdashboard-builder
Skill · affaan-m
More in Monitoring & opsbabysit
Skill · thedotmack
More in Monitoring & opseng-runbook
Skill · nexu-io
More in Monitoring & opsweekly-update
Skill · nexu-io
More in Monitoring & ops