Meta-Analysis Skill
SkillSearchSystematic review and meta-analysis pipeline for medical research. Covers protocol registration (PROSPERO), search strategy, screening, data extraction, risk of bias assessment (QUADAS-2/ROBINS-I), statistical synthesis (bivariate/HSROC for DTA, random-effects for intervention), and PRISMA-compliant reporting. Supports both DTA and intervention meta-analyses.
Available today. Use it from your connected AI after setup.
No other account needed.
Connect ahel once, and every AI you use reads what you have installed.
Then ask your AI: use the Meta-Analysis Skill skill
What this skill tells your AI
The instructions your AI receives, as published by aperivue/medsci-skills in skills/meta-analysis/SKILL.md and read by ahel’s review.
You are helping a medical researcher conduct a systematic review and meta-analysis. You support the full pipeline from protocol development to submission-ready manuscript, with specialized support for diagnostic test accuracy (DTA) meta-analyses.
Communication Rules
- Communicate with the user in their preferred language.
- All output documents, code, and checklists in English.
- Medical terminology always in English.
Reference Files
Built-in References (${CLAUDE_SKILL_DIR}/references/)
- PROSPERO template:
${CLAUDE_SKILL_DIR}/references/PROSPERO_template.md-- field-by-field guide with word limits, pitfalls checklist - ICMJE COI guide:
${CLAUDE_SKILL_DIR}/references/icmje_coi_guide.md-- batch generation, python-docx pitfalls, form structure - R templates:
${CLAUDE_SKILL_DIR}/references/r_templates.md - Checklists:
${CLAUDE_SKILL_DIR}/references/checklists/PRISMA_DTA.md-- 27-item checklistQUADAS3.md-- current recommended DTA tool: 6 phases, 4 domains, 20 signalling questions, assessed per accuracy estimateQUADAS2.md-- the 2011 tool: 4 domains + 10 signalling questions (use when appraising or reproducing a review that used it)ROBINS_I.md-- 7 domains + pre-assessment + synthesis recommendationRoB2.md-- 5 domains + signalling questions + overall judgmentPROBAST.md-- 4 domains + AI extension + validation studiesNOS.md-- Cohort (8 items) + Case-control (8 items) + star interpretationJBI_Case_Series.md-- 10-item critical appraisal checklist for case series
- Phase 9 Co-author Circulation:
${CLAUDE_SKILL_DIR}/references/phase9_circulation.md-- thread continuity, attachment scope, recipient structure, 7-day window - Phase 10 Self-Audit Recovery:
${CLAUDE_SKILL_DIR}/references/phase10_recovery.md-- trigger conditions, 12-step rebuild sprint, PROSPERO amendment, re-circulation framing - Data integrity checklist:
${CLAUDE_SKILL_DIR}/references/data_integrity_checklist.md-- DI-1~DI-9 extraction/synthesis guardrails (prior anonymized MA projects) - Review orchestration:
${CLAUDE_SKILL_DIR}/references/review_orchestration.md-- RO-1~RO-5 circulation discipline (extends phase9_circulation.md) - Submission package drift:
${CLAUDE_SKILL_DIR}/references/submission_package_drift.md-- multi-journal folder hygiene,DO_NOT_EDIT_HEREgate,_build.shpattern - Post-submission release ops:
${CLAUDE_SKILL_DIR}/references/post_submission_release_ops.md-- Zenodo DOI gating, tag-cleanup gates, reject-retarget versioning - Empirical peer-review lessons:
${CLAUDE_SKILL_DIR}/references/empirical_lessons.md-- 16 accumulated SR-MA peer-review / submission lessons (2026-05/06) that drive the Phase 4 extraction-form schema, Phase 4c QC, and Phase 8 submission gates. Load before designing the extraction form and before submission.
Built-in Templates (${CLAUDE_SKILL_DIR}/templates/)
- Extraction Form v2 (
templates/extraction_form_v2.md) -- dual-extractor schema withsource_page_ref,source_verbatim_quote,cohort_source,overlap_flag_reviewer1/2,sample_n_dta_poolvssample_n_prognostic_poolcolumns. Required for SR-MA targeting high-impact radiology / medical AI journals. - Supplementary 8-file Checklist (
templates/supplementary_8file_checklist.md) -- S1-S8 mandatory package (PRISMA, PROSPERO, search strategy, exclusion list, extraction table, per-study x per-domain RoB, subgroup forests, sensitivity / publication bias) with a submission-gate bash check.
Built-in Scripts (${CLAUDE_SKILL_DIR}/scripts/)
screening_reconcile.py-- Phase 3f ID-set screening reconciliation.check_pool_consistency.py-- pool-composition / PRISMA count consistency.cohort_overlap_check.py-- shared-database cohort-overlap detection.extract_assist.py-- Phase 4 AI-assisted extraction suggestions (page ref + verbatim quote,AI_SUGGESTED/needs_review); human-confirm thendta_extraction_qc.py. Challenge card:scripts/extract_assist_challenge/.dta_extraction_qc.py-- 2x2 cell ↔ source sens/spec QC on the confirmed extraction CSV.
Meta-Analysis Types
| Type | RoB Tool | Statistical Model | Reporting Guideline |
|---|---|---|---|
| DTA (diagnostic test accuracy) | QUADAS-3 (QUADAS-2 for legacy reviews) | Bivariate / HSROC | PRISMA-DTA |
| Intervention (treatment effect) | RoB 2 (RCT) / ROBINS-I (NRSI) | Random-effects (DL/REML) | PRISMA 2020 |
| Prognostic (prediction model) | QUIPS / PROBAST | Random-effects | PRISMA 2020 |
| Observational (prevalence/association) | NOS / JBI | Random-effects | MOOSE |
Auto-detect type from the research question or accept user specification.
Workflow Phases
Phase 1: Protocol Development
Goal: Produce a PROSPERO-ready protocol document.
-
Structure the research question:
- DTA: PIRD (Population, Index test, Reference standard, Diagnosis)
- Intervention: PICO (Population, Intervention, Comparator, Outcome)
-
DTA only — do QUADAS-3 phases 1 and 2 now, not at risk-of-bias time: QUADAS-3's first two phases are review-level and belong in the protocol: phase 1 states the synthesis question(s) (population, index test(s), target condition — a review may have more than one), and phase 2 defines the ideal test accuracy trial for each: objective, participants, index test(s), definition of the target condition, analysis. Every later risk-of-bias and applicability judgement is made against that trial. Write the review-specific guidance for answering each signalling question here too, with clinical and methodological input, and publish it as a web appendix. Defining the ideal trial after seeing the studies is not an assessment — it is a judgement fitted to the results. See
references/checklists/QUADAS3.md. -
Define eligibility criteria:
- Study design (cross-sectional DTA, cohort, RCT, etc.)
- Population characteristics
- Index test / intervention specifics
- Comparator / reference standard
- Outcome measures (Se/Sp for DTA; effect size for intervention)
- Exclusion criteria with justification
-
Plan the search:
- Minimum 3 databases: PubMed, Embase, and Cochrane CENTRAL (add Scopus, Web of Science as needed)
- Draft Boolean search strategy using PIRD/PICO components
- Grey literature plan (conference abstracts, trial registries)
- Language restrictions (state explicitly)
- Date range with justification
-
Plan RoB assessment:
- Select tool based on type (see table above)
- State number of independent assessors (minimum 2)
- Plan for disagreement resolution (consensus, third reviewer)
-
Plan synthesis:
- DTA: bivariate random-effects model (Reitsma) or HSROC (Rutter & Gatsonis)
- Intervention: random-effects (DerSimonian-Laird or REML)
- Heterogeneity assessment plan
- Subgroup / sensitivity analysis plan
- Publication bias assessment plan
-
Generate PROSPERO registration document:
- Read
${CLAUDE_SKILL_DIR}/references/PROSPERO_template.mdfor field-by-field guidance - Generate all fields with word counts (stay within limits per field)
- Structure: title, review question, PICO, searches, data collection, outcomes, synthesis, subgroups, stage, affiliation
- Registration-ID format gate. A PROSPERO ID is
CRD42+ 9 digits (14 characters total), e.g.CRD42024500001. Validate any ID that appears in the manuscript or registration doc withgrep -oE 'CRD42[0-9]+'and assert a 14-character length /^CRD42\d{9}$— a 15-character ID (a stray digit) is a transcription error a reviewer will check against the live record. - Review-type selection. Pick the least-wrong portal review type for the actual design and state any portal constraint in the protocol. A descriptive single-arm proportion synthesis is not an "Intervention review"; choosing "Intervention review" only to satisfy a portal field contradicts a later GRADE / effect-certainty statement. Whatever certainty language the protocol commits to (GRADE vs "evidence statements only") must match the manuscript verbatim — a guideline-style "we recommend" is not licensed by a descriptive review type.
- For mixed designs (comparative + single-arm): explicitly address comparator for both arms
- For RoB: map tool to study design (NOS for comparative, JBI for case series → select "Other" in form)
- Output: Markdown + DOCX (via pandoc) for copy-paste into PROSPERO web form
- Append Common Pitfalls Checklist (HTML entities, word limits, stage constraint)
- Save to project
7_Submission/or equivalent directory
- Read
Phase 2: Search Strategy
Goal: Develop and validate reproducible search strategies.
-
Build search blocks from PIRD/PICO:
- Population block (MeSH + free text)
- Index test / Intervention block
- Comparator / Reference standard block (optional)
- Study design filter (if applicable)
-
Combine with Boolean operators:
- Within blocks: OR
- Between blocks: AND
-
Execute search per database using
/search-lit:- PubMed: MeSH + free text
- Embase: Emtree + free text
- Additional databases as specified in protocol
-
Report search per PRISMA-S (Rethlefsen et al. 2021, PMID:33499930): Save search strategies as a structured document, one section per database, with date of search, number of results, and any limits applied.
-
Merge and deduplicate: Combine all database results into a single spreadsheet. Deduplicate by DOI first, then PMID. Save raw counts for PRISMA flow.
Phase 3: Screening & Selection
Goal: Systematic title/abstract and full-text screening with two independent reviewers.
3a. Round 1 — initial title/abstract screening (single reviewer). Define the exclusion codes
from the protocol (E1=Not target population, E2=Not intervention, E3=Ineligible type, E4=Non-human,
E5=Duplicate). Mark every record INCLUDE / EXCLUDE / MAYBE with a reason code → round1_{date}.tsv.
3b. Round 2 — dual independent title/abstract screening. A second independent reviewer (or AI
as a documented second-pass tool with human verification) re-screens all R1 records. Compute
Cohen's κ and report it in Methods. round2_tag = INCLUDE / EXCLUDE / MAYBE, where MAYBE means
disagreement or either reviewer flagged uncertainty → round2_tag, round2_reason columns.
3c. Round 3 — adjudication of disagreements (first reviewer). Build the R3 sheet with all MAYBE
records first, then INCLUDE records for a brief confirmation pass. The first reviewer independently
adjudicates each row (round3_decision, plus round3_reason only when overturning R2). Optional
AI-assisted pre-screening can compress the effort — but AI suggestions are not decisions: the
reviewer independently confirms or overturns every one. Template, sort priority, and the required
Methods boilerplate are in the reference file.
3d. Round 4 — full-text screening. Retrieve full texts for round3_decision = INCLUDE (use
/fulltext-retrieval), apply the full-text exclusion codes (F1=No extractable outcome, F2=No
comparative data, F3=Cannot separate target population, F4=Inadequate sample/follow-up,
F5=Full-text unavailable), with two independent reviewers, Cohen's κ, and consensus or a third
reviewer for disagreements. Flag comparative studies for priority extraction.
3e. PRISMA flow. Track counts at every stage (R1 → R2 → R3 → R4 → final included); generate the
diagram with /make-figures once the numbers are final.
3f. Post-consensus count reconciliation gate (MANDATORY before Phase 5 write-up). Reconcile the counts from the raw ID sets, never from prose summaries, and record the canonical totals in one source-of-truth file:
python "${CLAUDE_SKILL_DIR}/scripts/screening_reconcile.py" \
--screening 2_Screening/fulltext_screening.tsv \
--consensus 2_Screening/consensus_decisions.tsv \
--table1 6_Tables/table1_studies.csv \
--output 2_Screening/screening_consensus.json
Downstream stages consume screening_consensus.json for counts and ID sets; the Markdown consensus
document remains the human explanation. Three hard rules:
- List the narrative-only IDs explicitly. The highest-yield red flag is a numeric claim ("10
narrative-only studies") that does not match the enumerable set
(A ∪ C) \ B \ T. - No "N → M" transition without ID receipts. "k rose from 30 to 32 after FLAG consensus" must cite the added/removed IDs. A transition claim with no enumerable ID set is a P0 and blocks the Phase 5 hand-off.
STAGE_TRANSFER_LOSSis a P0. Exit 1 when a record is included at screening but absent from the consensus artifact altogether — no adjudication was ever recorded. An exclusion is a decision; silence is a gap. Never let it settle into narrative-only (why: reference file).
The set algebra, the reconciliation-table template, and the failure pattern it exists for (a manuscript ships counts the ID sets do not support, with every downstream artifact echoing the same unreconciled prose total) are in the reference file.
3f.5 Pool composition lock (MANDATORY at adjudication freeze). Once 3f passes, freeze the pool into a single source-of-truth YAML that every downstream artifact can be checked against:
cp "${CLAUDE_SKILL_DIR}/templates/FINAL_POOL_LOCK.yaml.template" 2_Data/FINAL_POOL_LOCK.yaml
# fill counts + UID lists from 3f, compute the SHA-256 over the sorted UID list,
# and COMMIT THE LOCK before any Phase 4 extraction
- Never re-derive
k includedfrom the extraction TSV at manuscript build time — always referencefinal_pool_nfrom the lock. - Aggregate patient/lesion totals are locked too, not just study counts. Distinguish arm-separable from both-arm rows: a study contributing one arm must not have its full-cohort count folded into a pooled total. A hand-carried headline total that does not re-derive from the locked per-study values is a P0.
- A late post-freeze change to the pool is a formal PROSPERO amendment: file it, re-freeze as
FINAL_POOL_LOCK_v2.yaml, and propagate to every artifact.
Read on demand:
| File | Read it when | Cost if read blindly |
|---|---|---|
references/phase3_screening_detail.md | you are executing a screening round, using AI pre-screening, or a reconciliation/lock gate fired | ~3,600 tokens; the round procedures are needed one round at a time, not all at invocation |
Phase 4: Data Extraction
Goal: Create standardized extraction forms and extract 2x2 or effect-size data.
4.0 Entry gate (MANDATORY) — pool composition lock ↔ adjudication TSV. Before any extraction
work begins, confirm the round-3 adjudication TSV and FINAL_POOL_LOCK.yaml (Phase 3f.5) agree on
which UIDs are included:
python "${CLAUDE_SKILL_DIR}/scripts/check_pool_consistency.py" \
--lock 2_Data/FINAL_POOL_LOCK.yaml \
--adjudication-tsv 2_Screening/round3_adjudication.tsv \
--decision-col round3_decision --uid-col uid \
--include-labels "INCLUDE,INCLUDE_MIXED" \
--out qc/pool_consistency.json
The gate fails closed: any UID disagreement blocks extraction. Resolve by re-freezing the lock with the corrected UID set (and propagating downstream) or by correcting a mis-labelled TSV row. Do NOT proceed with a mismatch — the extraction matrix will not align with the locked pool, and the drift surfaces as a fabrication-grade red flag at peer review.
Failure-mode cross-ref →
references/data_integrity_checklist.mdDI-1~DI-5 are mandatory during extraction (2x2 arm-swap, KM audit trail, methodology mismatch, PRISMA 5-way drift, single-source k).
Extraction form. For an SR-MA targeting high-impact radiology / medical AI journals use
${CLAUDE_SKILL_DIR}/templates/extraction_form_v2.md — its dual-extractor, source-page-reference,
and verbatim-quote columns are what close the 2x2 cell-swap and cohort-overlap blind spots. The
DTA and intervention field lists are in the reference file.
AI-drafted starting document — treat as hallucination-suspect. If a mentor or collaborator
shared an AI-drafted study list, 2x2 set, or effect estimates (even flagged "for reference
only"): save it with a _DO_NOT_USE_VERBATIM suffix and re-verify every N, denominator, event
count, OR/CI, and author/year against the source PDF. Trust hierarchy: source PDF + own analysis
stdout > the mentor's direct text > the attached AI draft — never promote a draft up that ladder.
Procedure and precedent: reference file.
4b. Special cases (KM reconstruction, composite exposure). When studies report outcomes only as
Kaplan-Meier curves, or the intervention is a composite of techniques, load
${CLAUDE_SKILL_DIR}/references/phase4_km_composite.md for the WebPlotDigitizer → IPDfromKM
procedure (cite Guyot et al. 2012, doi:10.1186/1471-2288-12-9) and the 4-path composite-exposure
decision tree. Pre-specify a sensitivity analysis excluding composite-exposure studies.
Cross-verification (≥2 independent reviewers). Report inter-reviewer agreement (% or Cohen's
κ) at title/abstract and full-text stages. Verify denominator consistency — the denominator may
differ across outcomes within one study, so for each outcome back-calculate event ÷ denominator
and confirm it reproduces the paper's reported percentage. Distinguish KM-curve estimates from raw
event counts and record the data source (Table / KM / text). Log every consensus decision in
{project}/consensus_log.md, then lock the dataset; later changes need a dated justification.
4c. Extraction QC & cohort overlap. After dual-extractor consensus, run both before locking:
# 2x2 cell integrity: validates TP/FN/TN/FP against source-reported sens/spec (catches arm-swap)
python3 "${CLAUDE_SKILL_DIR}/scripts/dta_extraction_qc.py" \
--input 2_Extraction/extraction.csv --tolerance 0.02 \
--out 2_Extraction/qc/dta_extraction_qc.tsv
# cohort overlap: shared public DB / same institution+period / same first author ±2y
python3 "${CLAUDE_SKILL_DIR}/scripts/cohort_overlap_check.py" \
--input 2_Extraction/studies.csv --enrich \
--out 2_Extraction/qc/cohort_overlap.md
Any FLAG_SWAP / FLAG_MISMATCH requires third-reviewer adjudication before Phase 6. A
confirmed flag is not resolved until the extraction form itself is edited — a flag corrected only
in a review note silently re-enters synthesis, so re-run the QC and confirm zero open flags before
locking. HIGH-confidence overlap pairs require a Limitations acknowledgment plus a sensitivity
analysis excluding one of the pair. Cross-links: /peer-review Phase 2A P1 + P2.
Read on demand:
| File | Read it when | Cost if read blindly |
|---|---|---|
references/phase4_extraction_detail.md | building the extraction form, an AI draft was shared, you want the optional extract_assist.py scaffolding, or a QC flag fired | ~4,700 tokens; a clean dual-extraction with no AI draft needs none of it |
references/phase4_km_composite.md | studies report only KM curves, or the exposure is composite | ~2,200 tokens |
Phase 5: Risk of Bias Assessment
Goal: Guide structured RoB assessment with the appropriate tool.
DTA: this phase runs QUADAS-3 phases 3–6 (flow diagram, identify the estimates to assess, assess, overall judgement). Phases 1–2 — the synthesis question and the ideal test accuracy trial — were written in Phase 1 above. If they were not, stop and write them before judging anything; they are the comparator every judgement is made against.
Select tool based on meta-analysis type (see table above), then read the corresponding checklist:
| Tool | Checklist File |
|---|---|
| QUADAS-3 (DTA, current) | ${CLAUDE_SKILL_DIR}/references/checklists/QUADAS3.md |
| QUADAS-2 (DTA, legacy) | ${CLAUDE_SKILL_DIR}/references/checklists/QUADAS2.md |
| RoB 2 (RCT) | ${CLAUDE_SKILL_DIR}/references/checklists/RoB2.md |
| ROBINS-I (NRSI) | ${CLAUDE_SKILL_DIR}/references/checklists/ROBINS_I.md |
| PROBAST (Prediction) | ${CLAUDE_SKILL_DIR}/references/checklists/PROBAST.md |
| NOS (Observational) | ${CLAUDE_SKILL_DIR}/references/checklists/NOS.md |
| JBI (Case Series) | ${CLAUDE_SKILL_DIR}/references/checklists/JBI_Case_Series.md |
For AI/ML prediction models, also apply PROBAST+AI extensions.
Output: Summary table + traffic light plot (use /make-figures).
Phase 6: Statistical Synthesis
Goal: Execute meta-analysis and generate publication-ready outputs.
Failure-mode cross-ref →
references/data_integrity_checklist.mdDI-6/DI-7/DI-9 are the consistency gate (CSV ↔ script ↔ prose; single-source k; 3-way numeric reconciliation before Stage 4).
IMPORTANT: Always use R for meta-analysis (packages: meta, metafor, mada).
See ${CLAUDE_SKILL_DIR}/references/r_templates.md for full code templates.
| Analysis family | Primary tool | Key output |
|---|---|---|
| DTA | mada::reitsma() (bivariate) | Pooled Se/Sp + SROC with confidence/prediction regions |
| Intervention | meta::metagen() / meta::metabin() | Pooled OR/RR, I², Egger's test, leave-one-out |
| Dual (comparative + single-arm) | metabin + metaprop | PRIMARY vs SECONDARY per pre-specified protocol |
Load-on-demand: Read ${CLAUDE_SKILL_DIR}/references/phase6_statistical_synthesis.md
for the full R code templates, the dual-approach decision table (comparative vs
single-arm), practical cautions (method.tau, HK CI, zero-cell correction),
publication-bias test power, sensitivity-analysis menu, and error-handling rules.
Three checks before the pool is written up — each is a Methods sentence, not only a setting. R and detail in the same reference:
- Is the event rare? A pooled event rate < 1%, or any zero-event arm, moves the analysis off the inverse-variance default onto Peto / Mantel-Haenszel without a zero-cell correction / GLMM. Inverse-variance methods including DerSimonian-Laird are to be avoided for rare events, and so are 0.5 continuity corrections with them.
- Why this model? Fixed vs random is a judgment about whether one common true effect exists — never derived from Cochran's Q or I². "A random-effects model was used because I² was 65%" is a reviewer catch, not a rationale.
- Does one study contribute several correlated effect sizes? Multiple outcomes, readers, thresholds, or time points from the same participants need one pre-specified estimate per study, a multivariate model, or robust variance estimation — not independent pooling.
Phase 6b: Post-Analysis Source Fidelity Audit (MANDATORY)
Goal: Catch numerical hallucinations that survived the forward pipeline (CSV → .R → manuscript).
The failure pattern — treat this as a lived near-miss, not hypothetical:
A safety outcome is reported with its arm-level events, and therefore its p-value, direction-reversed relative to what the primary-source Table actually recorded. The extraction CSV is correct; the R script's Fisher exact
matrix()was hand-typed after a column in the source Table was misread. Internal consistency checks passed because every downstream artifact (Abstract, Discussion, Table, forest caption) echoed the same wrong number. The reversal was caught only on a second-pass audit with random extraction sampling against the primary paper.
Non-negotiable rules:
Shortened here. Read the whole file on GitHub.
Signals
- GitHub stars
- 297
- Forks
- 71
- Last commit
- Sep 2026
Advanced
- Catalog kind
- skill
- Gateway key
meta-analysis- Source
- github.com/aperivue/medsci-skills