Investigating metric anomalies
SkillMonitoring & opsInvestigating-metric-anomalies is a skill that guides an AI agent through investigating server and infrastructure metric anomalies in PostHog Metrics.
Use Investigating metric anomalies in Claude, ChatGPT or Ahel Desktop
Free. Sign in, add Investigating metric anomalies and connect your AI. About a minute.
Also: Claude Code · Cursor · Codex
Then ask your AI: use the Investigating metric anomalies skill
Details
Instructions available. Your AI can read the instructions. Execution depends on the setup they require.
Account requirements not reviewed. Check the skill instructions before use; ahel provides instructions and does not run this skill.
No other account needed.
Have access to PostHog Metrics, including the metric-names-list, characterize-metric-anomaly, and query-metrics tools.
What your AI can do with it
- Finds the right metric from a symptom using metric-names-list
- Characterizes an anomaly with characterize-metric-anomaly: severity, onset time, and top
- Tests hypotheses with targeted query-metrics calls, including filters and formulas
- Drills into top movers by pinning label values and grouping by a second key
- Correlates with logs via query-logs and traces via APM span tools at the onset time
- Produces a conclusion with symptom, probable cause, blast radius, and confidence
Getting started
- Have access to PostHog Metrics, including the metric-names-list, characterize-metric-anomaly, and query-metrics tools.
- For cross-signal correlation, also have the query-logs tool and APM span tools available.
- Give the agent a metric symptom, such as a rising or dropping metric, or an alert fire time on an OTel/Prometheus metric.
- If the exact metric name is unknown, the agent starts with metric-names-list using a substring from the symptom.
- Provide the approximate time the anomaly began or the alert fire time so the agent can set the investigation window.
What this skill tells your AI
The instructions your AI receives, as published by posthog/posthog in products/metrics/skills/investigating-metric-anomalies/SKILL.md and read by ahel’s review.
The job: go from a metric symptom ("ingestion lag is rising") to a probable cause with evidence, fast. The metric tells you what and when; logs and traces tell you why. Follow the loop below — it front-loads the cheap, high-information calls and only fans out when the blast radius is unclear.
The loop
1. Pin down the metric
If you have the exact metric name, skip ahead. Otherwise call metric-names-list with a substring from the symptom (lag, error, latency, queue). The returned metric_type decides the lens: counters (sum) are only meaningful as rate/increase, gauges as avg, histograms as histogram_quantile.
2. Characterize first — one call, three answers
Call characterize-metric-anomaly with the metric name and anomalyFrom (the alert fire time, or when the user says it started looking wrong; subtract some margin if unsure). It compares against the preceding window by default and answers:
- How bad:
direction,change_ratio,anomaly_peakvsbaseline_mean. Ifdirectionisflat, your window or metric is wrong — widen the window, or compare against the same window yesterday viabaselineFrom/baselineTo(daily-pattern metrics often look "anomalous" against the immediately-preceding hours). - When:
onset_time— treat this timestamp as the pivot for everything that follows. - Where:
top_movers— label values whose behavior changed. One mover (a single pod, shard, or endpoint) means a localized culprit; everything moving together means a shared cause (an upstream dependency, a deploy, infra).
3. Sharpen with targeted metric queries
Use query-metrics to test the hypotheses the report raises:
- Drill a mover: re-query with
filterspinning the suspicious label value, grouped by a second key, to localize further (pod → container, endpoint → status code). - Normalize: a rising error count means nothing if traffic doubled — use
clauses+formula(errors / requests) to separate rate changes from volume changes. - Check the neighbors: query the obvious companion metrics over the same window (for lag: throughput and error counters of the same service; for latency: request rate and saturation gauges). Use the same
intervalso the grids align visually.
4. Correlate across signals at the onset
Pivot into logs and traces using the same service and a window bracketing onset_time (a few buckets before, through the peak):
- Logs: use
query-logs(follow its own discover-first workflow) filtered to the implicatedservice.nameand window, severityerrorfirst, thenwarn. Restarts, crash loops, connection errors, and deploy markers right before onset are the classic causes. Widen to other services in the request path if the service's own logs are clean. - Traces: use the APM span tools (
query-apm-spansetc.) for the same service/window — slow or erroring spans show which dependency degraded, and atrace_idfrom an exemplar or log line links a concrete request across all three signals.
5. Conclude with evidence, not vibes
State: the symptom (metric, magnitude, onset), the probable cause (what you found in logs/traces and how its timing aligns with the onset), the blast radius (which services/labels are affected, from the movers and grouped queries), and the confidence level. If the cause is still ambiguous, say which hypothesis the evidence favors and what would disambiguate (e.g. "the lag began draining at 20:12 — consistent with a consumer restart; check who restarted it").
Worked example: "ingestion lag is rising — why?"
metric-names-listwithvalue: "lag"→logs_rate_limiter_message_lag_seconds(histogram) and friends.characterize-metric-anomalyon it withanomalyFrom= alert time → directionup, change ratio 40x,onset_time20:10, top moverservice_name = logs-ingestion(the other services' lag stayed flat) — so the logs consumer specifically is behind, not the whole pipeline.query-metrics:rateof the consumer's throughput counter over the same window → throughput was zero during the gap and spiked after onset: the consumer wasn't slow, it was down, and the "rising lag" is it draining the backlog.query-logsforservice.name = logs-ingestion(and its neighbors) around 20:00–20:15 → process exit + restart lines at the gap boundaries.- Conclusion: consumer outage 20:01–20:10 (restart visible in logs); lag spike is backlog drain, self-recovering; affected signal: logs freshness only — no data loss (topic retained messages). Evidence: zero throughput during the window, message-age spike equal to outage duration, restart log lines.
Pitfalls
- Counters reset on process restart —
rate/increasealready handle this; never eyeball raw cumulative counter values. - A metric that stops reporting is not "zero" — a gap in
serieswithtop_moversshowing a vanished label value means the emitter died; pivot to logs immediately. - Don't trust a single aggregation: a flat
avgcan hide a screamingp95. For latency-like gauges and histograms, characterize the tail too. - Scraped Prometheus metrics arrive ~15–60s behind real time; don't read the last bucket of a series as "current".
Signals
- GitHub stars
- 40k
- Forks
- 3k
- Last commit
- Sep 2026
Others that do the same job
Questions
- When should this skill be used?
- When asked why a metric looks wrong (ingestion lag rising, error rate spiking, latency degrading, queue depth growing, throughput dropping), when an alert fires on an OTel/Prometheus metric, or for incident triage that starts from a metric symptom.
- What does characterize-metric-anomaly return?
- One call answers three things: how bad (direction, change_ratio, anomaly_peak vs baseline_mean), when (onset_time), and where (top_movers, label values whose behavior changed).
Advanced
- Item type
- skill
- Key
investigating-metric-anomalies- Source
- github.com/posthog/posthog
More in Monitoring & ops
Skill · anthropics
More in Monitoring & opsagent-eval
Skill · affaan-m
More in Monitoring & opsdashboard-builder
Skill · affaan-m
More in Monitoring & opsbabysit
Skill · thedotmack
More in Monitoring & opseng-runbook
Skill · nexu-io
More in Monitoring & opsweekly-update
Skill · nexu-io
More in Monitoring & ops