Investigating logs
SkillCloud & infraThis is a skill that lets an AI agent investigate logs in a PostHog project. It guides the agent to summarize before reading, using pattern mining and before/after diffs to check service health, explain error spikes, and triage incidents. The agent starts from summaries instead of raw rows or hand-written SQL over the logs table.
Use Investigating logs in Claude, ChatGPT or Ahel Desktop
Free. Sign in, add Investigating logs and connect your AI. About a minute.
Also: Claude Code · Cursor · Codex
Then ask your AI: use the Investigating logs skill
Details
Instructions available. Your AI can read the instructions. Execution depends on the setup they require.
Account requirements not reviewed. Check the skill instructions before use; ahel provides instructions and does not run this skill.
No other account needed.
Have a PostHog project with logs available and the logs MCP tools configured for your agent.
What your AI can do with it
- Check whether a service or deployment is healthy using PostHog logs.
- Explain an error spike by diffing log patterns between two windows.
- Triage an incident by finding what changed in the log stream.
- Orient in an unfamiliar log stream by mining message templates.
- Localize log volume with bucketed counts before pulling raw rows.
Getting started
- Have a PostHog project with logs available and the logs MCP tools configured for your agent.
- Ask your agent to check the logs for a service, deployment, or time window.
- Let the agent run pattern mining and pattern diffing before it reads raw log rows.
- Review the summary of new, rate-shifted, or gone patterns to decide the next step.
What this skill tells your AI
The instructions your AI receives, as published by posthog/posthog in products/logs/skills/investigating-logs/SKILL.md and read by ahel’s review.
Investigation is a narrowing problem: summarize before you read.
One posthog:logs-patterns call compresses millions of lines into at most 200 templates,
and one posthog:logs-patterns-diff call answers "what is different about now vs. before" directly.
Raw rows (posthog:query-logs) are the last step of an investigation, never the first.
When to use this skill
- "Check the logs" / "is service X healthy?" / "did my deploy (or model bump, config change, migration) break anything?"
- "Why are errors up?" / "explain this spike" / incident triage — "what changed?"
- "What is this service logging?" — orienting in an unfamiliar or noisy stream.
- Finding the log evidence for a failure reported elsewhere (an alert, an error-tracking issue, a user complaint).
When not to use this skill
- Creating or tuning log alerts — that's
authoring-log-alerts. - Analytics over product events, persons, or insights — that's
querying-posthog-data. - HogQL exposes a
logstable viaposthog:execute-sql, but do not investigate through it: hand-written SQL over logs routinely hits read-byte caps and re-derives what the tools below do in one cheap call. Reserve SQL for the rare case of joining log-derived facts with non-log data.
Tools
| Tool | Job |
|---|---|
posthog:logs-services-create | Top-25 services with log_count, error_count, error_rate, sparkline. Orientation. |
posthog:logs-patterns | Mine one window's message templates, ordered by frequency. "What is this stream saying?" |
posthog:logs-patterns-diff | Diff templates between two windows: new / rate-shifted / gone. "What changed?" |
posthog:logs-count / posthog:logs-count-ranges | Scalar and time-bucketed counts for a filter. Localize volume before pulling rows. |
posthog:logs-sparkline-query | Volume over time broken down by severity or service (the one bucketed view with a breakdown). |
posthog:logs-facet-values-create | Distribution of severity/service (or a resource attribute) under a filter. |
posthog:logs-attributes-list / posthog:logs-attribute-values-list | Discover attribute keys and values before building filters. |
posthog:query-logs | Raw rows. Endpoint of every drill-down, entry point of none. |
Each tool's own description documents its parameters and response shape — read it before calling.
Pick the workflow by question shape
"Is it healthy?" — post-deploy / post-change verification
The user changed something (deploy, model bump, config, migration) and wants to know the logs still look right.
- Pin down the change time and the affected service(s). Ask if the user hasn't said; the diff is meaningless without a boundary.
- Orient with
posthog:logs-services-create: is the service still logging at all, and what is its error_rate now? A service that went silent fails verification just as hard as one that started erroring. posthog:logs-patterns-diffwithquery.dateRangefrom the change time to now andbaselineDateRangeset to a comparable window just before the change, scoped toserviceNames. New error/fatal templates right after a change are the classic regression signature; largerate_ratioshifts on existing error templates are the second thing to check.- Check volume continuity with
posthog:logs-count-rangesspanning before and after the boundary: a rate discontinuity (crash loop, restart storm, silence) shows up here even when message content looks unchanged. - Drill only the suspects: pivot each suspicious pattern to raw lines via its
match_regexwithposthog:query-logs.
A pass verdict needs all three: no new error templates, no large error rate_ratio shifts, and continuous volume. Say which windows you compared — "healthy" is only as strong as the baseline.
"Explain this spike"
- Localize it:
posthog:logs-count-rangesover the user's window, then recurse into the dense bucket(s) — each bucket'sdate_from/date_tofeeds the next call. Stop after 3–4 levels. - Explain it:
posthog:logs-patterns-diffwith the spike asquery.dateRangeand the window just before asbaselineDateRange. The topnewandrate_shiftentries are the explanation. Do not mine both windows separately and diff by hand — the diff is one call.
Incident triage — "what broke?"
posthog:logs-patterns-diff first: incident window vs. a known-good window just before (or omit the baseline for
same-window-last-week). Suspects are new entries and the biggest rate_ratio shifts; pivot each to raw lines.
If the failing service is unknown, find it first with posthog:logs-facet-values-create faceting service_name
under severityLevels: ["error", "fatal"].
"What is this stream saying?" — unfamiliar service
posthog:logs-patterns over the last hour, scoped to the service. Scan templates by estimated_count and
non-zero error share in severity_counts. Widen the window or add searchTerm only if the answer isn't there.
Known needle — a specific message, attribute, or person
When the target is already precise (an error string, a request id, a distinct_id), skip pattern mining:
discover the right keys with posthog:logs-attributes-list / posthog:logs-attribute-values-list,
size the result with posthog:logs-count, then pull rows with posthog:query-logs.
Rules that keep investigations honest and cheap
- Scope
serviceNames(or a resource-attribute filter) on every call once the target service is known. Unscoped calls scan the whole team's stream and starve the pattern sample budget. posthog:query-logsrequires an explicitquery.dateRange— omitting it is a 400, not a default window.- Pattern counts are sampled estimates (
sampled: true); templates rarer than ~1 in 10,000 rows can be invisible. Absence of a rare template is not evidence it stopped. - Before trusting a wall of
newentries in a diff, checkbaseline.total_count— a tiny or empty baseline (logging only just started) makes everything look new. severityLevelsmatches the six canonical lowercase buckets againstseverity_textexactly. Zero rows on a severity filter → check the stored values withposthog:logs-attribute-values-list { key: "severity_text" }.- Budget: one services call, at most one patterns-diff per window pair, 3–4 count-ranges levels,
and
query-logsonly for confirmed suspects withlimit≤ 100.
Output
Lead with the verdict, then the evidence:
- Verdict: healthy / regressed / inconclusive, with the windows compared.
- Suspects (if any): template, classification (
new/rate_shift), estimated counts orrate_ratio, services, and 1–2 sample raw lines. - What was checked and what wasn't: services covered, windows, and any sampling or baseline caveats that limit confidence.
The user should be able to act on the verdict without re-running the investigation.
Related skills
authoring-log-alerts— turn the check behind a verdict into a continuous, low-noise alertexploring-apm-traces— follow a suspicious log line into the request traces around it
Signals
- GitHub stars
- 40k
- Forks
- 3k
- Last commit
- Sep 2026
Others that do the same job
Questions
- What is this skill for?
- It guides an agent through log investigations in PostHog: verifying a service or deployment is healthy, explaining an error spike, triaging an incident, or understanding what a log stream is saying.
- When should I not use this skill?
- Do not use it for creating or tuning log alerts, for analytics over product events, persons, or insights, or for investigating through hand-written SQL over the logs table.
- Does it pull raw log rows first?
- No. It summarizes before reading: pattern mining and pattern diffing come first, and raw rows are the last step of an investigation.
Advanced
- Item type
- skill
- Key
investigating-logs- Source
- github.com/posthog/posthog
More in Cloud & infra
Skill · vercel-labs
More in Cloud & infravercel-react-best-practices
Skill · vercel-labs
More in Cloud & infraturborepo
Skill · vercel
More in Cloud & inframicrosoft-foundry
Skill · microsoft
More in Cloud & infraazure-diagnostics
Skill · microsoft
More in Cloud & infrauncloud
Skill · affaan-m
More in Cloud & infra