Improving MCP tools
SkillAI & modelsimproving-mcp-tools is a skill that runs an improve-my-MCP campaign: a measure-fix-revalidate loop for MCP tool quality.
Use Improving MCP tools in Claude, ChatGPT or Ahel Desktop
Free. Sign in, add Improving MCP tools and connect your AI. About a minute.
Also: Claude Code · Cursor · Codex
Then ask your AI: use the Improving MCP tools skill
Details
Instructions available. Your AI can read the instructions. Execution depends on the setup they require.
Account requirements not reviewed. Check the skill instructions before use; ahel provides instructions and does not run this skill.
No other account needed.
Have the eval harness at services/mcp/evals/ and a fixed benchmark task set in benchmark/tasks.yaml, since scores are only comparable across runs of the same benchmark version.
What your AI can do with it
- Runs an autoresearch-style loop that measures the MCP agent experience with the eval
- Pulls production failure data via query-mcp-tool-stats, query-mcp-tool-failures
- Picks the highest-impact tool problem by ranking issues on reach and severity
- Makes one bounded fix per iteration, such as rewriting a tool description or tightening
- Turns a change into a draft pull request only when before-and-after scores prove
- Keeps a benchmark task suite and journal so iterations stay comparable and resumable
Getting started
- Have the eval harness at services/mcp/evals/ and a fixed benchmark task set in benchmark/tasks.yaml, since scores are only comparable across runs of the same benchmark version.
- Set up a seeded local or devbox stack to run the harness against, never a customer project, and provide a personal API key as LIVE_MCP_TOKEN.
- Run a baseline measurement with the harness, then pull production evidence using the MCP analytics tools and the signals scout cookbook queries.
- Pick one issue per iteration, skipping anything the journal already shows with two failed attempts, and make a bounded fix limited to the allowlisted files.
- Re-run the affected benchmark slice plus a no-regression sample, and keep the change only if the target metric improves and nothing else degrades.
What this skill tells your AI
The instructions your AI receives, as published by posthog/posthog in .agents/skills/improving-mcp-tools/SKILL.md and read by ahel’s review.
An MCP server gets better only in ways you can measure. This skill is the campaign procedure: score the current agent experience, fix the biggest problem, re-score, and only ship changes the numbers justify. It is the operating manual for the "improve my MCP" loop — one iteration per pass, journaled so a later iteration (or a different agent) can resume without repeating work.
The objective function
services/mcp/evals/ is the harness. benchmark/tasks.yaml is a fixed set of
agent tasks with expected_tools and success_criteria; scores are only
comparable across runs of the same benchmark version.
- Probe mode (deterministic, no LLM):
LIVE_MCP_URL=... LIVE_MCP_TOKEN=... pnpm exec tsx evals/runner/probe.ts --out score.jsonfromservices/mcp/. Reports tool-presence misses (discoverability), probe failures, and latency p50/p95. Non-zero exit = regression. - Agent mode (LLM replay + judge): scores task success and tool-selection accuracy. Use it for description/discoverability changes — probes cannot detect that an agent picks the wrong tool.
Run the harness against a seeded local or devbox stack, never against a
customer project. Local recipe: NODE_ENV=development PORT=9876 POSTHOG_API_BASE_URL=http://localhost:8000 pnpm dev:hono, personal API key as
LIVE_MCP_TOKEN.
One iteration
- Measure. Run the harness for a baseline. Pull production evidence with
the MCP analytics tools (
query-mcp-tool-stats,query-mcp-tool-failures,query-mcp-tool-descriptions,query-mcp-tool-sample-intents) and the lenses in the signals scout cookbook (products/signals/skills/signals-scout-mcp-tool-calls/references/queries.md): failure leaderboard, retry/struggle, latency, intents that matched no tool. - Pick one issue. Rank by reach × severity. Skip anything the journal shows with two failed attempts. One issue per iteration — a PR that fixes three things can't be attributed to any of them when scores move.
- Fix, bounded. Only files inside the allowlist (below). Typical fixes: sharpen a tool description so the right intent finds it, tighten an input schema that agents keep getting wrong, fix an annotation, update a skill.
- Validate. Re-run the affected benchmark slice plus a no-regression sample. Keep the change only if the target metric improves and nothing else degrades. A discarded change is a normal outcome — journal it and move on.
- Ship. One PR per iteration with before/after scores in the body (format
in references/campaign-journal.md). Keep
it stampable: ≤400 changed lines, only files inside the allowlist below,
request a stamphog review (MCP first, label fallback, see
/merging-prs). Autonomy level comes from the campaign config — default is draft PR for human review; only arm auto-merge when the operator has explicitly enabled the self-driving experiment (see guardrails). - Journal. Append the iteration record before ending the pass.
Hard guardrails
These are not suggestions; violating any of them ends the campaign pass.
- Allowlist — a campaign PR may only touch:
products/*/mcp/tools.yaml,products/*/skills/**,services/mcp/evals/**, the codegen outputs ofpnpm generate-tools/scaffold-yaml(services/mcp/src/tools/generated/**andservices/mcp/schema/generated-tool-definitions.json), and docs. Anything else (handler code, package manifests, workflows, migrations, auth paths) → stop and hand the finding to a human as a draft PR or report instead. - Read-only against data. The harness and all production queries are read-only. Never create, mutate, or delete customer-visible objects while measuring.
- Evidence or it didn't happen. No PR without a baseline score, an after score, and the exact harness commands used.
- Benchmark integrity. Never edit
benchmark/tasks.yamlin the same PR as a fix it validates — changing the exam and the answer together proves nothing. Benchmark changes are their own PR and bumpversion. - Budgets. Respect the operator's iteration/token/PR caps (default: stop after 3 open unmerged campaign PRs). Two failed attempts on an issue parks it permanently.
- Kill switch. If the campaign config, its feature flag, or the operator says stop — stop mid-iteration, journal state, end cleanly.
Failure modes to expect
- A description change that helps one intent can steal traffic from the right
tool for another — that's why the no-regression sample is mandatory. The
intent-cluster snapshot's
tool_overlaps(seeexploring-mcp-intent-clusters) lists exactly which pairs compete for which intents: snapshot it before a description rewrite and recompute after, and treat a capture shift in an overlapping pair as the regression signal. - Probe latency varies with stack warmth; compare medians across ≥3 runs before attributing a latency change to your fix.
- Tool-presence misses can be feature-flag gating, not catalog absence —
check
getToolsForFeaturesgating before "fixing" discoverability.
Signals
- GitHub stars
- 40k
- Forks
- 3k
- Last commit
- Sep 2026
Others that do the same job
Questions
- When should this skill be used?
- When asked to "improve my MCP", run an MCP improvement campaign, fix tool discoverability or descriptions based on evidence, or prepare an eval-backed PR for a tool change.
- How many fixes does it make per iteration?
- One. A PR that fixes three things cannot be attributed to any of them when scores move, so each iteration addresses a single issue.
Advanced
- Item type
- skill
- Key
improving-mcp-tools- Source
- github.com/posthog/posthog
Related picks
Skill · mattpocock
The pick for TypeScripttypescript-pro
Skill · jeffallan
The pick for TypeScriptnodejs-backend-patterns
Skill · wshobson
The pick for Noderun-node-tests
Skill · hiroro-work
The pick for Nodeskill-creator
Skill · anthropics
More in AI & modelswayfinder
Skill · mattpocock
More in AI & models