CodeGraph Quality Audit
SkillAI & modelsagent-eval is a skill that benchmarks CodeGraph retrieval quality on a real codebase. It compares agent behavior with and without CodeGraph indexing on a selected open-source repository, so a person can test, audit, or validate a CodeGraph version, a local dev build or a published npm version, against a language's repo.
Use CodeGraph Quality Audit in Claude, ChatGPT or Ahel Desktop
Free. Sign in, add CodeGraph Quality Audit and connect your AI. About a minute.
Also: Claude Code · Cursor · Codex
Then ask your AI: use the CodeGraph Quality Audit skill
Details
Instructions available. Your AI can read the instructions. Execution depends on the setup they require.
Account requirements not reviewed. Check the skill instructions before use; ahel provides instructions and does not run this skill.
No other account needed.
Have a CodeGraph version to test, either a local dev build or a published npm version.
What your AI can do with it
- Runs a benchmark script comparing agent performance with and without CodeGraph indexing
- Walks through choosing a CodeGraph version, language, target repo, and test mode
- Launches the audit in the background
- Reports costs and tool-call counts
- Produces comparison tables of results
Getting started
- Have a CodeGraph version to test, either a local dev build or a published npm version.
- Pick a language and an open-source repository to benchmark against.
- Run /agent-eval or ask the agent to test, benchmark, audit, or validate the codegraph version.
- Follow the skill's prompts to choose the CodeGraph version, language, target repo, and test mode.
- Let the audit run in the background and review the reported costs, tool-call counts, and comparison tables.
What this skill tells your AI
The instructions your AI receives, as published by colbymchenry/codegraph in .claude/skills/agent-eval/SKILL.md and read by ahel’s review.
Measures how much CodeGraph helps an agent versus plain grep/read, for a chosen
codegraph version on a chosen real-world repo. Drives the harness in
scripts/agent-eval/.
Prerequisites
tmux3+, a logged-inclaudeCLI,node,git(macOS/Linux).- Run from the codegraph repo root.
Workflow
Copy this checklist:
- [ ] 1. Pick version (local or npm)
- [ ] 2. Pick language
- [ ] 3. Pick repo by size
- [ ] 4. Pick harness (headless / tmux / both)
- [ ] 5. Run audit.sh in the background
- [ ] 6. Report results
Step 1 — version. Ask with AskUserQuestion: which codegraph version to test.
Offer "Local dev build" and "Latest published"; the free-text "Other" lets the
user type a specific version (e.g. 0.7.10). Map the answer to a VERSION token:
- "Local dev build" →
local - "Latest published" →
latest - a typed version → that string (e.g.
0.7.10)
Step 2 — language. Read .claude/skills/agent-eval/corpus.json. Ask with
AskUserQuestion which language to test, listing the languages that have entries.
Step 3 — repo. From the chosen language's entries, ask which repo. Label each
option with its size and file count, e.g. excalidraw — Medium (~600 files).
Each entry carries the repo URL and a representative question.
Step 4 — harness. Ask with AskUserQuestion which harness to run, and map
the answer to a MODE token:
- "Headless" →
headless—claude -pwith stream-json: exact tokens/cost and a clean tool sequence (2 runs, fast, no TTY). - "Interactive (tmux)" →
tmux— drives the real Claude TUI in tmux: faithful Explore-subagent behavior, metrics from session logs (2 runs, slower). - "Both" →
all— headless + interactive (4 runs).
Step 5 — run. Launch in the background (sets the version, clones if missing, wipes + re-indexes, runs the chosen arms — several minutes):
scripts/agent-eval/audit.sh <VERSION> <repo-name> <repo-url> "<question>" <MODE>
Step 6 — report. When the job finishes, read the log and report per arm:
- Headless (
parse-run.mjs): total tool calls, fileReads, Grep/Bash, codegraph-tool calls, duration, total cost. - Interactive (
parse-session.mjs): theVERDICT: codegraph_explore used Nx | Read N | Grep/Bash NandTOKENS:lines. - Both paths also print the three feedback metrics — residual context occupancy,
explore sufficiency, allocation efficiency — and a headless A/B ends with a
side-by-side
ARM COMPARISONtable. Report that table, and check its contamination row first:CLI calls that RETURNED output> 0 means the arm reached codegraph through Bash and its numbers are void. How to read the rest:docs/benchmarks/agent-eval-feedback-metrics.md.
Lead with cost + tool/Read counts — they are the reliable signals; raw token in/out are confounded by subagent delegation and prompt caching. State whether codegraph reduced effort and whether both arms reached a correct answer.
Notes
- The index is rebuilt every run (
audit.shwipes.codegraph) — different versions extract differently, so an index must be served by the same binary that built it. audit.shtemporarily mutates the globalcodegraphinstall for the test, then restores your dev link vialocal-install.sh.- Corpus repos are cloned to
/tmp/codegraph-corpus(reused if already present). - Add or edit repos in
corpus.json(fields:name,repo,size,files,question).
Signals
- GitHub stars
- 73k
- Forks
- 5k
- Last commit
- Oct 2026
- Hacker News mentions
- 2
Questions
- How to evaluate Claude skill?
- agent-eval evaluates retrieval quality by running a benchmark script that compares agent behavior with and without CodeGraph indexing on a real open-source repository, then reports costs, tool-call counts, and comparison tables.
- What is Claude Eval?
- In this context it refers to the agent-eval skill, which audits a CodeGraph version against a language's repo to measure how well CodeGraph retrieves answers compared to basic searching.
- What should be a Claude skill?
- A skill guides an agent through a task. agent-eval is an example: it walks through choosing a CodeGraph version, language, repo, and test mode, then launches and reports the benchmark.
Advanced
- Item type
- skill
- Key
codegraph-agent-eval- Source
- github.com/colbymchenry/codegraph
github.com/colbymchenry/codegraph
More in AI & models
Skill · anthropics
More in AI & modelswayfinder
Skill · mattpocock
More in AI & modelswizard
Skill · mattpocock
More in AI & modelsalgorithmic-art
Skill · anthropics
More in AI & modelscode-review-and-quality
Skill · addyosmani
More in AI & modelsai-first-engineering
Skill · affaan-m
More in AI & models