Benchmark — Performance Baseline & Regression Detection
SkillDev toolsBenchmark lets your AI measure how your pages, APIs, and builds perform and compare the numbers before and after code changes. That way slowdowns and regressions get caught early instead of after they ship. It can also compare stack alternatives so you can choose the option that performs best.
Available today. Use it from your connected AI after setup.
No other account needed.
After adding it, ask your AI to record a baseline for the pages, APIs, or builds you care about. From then on, it can compare new measurements against that baseline whenever your code changes.
Then ask your AI: use the Benchmark — Performance Baseline & Regression Detection skill
What your AI can do with it
- Measure page, API, and build performance
- Set performance baselines for your project
- Detect regressions by comparing results before and after code changes
- Spot regressions around pull requests
- Compare stack alternatives on performance
What this skill tells your AI
The instructions your AI receives, as published by affaan-m/ecc in skills/benchmark/SKILL.md and read by ahel’s review.
When to Use
- Before and after a PR to measure performance impact
- Setting up performance baselines for a project
- When users report "it feels slow"
- Before a launch — ensure you meet performance targets
- Comparing your stack against alternatives
How It Works
Mode 1: Page Performance
Measures real browser metrics via browser MCP:
1. Navigate to each target URL
2. Measure Core Web Vitals:
- LCP (Largest Contentful Paint) — target < 2.5s
- CLS (Cumulative Layout Shift) — target < 0.1
- INP (Interaction to Next Paint) — target < 200ms
- FCP (First Contentful Paint) — target < 1.8s
- TTFB (Time to First Byte) — target < 800ms
3. Measure resource sizes:
- Total page weight (target < 1MB)
- JS bundle size (target < 200KB gzipped)
- CSS size
- Image weight
- Third-party script weight
4. Count network requests
5. Check for render-blocking resources
Mode 2: API Performance
Benchmarks API endpoints:
1. Hit each endpoint 100 times
2. Measure: p50, p95, p99 latency
3. Track: response size, status codes
4. Test under load: 10 concurrent requests
5. Compare against SLA targets
Mode 3: Build Performance
Measures development feedback loop:
1. Cold build time
2. Hot reload time (HMR)
3. Test suite duration
4. TypeScript check time
5. Lint time
6. Docker build time
Mode 4: Before/After Comparison
Run before and after a change to measure impact:
/benchmark baseline # saves current metrics
# ... make changes ...
/benchmark compare # compares against baseline
Output:
| Metric | Before | After | Delta | Verdict |
|--------|--------|-------|-------|---------|
| LCP | 1.2s | 1.4s | +200ms | WARNING: WARN |
| Bundle | 180KB | 175KB | -5KB | ✓ BETTER |
| Build | 12s | 14s | +2s | WARNING: WARN |
Output
Stores baselines in .ecc/benchmarks/ as JSON. Git-tracked so the team shares baselines.
Integration
- CI: run
/benchmark compareon every PR - Pair with
/canary-watchfor post-deploy monitoring - Pair with
/browser-qafor full pre-ship checklist
Signals
- GitHub stars
- 256k
- Forks
- 38k
- Last commit
- Sep 2026
- Hacker News mentions
- 20
Others that do the same job
Advanced
- Catalog kind
- skill
- Gateway key
benchmark- Source
- github.com/affaan-m/ecc