Test-Driven Development
SkillDev toolsThe tdd skill guides an AI agent through test-driven development. It enforces the red-green-refactor loop: agree on test seams, write one failing test, then the minimal code to pass it, repeating in vertical slices. It also defines what a good test is, where tests belong, and common anti-patterns.
Use Test-Driven Development in Claude, ChatGPT or Ahel Desktop
Free. Sign in, add Test-Driven Development and connect your AI. About a minute.
Also: Claude Code · Cursor · Codex
Then ask your AI: use the Test-Driven Development skill
Details
Instructions available. Your AI can read the instructions. Execution depends on the setup they require.
Account requirements not reviewed. Check the skill instructions before use; ahel provides instructions and does not run this skill.
No other account needed.
Have a project with a codebase and a way to run tests.
What your AI can do with it
- Guide the agent through the red-green-refactor loop
- Agree on test seams before writing any test
- Write one failing test, then minimal code to pass it
- Define what a good test is and where tests belong
- Identify anti-patterns like implementation-coupled tests
- Keep refactoring out of the loop for a separate review stage
Getting started
- Have a project with a codebase and a way to run tests.
- Add the tdd skill to your agent's available skills.
- When starting a feature or bug fix, tell the agent to work test-first.
- Confirm the test seams the agent proposes before any test is written.
- Let the agent cycle through failing test, minimal code, and repeat.
What this skill tells your AI
The instructions your AI receives, as published by mattpocock/skills in skills/engineering/tdd/SKILL.md and read by ahel’s review.
TDD is the red → green loop. This skill is the reference that makes that loop produce tests worth keeping: what a good test is, where tests go, the anti-patterns, and the rules of the loop. Every section applies on every cycle: consult them before and during the loop, not after.
When exploring the codebase, read GLOSSARY.md (if it exists) so test names and interface vocabulary match the project's domain language, and respect ADRs in the area you're touching.
What a good test is
Tests verify behavior through public interfaces, not implementation details. Code can change entirely; tests shouldn't. A good test reads like a specification: "user can checkout with valid cart" tells you exactly what capability exists, and it survives refactors because it doesn't care about internal structure.
See tests.md for examples and mocking.md for mocking guidelines.
Seams: where tests go
A seam is the public boundary you test at: the interface where you observe behavior without reaching inside. Tests live at seams, never against internals.
Test only at pre-agreed seams. Before writing any test, write down the seams under test and confirm them with the user. No test is written at an unconfirmed seam. You can't test everything, so agreeing the seams up front is how testing effort lands on the critical paths and complex logic instead of every edge case.
Ask: "What's the public interface, and which seams should we test?"
When the shape of that interface is itself in question (how deep the module is, where the seam belongs, what the interface should expose), call the Skill tool with "codebase-design" for the vocabulary. It is the shared source of the module, interface, depth, seam, adapter, leverage and locality terms, and it is a reference to consult, not a session to run.
Anti-patterns
- Implementation-coupled: mocks internal collaborators, tests private methods, or verifies through a side channel (querying the database instead of using the interface). The tell: the test breaks when you refactor but behavior hasn't changed.
- Tautological: the assertion recomputes the expected value the way the code does (
expect(add(a, b)).toBe(a + b), a snapshot derived by hand the same way, a constant asserted equal to itself), so it passes by construction and can never disagree with the code. Expected values must come from an independent source of truth: a known-good literal, a worked example, the spec. - Horizontal slicing: writing all tests first, then all implementation. Bulk tests verify imagined behavior: you test the shape of things rather than user-facing behavior, the tests go insensitive to real changes, and you commit to test structure before understanding the implementation. Work in vertical slices instead: one test → one implementation → repeat, each test a tracer bullet that responds to what the last cycle taught you.
Rules of the loop
- Red before green. Write the failing test first, then only enough code to pass it. Don't anticipate future tests or add speculative features.
- One slice at a time. One seam, one test, one minimal implementation per cycle.
- Refactoring is not part of the loop. It belongs to the review stage (see the
code-reviewskill), not the red → green implementation cycle.
Signals
- GitHub stars
- 273k
- Forks
- 23k
- Last commit
- Sep 2026
- Installs
- 1.0M installs
Questions
- What is TDD with an example?
- TDD is the red-green loop: write one failing test, then the minimal code to pass it, repeating in vertical slices. For example, write a test that a user can checkout with a valid cart, watch it fail, then implement just enough to pass.
- What is a test seam?
- A seam is the public boundary you test at, where you observe behavior without reaching inside. Tests live at seams, never against internals. Agree on seams before writing any test.
- What are common anti-patterns in TDD?
- Implementation-coupled tests that mock internals or verify through side channels, tautological tests whose assertions recompute the expected value the way the code does, and horizontal slicing where all tests are written before any implementation.
Advanced
- Item type
- skill
- Key
tdd-mattpocock- Source
- github.com/mattpocock/skills