skill-comply: Automated Compliance Measurement

SkillAI & models

skill-comply is a skill that checks whether agents actually follow the skills, rules, and agent definitions they are given. It turns a markdown target into an expected behavior sequence, generates test scenarios at three prompt strictness levels, runs the agent while recording its tool calls, and classifies those calls against the spec. It then reports compliance rates with full tool call timelines.

Use skill-comply: Automated Compliance Measurement in Claude, ChatGPT or Ahel Desktop

Free. Sign in, add skill-comply: Automated Compliance Measurement and connect your AI. About a minute.

Also: Claude Code · Cursor · Codex

Then ask your AI: use the skill-comply: Automated Compliance Measurement skill

Details

Instructions available. Your AI can read the instructions. Execution depends on the setup they require.

Have a markdown file that defines the skill, rule, or agent behavior you want to test.

skill-comply: Automated Compliance MeasurementStart free

What your AI can do with it

  • Turn a markdown skill, rule, or agent definition into an expected behavior sequence
  • Generate test scenarios at three prompt strictness levels
  • Run an agent and record its tool calls
  • Classify tool calls against the expected sequence
  • Report compliance rates with full tool call timelines

Getting started

  1. Have a markdown file that defines the skill, rule, or agent behavior you want to test.
  2. Add skill-comply to your agent setup so it can read that markdown target.
  3. Point skill-comply at the markdown target to generate scenarios at three prompt strictness levels.
  4. Run the generated scenarios so the agent's tool calls are recorded and classified.
  5. Read the self-contained report to see compliance scores and tool call timelines.

What this skill tells your AI

The instructions your AI receives, as published by affaan-m/ecc in skills/skill-comply/SKILL.md and read by ahel’s review.

Measures whether coding agents actually follow skills, rules, or agent definitions by:

  1. Auto-generating expected behavioral sequences (specs) from any .md file
  2. Auto-generating scenarios with decreasing prompt strictness (supportive → neutral → competing)
  3. Running claude -p and capturing tool call traces via stream-json
  4. Classifying tool calls against spec steps using LLM (not regex)
  5. Checking temporal ordering deterministically
  6. Generating self-contained reports with spec, prompts, and timelines

Supported Targets

  • Skills (skills/*/SKILL.md): Workflow skills like search-first, TDD guides
  • Rules (rules/common/*.md): Mandatory rules like testing.md, security.md, git-workflow.md
  • Agent definitions (agents/*.md): Whether an agent gets invoked when expected (internal workflow verification not yet supported)

When to Activate

  • User runs /skill-comply <path>
  • User asks "is this rule actually being followed?"
  • After adding new rules/skills, to verify agent compliance
  • Periodically as part of quality maintenance

Usage

# Full run
uv run python -m scripts.run ~/.claude/rules/common/testing.md

# Dry run (no cost, spec + scenarios only)
uv run python -m scripts.run --dry-run ~/.claude/skills/search-first/SKILL.md

# Custom models
uv run python -m scripts.run --gen-model haiku --model sonnet <path>

Key Concept: Prompt Independence

Measures whether a skill/rule is followed even when the prompt doesn't explicitly support it.

Report Contents

Reports are self-contained and include:

  1. Expected behavioral sequence (auto-generated spec)
  2. Scenario prompts (what was asked at each strictness level)
  3. Compliance scores per scenario
  4. Tool call timelines with LLM classification labels

Advanced (optional)

For users familiar with hooks, reports also include hook promotion recommendations for steps with low compliance. This is informational — the main value is the compliance visibility itself.

Signals

GitHub stars
270k
Forks
40k
Last commit
Sep 2026

Questions

What are compliance skills?
Compliance skills measure whether an agent follows the skills, rules, or agent definitions it was given. skill-comply generates scenarios, runs the agent, classifies its tool calls against the expected sequence, and reports compliance rates.
What kinds of files can I test?
Any markdown target that defines a skill, a rule, or an agent definition. skill-comply turns that markdown into an expected behavior sequence and checks the agent's tool calls against it.
Advanced
Item type
skill
Key
skill-comply
Source
github.com/affaan-m/ecc