Skill Tester

SkillAI & models

Skill-tester checks the quality of the skills your AI writes or works with. Once it is added, your AI can validate a skill's structure, safely test its Python scripts, and score how well the skill is built. Each skill gets a letter grade and a quality tier so you know whether it is ready to rely on.

Available today. Use it from your connected AI after setup.

After adding skill-tester, ask your AI to check a skill it has written or one from the claude-skills collection. You will get the test results along with the grade and tier.

Then ask your AI: use the Skill Tester skill

What your AI can do with it

  • Validate a skill's structure and catch problems
  • Test a skill's Python scripts for syntax errors, missing imports, and runtime failures
  • Check that a script's output comes back in the expected format
  • Run a skill's scripts safely
  • Score overall quality with a letter grade
  • Assign a tier such as BASIC or STANDARD based on the score

What this skill tells your AI

The instructions your AI receives, as published by borghei/claude-skills in engineering/skill-tester/SKILL.md and read by ahel’s review.

Validate skill packages for structure compliance, test Python scripts for syntax and stdlib-only imports, and score quality across four dimensions (documentation, code quality, completeness, usability) with letter grades and improvement recommendations. Supports BASIC, STANDARD, and POWERFUL tier classification.

Core Capabilities

  • Structure validation — check required files, directory layout, YAML frontmatter, and required SKILL.md sections against tier thresholds.
  • Script testing — AST syntax checks, stdlib-only import analysis, argparse and __main__-guard detection, runtime --help and sample-data execution.
  • Quality scoring — four equally weighted dimensions producing an overall score, letter grade (A+ through F), and a prioritized improvement roadmap.
  • Tier classification — BASIC / STANDARD / POWERFUL requirements for SKILL.md depth, script count/LOC, argparse, output formats, and error handling.
  • Dual output — human-readable reports and --json for CI/CD gating with meaningful exit codes.

When to Use

  • Creating a new skill and validating it before publishing.
  • Auditing an existing skill's structure, scripts, and quality.
  • Embedding a quality gate into a CI/CD pipeline.

Clarify First

Before validating, confirm these inputs. If any is unknown or vague, ASK — do not assume:

  • Target skill path — the skill directory to validate, test, and score (the subject of all three tools)
  • Target tier — BASIC / STANDARD / POWERFUL (--tier; sets the required sections, script count, and structural thresholds)
  • Pass bar — the minimum quality score / whether failures gate CI (--minimum-score, exit codes)

Stop rule: ask only the 2-3 that most change the output. If the user says "just draft it," proceed and list your assumptions at the top of the artifact.

Quick Start

# Validate skill structure and documentation
python skill_validator.py engineering/my-skill --tier POWERFUL --json

# Test all Python scripts in a skill
python script_tester.py engineering/my-skill --timeout 30

# Score quality with improvement roadmap
python quality_scorer.py engineering/my-skill --detailed --minimum-score 75

Tools

ToolPurposeCommand
skill_validator.pyValidate structure, frontmatter, required sections, and scripts against tier rulespython scripts/skill_validator.py engineering/my-skill --tier POWERFUL --json
script_tester.pyStatic + runtime tests of scripts (syntax, imports, argparse, --help, samples)python scripts/script_tester.py engineering/my-skill --timeout 60 --json
quality_scorer.pyScore four quality dimensions with letter grade and improvement roadmappython scripts/quality_scorer.py engineering/my-skill --detailed --minimum-score 75 --json

See references/tool-reference.md for full parameter tables, output formats, and exit codes.

References

Load the reference that matches the task — keep this file lean and pull detail on demand:

  • references/workflows-and-cicd.md — the three core validation workflows, the tier-requirements and quality-scoring tables, CI/CD integration, anti-patterns, troubleshooting, and success criteria. Read when running a validation pass or wiring a CI gate.
  • references/tool-reference.md — full parameter tables, output formats, and exit codes for the three Python tools. Read when scripting the tools or interpreting JSON output.
  • references/skill-structure-specification.md — the authoritative specification for skill directory structure, required files, and frontmatter. Read when defining what "valid structure" means.
  • references/tier-requirements-matrix.md — the full BASIC/STANDARD/POWERFUL requirements matrix with detailed criteria per tier. Read when classifying or upgrading a skill's tier.
  • references/quality-scoring-rubric.md — the detailed scoring rubric with per-component weights and grading bands. Read when interpreting or tuning quality scores.

Scope & Limitations

Covers:

  • Structural validation of skill directories against tier-specific requirements (BASIC, STANDARD, POWERFUL)
  • Static analysis of Python scripts including syntax checking, import validation, argparse detection, and main guard verification
  • Multi-dimensional quality scoring across documentation, code quality, completeness, and usability
  • Dual output formatting (JSON for CI/CD pipelines, human-readable for developer consumption)

Does NOT cover:

  • Functional correctness of script logic or algorithm accuracy — the tester verifies structure and conventions, not business logic
  • Performance benchmarking or memory profiling of scripts — see engineering/performance-profiler for runtime analysis
  • Security vulnerability scanning of script code — see engineering/skill-security-auditor for dependency and code security audits
  • Cross-skill dependency resolution or integration testing — skills are validated in isolation without verifying inter-skill compatibility

Integration Points

SkillIntegrationData Flow
engineering/skill-security-auditorRun security audit after validation passesskill_validator.py confirms structure compliance, then skill-security-auditor scans for vulnerabilities in the same skill path
engineering/ci-cd-pipeline-builderEmbed skill-tester as a quality gate stagePipeline builder generates workflow YAML that invokes skill_validator.py, script_tester.py, and quality_scorer.py sequentially
engineering/changelog-generatorFeed quality score deltas into changelog entriesCompare quality_scorer.py JSON output between releases to surface quality improvements or regressions
engineering/pr-review-expertAttach validation report to pull request reviewsskill_validator.py --json output is posted as a PR comment for reviewer context
engineering/performance-profilerComplement structural testing with runtime profilingAfter script_tester.py confirms execution succeeds, performance-profiler measures execution time and resource usage
engineering/tech-debt-trackerTrack quality score trends over timePeriodic quality_scorer.py --json output is ingested to detect score degradation and flag technical debt

Signals

GitHub stars
752
Forks
137
Last commit
Aug 2026

ahel review

  • K6low
    bundled executables the agent is told to run

Automated review, not a security audit. Ruleset v1+k2.

Advanced
Catalog kind
skill
Gateway key
skill-tester
Source
github.com/borghei/claude-skills