Data Quality Frameworks
SkillAI & modelsdata-quality-frameworks is a skill that guides an AI agent through implementing data quality validation with Great Expectations, dbt tests, and data contracts.
Use Data Quality Frameworks in Claude, ChatGPT or Ahel Desktop
Free. Sign in, add Data Quality Frameworks and connect your AI. About a minute.
Also: Claude Code · Cursor · Codex
Then ask your AI: use the Data Quality Frameworks skill
Details
Instructions available. Your AI can read the instructions. Execution depends on the setup they require.
Account requirements not reviewed. Check the skill instructions before use; ahel provides instructions and does not run this skill.
No other account needed.
Have Python available and install Great Expectations with pip.
What your AI can do with it
- Set up Great Expectations suites with expectations for completeness, uniqueness
- Build dbt test suites structured as a testing pyramid from schema to unit to integration
- Run validation checkpoints and generate pass/fail reports per table
- Establish data contracts between teams and version schema contracts
- Automate data validation in CI/CD and alert on failures
- Monitor data quality metrics, including timeliness checks
Getting started
- Have Python available and install Great Expectations with pip.
- Initialize a Great Expectations project and create a datasource.
- Ask the agent to create an expectation suite and add expectations such as non-null and unique column checks.
- Configure a checkpoint, such as a daily validation checkpoint, and run it to produce validation results.
- Consult the skill's references/details.md file when the top-level patterns are not enough for a specific case.
What this skill tells your AI
The instructions your AI receives, as published by wshobson/agents in plugins/data-engineering/skills/data-quality-frameworks/SKILL.md and read by ahel’s review.
Production patterns for implementing data quality with Great Expectations, dbt tests, and data contracts to ensure reliable data pipelines.
When to Use This Skill
- Implementing data quality checks in pipelines
- Setting up Great Expectations validation
- Building comprehensive dbt test suites
- Establishing data contracts between teams
- Monitoring data quality metrics
- Automating data validation in CI/CD
Core Concepts
1. Data Quality Dimensions
| Dimension | Description | Example Check |
|---|---|---|
| Completeness | No missing values | expect_column_values_to_not_be_null |
| Uniqueness | No duplicates | expect_column_values_to_be_unique |
| Validity | Values in expected range | expect_column_values_to_be_in_set |
| Accuracy | Data matches reality | Cross-reference validation |
| Consistency | No contradictions | expect_column_pair_values_A_to_be_greater_than_B |
| Timeliness | Data is recent | expect_column_max_to_be_between |
2. Testing Pyramid for Data
/\
/ \ Integration Tests (cross-table)
/────\
/ \ Unit Tests (single column)
/────────\
/ \ Schema Tests (structure)
/────────────\
Quick Start
Great Expectations Setup
# Install
pip install great_expectations
# Initialize project
great_expectations init
# Create datasource
great_expectations datasource new
# great_expectations/checkpoints/daily_validation.yml
import great_expectations as gx
# Create context
context = gx.get_context()
# Create expectation suite
suite = context.add_expectation_suite("orders_suite")
# Add expectations
suite.add_expectation(
gx.expectations.ExpectColumnValuesToNotBeNull(column="order_id")
)
suite.add_expectation(
gx.expectations.ExpectColumnValuesToBeUnique(column="order_id")
)
# Validate
results = context.run_checkpoint(checkpoint_name="daily_orders")
Detailed patterns and worked examples
Detailed pattern documentation lives in references/details.md. Read that file when the navigation tier above is insufficient.
Summary: {total_passed}/{total_tables} tables passed")
report.append("")
for table, result in results.items():
status = "✅" if result.passed else "❌"
report.append(f"### {status} {table}")
report.append(f"- Expectations: {result.total_expectations}")
report.append(f"- Failed: {result.failed_expectations}")
if not result.passed:
report.append("- Failed checks:")
for detail in result.details:
if not detail["success"]:
report.append(f" - {detail['expectation']}: {detail['observed_value']}")
report.append("")
return "\n".join(report)
Usage
context = gx.get_context() pipeline = DataQualityPipeline(context)
tables_to_validate = { "orders": "orders_suite", "customers": "customers_suite", "products": "products_suite", }
results = pipeline.run_all(tables_to_validate) report = pipeline.generate_report(results)
Fail pipeline if any table failed
if not all(r.passed for r in results.values()): print(report) raise ValueError("Data quality checks failed!")
## Best Practices
### Do's
- **Test early** - Validate source data before transformations
- **Test incrementally** - Add tests as you find issues
- **Document expectations** - Clear descriptions for each test
- **Alert on failures** - Integrate with monitoring
- **Version contracts** - Track schema changes
### Don'ts
- **Don't test everything** - Focus on critical columns
- **Don't ignore warnings** - They often precede failures
- **Don't skip freshness** - Stale data is bad data
- **Don't hardcode thresholds** - Use dynamic baselines
- **Don't test in isolation** - Test relationships too
Signals
- GitHub stars
- 40k
- Forks
- 4k
- Last commit
- Sep 2026
ahel review
K1binfo
installs-packages
Automated review, not a security audit. Ruleset v1+k2.
Others that do the same job
Questions
- When should this skill be used?
- When building data quality pipelines, implementing validation rules, or establishing data contracts. It also fits setting up Great Expectations validation, building dbt test suites, monitoring quality metrics, and automating validation in CI/CD.
- Which data quality dimensions does it cover?
- Completeness, uniqueness, validity, accuracy, consistency, and timeliness, each with example checks such as expect_column_values_to_not_be_null or expect_column_values_to_be_unique.
Advanced
- Item type
- skill
- Key
data-quality-frameworks-wshobson- Source
- github.com/wshobson/agents
Related picks
Skill · mattpocock
The pick for TypeScripttypescript-pro
Skill · jeffallan
The pick for TypeScriptpython-performance-optimization
Skill · wshobson
The pick for Pythonpython-pro
Skill · jeffallan
The pick for Pythonbuilding-dbt-models
Skill · unknown-333
The pick for dbtusing-dbt-for-analytics-engineering
Skill · dbt-labs
The pick for dbt