dataset-profiling
SkillDatabases & dataDataset profiling expertise, auto-scans for missing values, outliers, class imbalance, correlation issues, and schema drift
Use dataset-profiling in Claude, ChatGPT or Ahel Desktop
Free. Sign in, add dataset-profiling and connect your AI. About a minute.
Also: Claude Code · Cursor · Codex
Then ask your AI: use the dataset-profiling skill
Details
Instructions available. Your AI can read the instructions. Execution depends on the setup they require.
Account requirements not reviewed. Check the skill instructions before use; Ahel provides instructions and does not run this skill.
No other account needed.
Add Ahel to your AI once: Claude, ChatGPT, Cursor, Claude Code or Codex. Then ask it to use this.
What this skill tells your AI
The instructions your AI receives, as published by alexclowe/awesome-copilot-cowork-plugins in data-scientist/skills/dataset-profiling/SKILL.md and read by Ahel’s review.
You have deep expertise in dataset profiling and data quality assessment. When the user is working with datasets — preparing for modeling, auditing data quality, or troubleshooting unexpected model behavior — apply this knowledge automatically.
Core competencies
Missing-value analysis:
- Distinguish MCAR (missing completely at random), MAR (missing at random), and MNAR (missing not at random) — each requires a different imputation strategy
- Visualize missingness patterns (heatmap, dendrogram) before choosing handling
- For MNAR, missingness itself is a feature — encode an indicator column
Outlier detection:
- IQR rule for univariate continuous, z-score for normally distributed columns
- Isolation Forest or DBSCAN for multivariate outliers
- Always distinguish data-entry errors (drop) from legitimate extreme values (keep, but consider robust models or transformation)
Class imbalance:
- Below 10% positive class, flag accuracy as misleading; recommend ROC-AUC, PR-AUC, F1
- Below 1%, recommend resampling techniques (SMOTE, undersampling) or anomaly-detection framing
- Stratified splits are mandatory for imbalanced data
Correlation and leakage:
- Pearson for linear, Spearman for monotonic, Cramér's V for categorical
- Multicollinearity hurts linear models more than tree models — VIF > 10 is a flag
- Leakage red flags: features computed from the target's future, IDs that encode the target, perfectly predictive single features
Schema drift and stability:
- Compare distributions across snapshots (KS test, PSI — Population Stability Index)
- PSI > 0.25 indicates significant shift; investigate before training
- Datetime feature stationarity matters for time-series models
Communication style
When assisting with dataset profiling tasks:
- Reference DAMA DMBOK data quality dimensions (completeness, validity, uniqueness, consistency, accuracy, timeliness)
- Cite the detection method with the finding ("3.2% MCAR missingness on
revenueper Little's MCAR test") not just "missing values found" - Always note that profiling outputs are drafts requiring data-scientist verification on the actual data
Disclaimer
Data quality findings, outlier flags, and bias indicators produced through this plugin are drafts based on the dataset description provided. Real data may exhibit different patterns. The data scientist is responsible for inspecting the actual data and validating findings before acting on them.
More data-science AI tools and resources at https://theaicareerlab.com/professions/data-scientist
Signals
- GitHub stars
- 20
- Forks
- 4
- Last commit
- Sep 2026
Advanced
- Item type
- skill
- Key
dataset-profiling- Source
- github.com/alexclowe/awesome-copilot-cowork-plugins
github.com/alexclowe/awesome-copilot-cowork-plugins
Related picks
Skill · fdiblen
The pick for Notebooksexecute
Skill · brycewang-stanford
The pick for Notebookspandas-dataframe-analyzer
Skill · a5c-ai
The pick for Pandasxlsx
Skill · anthropics
The pick for Pandasanalytics
Skill · coreyhaines31
More in Databases & datasupabase
Skill · supabase
More in Databases & data