Canonical Schema
SkillDatabases & dataStandardized data schema across all financial datasets. Use when defining or enforcing column names, types, and index conventions.
Use Canonical Schema in Claude, ChatGPT or Ahel Desktop
Free. Sign in, add Canonical Schema and connect your AI. About a minute.
Also: Claude Code · Cursor · Codex
Then ask your AI: use the Canonical Schema skill
Details
Instructions available. Your AI can read the instructions. Execution depends on the setup they require.
Account requirements not reviewed. Check the skill instructions before use; Ahel provides instructions and does not run this skill.
No other account needed.
Add Ahel to your AI once: Claude, ChatGPT, Cursor, Claude Code or Codex. Then ask it to use this.
What this skill tells your AI
The instructions your AI receives, as published by ml4t/skills in infrastructure/canonical-schema/SKILL.md and read by Ahel’s review.
When every dataset uses different column names - date vs ts_event vs timestamp, asset vs ticker vs symbol - every downstream notebook needs special-case handling.
The Problem
Provider A delivers date, ticker, close; Provider B uses ts_event,
symbol, price; Provider C uses timestamp, asset, adj_close. Without a
canonical schema, every downstream notebook needs provider-specific renames.
The Pattern
WRONG
import polars as pl
# Different column names per dataset - downstream code breaks constantly
etfs = pl.read_parquet("etfs.parquet") # has: date, ticker, close
futures = pl.read_parquet("futures.parquet") # has: ts_event, product, settle
crypto = pl.read_parquet("crypto.parquet") # has: timestamp, symbol, close
# Every notebook needs provider-specific column mapping
if "date" in df.columns:
df = df.rename({"date": "timestamp"})
elif "ts_event" in df.columns:
df = df.rename({"ts_event": "timestamp"})
# Repeat for every column, every dataset, every notebook
CORRECT
import polars as pl
# Canonical schema: enforced once at load time, trusted everywhere after
CANONICAL_COLUMNS = {
"time": "timestamp", # ALL frequencies: daily, hourly, minute
"entity": "symbol", # Exception: cme_futures uses "product"
}
def enforce_schema(df: pl.DataFrame, dataset: str) -> pl.DataFrame:
"""Rename provider columns to canonical names at load time."""
renames = {}
# Time column: accept common variants, output "timestamp"
for variant in ["date", "ts_event", "datetime", "time"]:
if variant in df.columns:
renames[variant] = "timestamp"
# Entity column: accept common variants, output "symbol"
if dataset != "cme_futures": # futures use "product"
for variant in ["asset", "ticker", "pair", "instrument"]:
if variant in df.columns:
renames[variant] = "symbol"
return df.rename(renames)
# Load once, use everywhere - no downstream renames needed
etfs = enforce_schema(pl.read_parquet("etfs.parquet"), "etfs")
assert "timestamp" in etfs.columns
assert "symbol" in etfs.columns
The Two Canonical Columns
| Column | Name | Type | Usage |
|---|---|---|---|
| Time | timestamp | Date or Datetime | Every dataset, every frequency |
| Entity | symbol | Utf8 | All datasets except CME futures |
| Entity (futures) | product | Utf8 | CME futures only (contract identifier) |
OHLCV Columns
Lowercase, no prefix: open, high, low, close, volume. If adjustments exist: adj_close.
Enforcement Point
Schema enforcement happens at load time, not downstream. This means:
- Data loaders validate and rename on return
- Notebooks never import raw provider data directly
- If a notebook gets a
ColumnNotFoundErrorfordateorasset, the notebook is wrong - fix the notebook to usetimestamporsymbol
Guardrails
- Never rename canonical columns back to legacy names in notebooks - fix the notebook
- Never add compatibility shims that accept both old and new names - migrate forward
- If a new provider uses a different name, add the rename in the loader, not in 50 notebooks
productis only for CME futures - do not generalize to other datasets- Check for legacy names (
asset,date,ticker,pair) during code review
Production Implementation
ml4t-data standardizes generic OHLCV fetches to canonical columns:
from ml4t.data import DataManager
dm = DataManager()
panel = dm.batch_load(
["SPY", "QQQ"],
start="2015-01-01",
end="2024-12-31",
provider="yahoo",
)
# Columns: timestamp, symbol, open, high, low, close, volume
Checklist
- All data loaded through loaders that enforce canonical names
- Time column is
timestampfor every dataset and frequency - Entity column is
symbol(orproductfor CME futures only) - OHLCV columns are lowercase:
open,high,low,close,volume - No legacy names (
date,asset,ticker) in notebooks; schema enforced at load time
Signals
- GitHub stars
- 22
- Forks
- 11
- Last commit
- Oct 2026
Advanced
- Item type
- skill
- Key
ml4t-canonical-schema- Source
- github.com/ml4t/skills
Related picks
Skill · fdiblen
The pick for Notebooksexecute
Skill · brycewang-stanford
The pick for Notebookspandas-dataframe-analyzer
Skill · a5c-ai
The pick for Pandasxlsx
Skill · anthropics
The pick for Pandaspython-performance-optimization
Skill · wshobson
The pick for Pythonpython-pro
Skill · jeffallan
The pick for Python