Register a Hugging Face dataset
SkillDatabases & dataAdd claude skill steps to register a Hugging Face dataset in Marin by checking its schema and writing a datasets module.
Use Register a Hugging Face dataset in Claude, ChatGPT or Ahel Desktop
Free. Sign in, add Register a Hugging Face dataset and connect your AI. About a minute.
Also: Claude Code · Cursor · Codex
Then ask your AI: use the Register a Hugging Face dataset skill
Details
Instructions available. Your AI can read the instructions. Execution depends on the setup they require.
Account requirements not reviewed. Check the skill instructions before use; ahel provides instructions and does not run this skill.
No other account needed.
Add ahel to your AI once: Claude, ChatGPT, Cursor, Claude Code or Codex. Then ask it to use this.
About this skill
Register a named Hugging Face dataset for Marin by inspecting its schema and adding the appropriate experiments/datasets module.
What this skill tells your AI
The instructions your AI receives, as published by marin-community/marin in .agents/skills/add-dataset/SKILL.md and read by ahel’s review.
Inspect the schema without downloading the full dataset:
uv run lib/marin/tools/get_hf_dataset_schema.py <dataset_name> [options]
For programmatic inspection:
from marin.tools.get_hf_dataset_schema import get_schema
schema = get_schema(dataset_name="wikitext", config_name="wikitext-103-v1")
Use repo-managed dependencies. For a one-off inspection without a provisioned
environment, add --with datasets --with pyyaml to uv run.
If the result says a config is required, select one of available_configs and
retry with --config_name. Add --trust_remote_code only after inspecting the
dataset repository and accepting its code-execution boundary. The tool streams;
do not replace it with a full dataset download.
If the dataset cannot be found, stop and report the identifier, path, or access failure instead of guessing a replacement.
Choose the text field from the reported schema. Prefer an exact text field,
then a field containing text, then another string field. Inspect sample_row
to verify the content; it may be empty for some datasets. The result also
reports splits, text_field_candidates, and features.
Add a leaf module under experiments/datasets/ using the lazy builders in
marin.experiment.data:
- expose
<name>_dataset()for one corpus; - expose
<name>_datasets() -> dict[str, ...]for a keyed family; - for Hugging Face subsets, follow
experiments/datasets/nemotron.pyand return one keyed handle per subset.
Validate the selected config, splits, text mapping, and one sample before adding tokenization or downstream experiment configuration.
Signals
- GitHub stars
- 4k
- Forks
- 311
- Last commit
- Oct 2026
Advanced
- Item type
- skill
- Key
add-dataset-marin-community- Source
- github.com/marin-community/marin
github.com/marin-community/marin
Related picks
Skill · wshobson
The pick for Pythonpython-pro
Skill · jeffallan
The pick for Pythonsupabase
Skill · supabase
More in Databases & dataconnect
Skill · composiohq
More in Databases & dataanalytics
Skill · coreyhaines31
More in Databases & dataazure-kusto
Skill · microsoft
More in Databases & data