Add Model to TabArena
SkillFiles & storageLets your agent add a Claude skill-style scaffold that integrates a new ML model into the TabArena benchmark.
Use Add Model to TabArena in Claude, ChatGPT or Ahel Desktop
Free. Sign in, add Add Model to TabArena and connect your AI. About a minute.
Also: Claude Code · Cursor · Codex
Then ask your AI: use the Add Model to TabArena skill
Details
Instructions available. Your AI can read the instructions. Execution depends on the setup they require.
Account requirements not reviewed. Check the skill instructions before use; Ahel provides instructions and does not run this skill.
No other account needed.
Add Ahel to your AI once: Claude, ChatGPT, Cursor, Claude Code or Codex. Then ask it to use this.
About this skill
Add a new ML model to the TabArena benchmark system. Use this skill whenever the user wants to integrate a new tabular ML model into TabArena, even if they just say "add X model", "integrate X", "support X", or "wrap X for the benchmark". Creates all required files: the AutoGluon model wrapper, the
What this skill tells your AI
The instructions your AI receives, as published by autogluon/tabarena in .claude/skills/add-model/SKILL.md and read by Ahel’s review.
This skill integrates a new tabular ML model into the TabArena benchmark.
Every model lives in one folder at packages/tabarena/src/tabarena/models/<ModelKey>/. That folder contains the wrapper, the HPO generator, and the metadata — and is auto-discovered by tabarena.models._registry.discover_models(). There is no separate benchmark/models/ag/ layout anymore.
Per model, you create up to 5 source files, then edit three existing files. There is no per-model test file — the model is fit-tested automatically by the registry-driven tests/tabarena/models/test_all_models.py.
First: single model or external system?
TabArena has two integration paths; this skill's steps implement the single model path, which is the default:
- Single model (default — everything below): one AutoGluon wrapper in
models/<ModelKey>/, fit by TabArena's shared harness. It gets the shared preprocessing + validation protocol, the bagged / holdout / outer execution modes, an HPO search space, registry auto-discovery, and leaderboard integration. - External system: a self-contained ML system that does its own preprocessing, validation,
HPO, and/or ensembling — AutoML frameworks, multi-model stacks, LLM/agent pipelines. Systems get
their own folder and registry,
packages/tabarena/src/tabarena/systems/<key>/(system.py= theExternalSystemModelsubclass,hpo.py= theSystemConfigGenerator,info.py= theSystemInfo), and run through a bundle insystem_experiments=Truemode. That path is theadd-systemskill, not this one. Runnable references:examples/benchmarking/run_quickstart_tabarena_system.py,examples/beyondarena/run_quickstart_beyondarena_system.pyandexamples/advanced/run_async_tabarena_api_system.py(async/API-driven).
Ask instead of assuming: if what the user wants to add looks like a system — it ensembles or stacks multiple models, runs its own HPO or validation splits, or is described as a "framework", "AutoML tool", "pipeline", or "agent" — ask which integration they want before proceeding. When nothing suggests a system, take the single-model path and continue with Step 0.
Step 0: Gather inputs
Parse $ARGUMENTS for the model name. Then collect (ask only for what's missing or unclear):
| Input | Example | Notes |
|---|---|---|
ModelName | "TabPFN-2.6" | Human-readable display name |
ModelKey | "tabpfnv26" | Snake_case folder/file key (derive from ModelName) |
ClassName | "TabPFNv26" | CamelCase class prefix (derive from ModelName) |
ag_key | "TA-TABPFN-2.6" | AutoGluon registry key; prefix with "TA-" |
ag_name | "TA-TabPFN-2.6" | AutoGluon display name; same as ag_key with proper casing |
pip_package | "tabpfn>=7.0.0" | Pip install spec for pyproject.toml |
doc_url | "https://..." | Documentation / GitHub / paper URL |
model_type | foundation | foundation, torch, or sklearn |
supports_gpu | true | Whether the model uses GPU |
problem_types | binary,multiclass,regression | Supported task types |
Deriving keys: "TabPFN-2.6" → key "tabpfnv26", class prefix "TabPFNv26". "TabSTAR" → key "tabstar", class prefix "TabStar". Strip hyphens, lowercase for key; CamelCase for class.
Step 1: Understand the model API
If doc_url was provided, fetch it with WebFetch to understand:
- Import path (e.g.,
from tabstar.tabstar_model import TabSTARClassifier) - Constructor parameters and their defaults
.fit(X, y, ...)signature.predict()/.predict_proba()signature- Key hyperparameters to expose
Step 2: Pick the right base class and reference model
Choose the most similar existing model to read for detailed inspiration:
| Model type | Base class | Read this reference model |
|---|---|---|
| Torch-based foundation / pre-trained GPU (e.g. TabPFN, TabICL) | AbstractTorchModel with a shared_weights declaration | packages/tabarena/src/tabarena/models/tabicl/model.py (the declaration of Step 3g); exaone_tabular/ for a library that loads inside its constructor, tabdpt/ for a loader the constructor calls |
| Torch NN trained from scratch (e.g. TabM, RealMLP) | AbstractTorchModel | packages/tabarena/src/tabarena/models/tabm/model.py |
| Non-torch GPU model (e.g. JAX/Flax like TabFM, or any lib that manages its own device) | AbstractModel | packages/tabarena/src/tabarena/models/tabstar/model.py |
| CPU / sklearn-like (e.g. KNN) | AbstractModel | packages/tabarena/src/tabarena/models/knn/model.py |
AbstractTorchModel is only for torch-based models. Its whole purpose is the torch device
machinery — get_device() / _set_device() are abstract and the load path calls torch.cuda.is_available()
(so a non-torch device string like "gpu" would crash it). If the model is not torch (JAX/Flax,
or any library that manages device placement itself at the process level, e.g. via
CUDA_VISIBLE_DEVICES / jax.devices()), inherit AbstractModel even though it runs on GPU, and
just add the GPU resource attributes (default_num_gpus, minimum_num_gpus,
_default_ag_args_ensemble_extra with sequential_local, plus _more_tags) — do not
implement get_device/_set_device. tabstar/model.py (a GPU foundation model on AbstractModel)
is the reference; tabfm/model.py is the JAX example.
Read the reference model file now (use the Read tool). Use it as a structural guide — you will adapt rather than copy.
Also read the annotated patterns in references/model_patterns.md — it contains templates for model.py, hpo.py, and info.py.
Step 3: Create new files
Create these files (paths relative to the repo root):
3a. packages/tabarena/src/tabarena/models/{ModelKey}/__init__.py
Re-export the public symbols so from tabarena.models.{ModelKey} import ... works:
from __future__ import annotations
from tabarena.models.{ModelKey}.hpo import gen_{ModelKey}
from tabarena.models.{ModelKey}.info import {ModelKey}_info, {ModelKey}_method_metadata
__all__ = ["gen_{ModelKey}", "{ModelKey}_info", "{ModelKey}_method_metadata"]
3b. packages/tabarena/src/tabarena/models/{ModelKey}/model.py
The AutoGluon wrapper class. Use the template in references/model_patterns.md section "Model wrapper template". Key points:
- Start with
from __future__ import annotations - Inherit from
AbstractTorchModel(torch-based models) orAbstractModel(CPU models and non-torch GPU models — see Step 2: JAX/Flax etc. useAbstractModel) - Set
ag_key,ag_name,ag_priority = 65,seed_name = "random_state", and_supported_problem_types = [...] - Implement
_fit()and_set_default_params() - Declare config as class attributes, not override methods (AutoGluon 1.6). Read
references/model_patterns.md→ "Declare config as class attributes". Overridingsupported_problem_types()is the one AutoGluon actively rejects:verify_modelraises, so the model's smoke test fails. The others (_get_default_resources,get_minimum_resources,_get_default_ag_args_ensemble,_get_default_auxiliary_params) still work but are the old style. Never mutateself.params/self.params_auxafter construction — it raises in 1.7. - Honor the
_fitcontract (readreferences/model_patterns.md→ "The_fitcontract"). The most common review findings on new wrappers are: ignoring the providedX_val/y_val(and instead auto-splitting a second holdout), ignoringtime_limit, hardcoding the thread count instead of wiringnum_cpus, and label-encoding +fillna(0)categoricals when the library handles them natively.models/realmlp/model.pyis the reference for all of these. (In-context-learning foundation models have no train loop / no eval set, so they legitimately ignoretime_limit+X_val— seesap_rpt_oss/tabstar/tabfm.) - For GPU models: also set
default_resources_physical_cores_only = True,default_num_gpus = 1,minimum_num_gpus = 1, and_default_ag_args_ensemble_extra(withfold_fitting_strategy: sequential_local— andrefit_folds: Truefor foundation/pre-trained TFMs; see the "Foundation models: setrefit_folds=True" note inreferences/model_patterns.md. From-scratch NNs omit it), plus_more_tags()(withcan_refit_full: True). Do not declare acan_estimate_memory_usage_statictag: AutoGluon derives it from whether you implement_estimate_memory_usage_static. Only torch models (AbstractTorchModel) additionally implementget_device()/_set_device(); non-torch GPU models onAbstractModelmust NOT (they have no.to(device)). - Docstring must include: description, paper title, authors, codebase URL, license
- Keep optional third-party imports (the wrapped library itself) inside
_fit/ per-method scope so importing this module never requires the optional dep at top-level - Decide the model's untimed warm-up (Step 3g) while you have the library docs in hand
- Foundation model with Hugging Face weights: always pin
revision=on everyhf_hub_download/snapshot_downloadcall the wrapper or itsprefetch_weightsmakes, and pass the pinned file to the library where it takes a path; never resolve against the repo's moving default branch. See "Foundation-model weights: always pin the HF checkpoint revision" inreferences/model_patterns.mdfor how to resolve the commit.
3c. packages/tabarena/src/tabarena/models/{ModelKey}/hpo.py
The search-space generator. By default use an empty search space (like TabPFN-2.6) — only add hyperparameters if the user explicitly asks or if the model has obvious tunable knobs. See template in references/model_patterns.md section "hpo.py template".
3d. packages/tabarena/src/tabarena/models/{ModelKey}/info.py
Defines {ModelKey}_method_metadata: MethodMetadata and {ModelKey}_info: ModelInfo. info.py is the single source the auto-discovery registry walks — populating it correctly is how the model becomes visible to discover_models(). See template in references/model_patterns.md section "info.py template".
When you set ag_key/model_key here, also classify the model in get_model_family (Step 4d) — those keys decide the model's leaderboard family, and skipping this makes it show as ❓ Other on the website.
3e. Multi-file support code (optional)
If the wrapper needs helper modules (preprocessors, vendored upstream code, large internal classes), put them in a private subfolder of packages/tabarena/src/tabarena/models/{ModelKey}/:
_internal/— for hand-written helpers (preprocessors, internal classes, adapters)_vendor/— only for code copied verbatim from an upstream project; keep the original layout/license alongside
Both subfolders need their own empty __init__.py. Import them from model.py via absolute paths, e.g. from tabarena.models.{ModelKey}._internal.preprocessing import Preprocessor.
3f. Test config (no per-model test file)
There is no per-model test file. tests/tabarena/models/test_all_models.py
is parametrized over the model registry, so it fits the new model automatically once
its info.py is discoverable. It skips on ImportError (optional dep missing) and for
GPU-only models without CUDA.
Only touch tests/tabarena/models/smoke_configs.py if the model's toy fit needs
a speed-up: add one entry to SMOKE_OVERRIDES, keyed by the model's MethodMetadata.method
(the registry key), e.g. "{ModelName}": ModelSmokeTest({"max_epochs": 1}), or
ModelSmokeTest(problem_types=("regression",)) for a regression-only model. If the model
fits fine with default hyperparameters on all problem types, add nothing. A wrapper that declares its
cheapness knobs as the cheap_hyperparameters ClassVar (Step 3g) needs no entry either: smoke_for
merges them into the smoke config and the warm-up dummy fit uses the same dict.
3g. Warm-up (untimed environment warm-up): decide, don't skip
TabArena runs an untimed warm-up before every timed fit (AbstractExecModel.warmup_fn, dispatched
per model class by tabarena.models.warmup.warmup_model_cls), so one-time per-environment costs
(library imports, JIT and kernel compilation, the CUDA context, pretrained weights) stay out of
time_train_s / time_infer_s and the fit time limit. The layers are additive and run in this
order for every model; a new model declares only what the generic layers cannot see:
- An optional
warmupclassmethod,warmup(cls, *, problem_type=None, num_cpus=None, num_gpus=None, hyperparameters=None, **kwargs) -> None, for work the declarative layers cannot express (a library's own kernel pre-compilation, allocator settings that must precede the CUDA context). References:chimeraboost,mitra_v2. - Torch import plus CUDA context for
AbstractTorchModelsubclasses (automatic). warmup_modules: ClassVar[tuple[str, ...]] = ("yourlib", "yourlib.submodule"): the modules_fitand_predictimport lazily, merged over the MRO. A"torch"entry also creates the CUDA context for a torch-backed model on plainAbstractModel.- The
ag_keymap inwarmup.pyfor AutoGluon built-ins (LightGBM, CatBoost, XGBoost, ...). - A dummy fit and predict on a small synthetic dataset (
make_synthetic_frames, fixed seed, no task data), which triggers the lazy imports, library caches and kernel loads a first fit pays. For a class that declaresshared_weights(below) it also builds and registers the network the timed fit reuses. Opt-outs on the class:warmup_dummy_fit = False,warmup_dummy_fit_kwargs(n_rows,n_features,n_categorical,time_limit) andcheap_hyperparameters(cheapness knobs such asn_estimators=1, merged over the config; never an input of the network loader, since the dummy fit must build the network the real fit looks up).
| Situation | Declare |
|---|---|
| sklearn-like or lightweight | Nothing; the dummy fit covers it. |
Torch model on AbstractTorchModel | warmup_modules with the heavy extra imports (transformers, ...). |
Torch-backed model on AbstractModel | warmup_modules = ("torch", "yourlib"). References: modernnca, xrfm, tabstar. |
| Library JIT-compiles kernels (numba, JAX, custom CUDA) | A warmup classmethod calling the library's pre-compile entry point when one exists (chimeraboost, whose warmup() needs chimeraboost>=0.14.1). Ask the user for the entry point and minimum version when the docs do not say. |
| Foundation model with pretrained weights | A shared_weights declaration (below) plus warmup_modules for the library. |
The library itself blocks the warm-up or the sharing (import-time global side effect, lazy load at the first predict, load inside __init__, leaked global state) | A developer fix (below): the smallest workaround, headed Developer fix, with the upstream ask written down. |
Fairness contract (the tabarena/models/warmup.py module docstring is the reference): data-independent
work only; never task data, never task- or data-specific state carried into the fit, never a global
random number generator advanced.
When the library gets in the way: the developer-fix pattern. The warm-up and the shared weights
assume a library that imports without side effects, loads its network in one separable call and
does not touch process-global state. Several do not, and the fix belongs upstream. Do not skip the
warm-up for such a model (warmup_modules = (), warmup_dummy_fit = False were the interim answer
for iLTM and are gone); write the smallest workaround in the wrapper instead, and mark it so it can be
found and removed later. The four shapes seen so far, with the in-tree reference for each:
| The library... | Developer fix | Reference |
|---|---|---|
flips a global torch flag or the root logger when imported or fitted (TF32, cuDNN, logging.basicConfig) | pin the flag explicitly in _fit right after the import (one policy for every fit, whatever imported first) and save/restore the rest in a context manager around the fit | iltm/model.py (_isolate_iltm_global_state, allow_tf32 = False) |
| builds its network only at the first predict | build it at the end of _fit, so the load is shared and the timed predict starts warm | nori/model.py (the eager _get_predictor() call) |
loads the checkpoint inside __init__, with no _load_model and no network= argument | a load_network(...) function in <model>/_estimators.py replicating the loading half of the constructor (declared as the shared_weights loader), plus a constructor replica when the constructor cannot take a prebuilt network, guarded against library bumps | exaone_tabular/_estimators.py (TabDPT had one until layer6ai-labs/TabDPT-inference#79 merged) |
declares no logger until __init__ ran, breaks after an unpickle, or has another bug the fit path hits | the narrowest patch at the call site, idempotent, applied where the code path enters the library | iltm/model.py (_ensure_iltm_logger_patched) |
The header is the contract: the module docstring, function docstring or comment opens with
Developer fix (or Developer fix: inline), then says what the library does, what it should offer
instead, the version the fix was written against, and the upstream issue or PR once filed. Before
writing one, ask the library's maintainers for the seam (a separable loader, a network= argument,
an import without side effects); file the issue or PR, link it from the header, and remove the fix
when the library ships it. TabDPT is the precedent: its constructor replica linked
layer6ai-labs/TabDPT-inference#79, and once that merged tabdpt/model.py declared the upstream
_load_model like TabICL, with the extra moved to the first release that carries it (tabdpt>=1.3.1).
grep -rn "Developer fix" packages/tabarena/src/tabarena/models lists what is outstanding.
Report every developer fix you add in Step 8.
Foundation models share one network. A bagged fit of an in-context model would otherwise build
the same frozen network once per fold child and once more for the refit child, inside the timed fit.
Every such library builds its network inside one call its fit makes: a method on the estimator
(tabpfn _initialize_model_variables, tabicl _load_model), a module function (causilo
load_pretrained_model) or a classmethod (Tab2D.from_pretrained). The wrapper names that call and
the inputs that decide which network it builds; AutoGluon's AbstractTorchModel does the rest:
from autogluon.core.models.abstract import SharedWeights
class {ClassName}Model(AbstractTorchModel):
shared_weights: ClassVar[SharedWeights] = SharedWeights(
loader=("somelib:SomeClassifier._load_model", "somelib:SomeRegressor._load_model"),
key=("checkpoint_version", "model_path"), # the loader's inputs that pick the network
disabled_by=("kv_cache",), # inputs under which the library writes into it
)
cheap_hyperparameters: ClassVar[dict] = {"n_estimators": 1}
loader is "package.module:Class.method" for a method the estimator's fit calls, or
"package.module:function" (also Class.classmethod) for a call that returns the network; several
when the library has one class per task. key names the estimator attributes (method) or call
arguments (function) that decide which network is built; the device is part of the key by itself,
from the loader's device input or, for a loader without one, from self.device when _fit sets it
before the fit, else the fit's num_gpus. disabled_by lists inputs (names, or a predicate over the
inputs) under which the library writes into or casts the network, so that fit builds its own.
copy_per_fit=True is for a fine-tuning model: every fit trains a deep copy and pickles it. Nothing in
_fit changes: the library's own code runs on the first call in the process and registers what it
produced; every later call with the same inputs and device reuses it, the fitted model pickles without
the network and takes it back on load, set_device swaps registry entries, and
get_info()["shared_weights"] records it. Sharing is off for a fit with
ag_args_fit={"share_pretrained_weights": False}, per class through
model_class_settings={"<ag_key>": {"share_weights": False}}, and for a configuration named in
disabled_by; fold fitting and refit_folds stay the wrapper's ordinary ensemble arguments.
The library contract. What a library has to offer is one loader call whose inputs are estimator
parameters (checkpoint name or path, device, and any flag that shapes the network) and whose effect
is to produce the network, as a return value or as attributes it sets. Check this before integrating:
if the library loads inside __init__ (EXAONE Tabular; TabDPT until its maintainers merged a
separable _load_model), ask the maintainers for a separable loader or a network= constructor
argument first, and only then fall back to an adapter in <model>/_estimators.py: a
load_network(...) function that replicates the loading half of the constructor (the declared
loader) and, when the constructor cannot take a prebuilt network, a constructor replica. Mark the
module as a developer fix in its docstring, say what the library should offer instead, and guard the
replica against library bumps (exaone_tabular/_estimators.py). A library whose __setstate__
re-runs its loader (tabicl) reloads shared for free; one that pickles its network (tabpfn) is covered
by the generic weightless pickle. A loader the constructor calls (tabdpt _load_model) works too, as
long as the keyed attributes are set before the call; an input the constructor resolves after it
(tabdpt compile) cannot go in disabled_by, so the wrapper overrides _shares_weights for it.
Wrappers that do not care about sharing simply declare nothing (TabDPT v1.1 and TabDPT-Turbo, whose
releases have no separable loader, TabPFN-Wide, SAP-RPT-OSS, iLTM whose library caches for itself).
references/model_patterns.md under "Shared pretrained weights" has the field-by-field reference;
tabicl/model.py is the plain in-tree case, causilo/model.py a function loader without a device
input, mitra_v2/model.py a fine-tuning model, tabdpt/model.py a loader the constructor calls and
exaone_tabular/ the adapter.
Tests: tests/tabarena/models/test_shared_weights_models.py discovers every registered class that
declares shared_weights and checks the declaration (well-formed, the loader resolves when the library
is installed, the keyed inputs are the loader's parameters, the named disablers switch sharing off);
nothing to register per wrapper. pytest -m models tests/tabarena/models/test_shared_weights_models.py -k <Model> fits the real library twice, asserts one network, a weightless pickle that reloads and
predicts the same, and equal predictions with sharing switched off.
Inference side. The exec model persists the fitted model in memory around the predict timer
(AGWrapper.persist, through persist_inference.persist_for_inference) and calls an optional
instance method prepare_for_inference(self) -> None on every persisted object, bagged children
included. Its contract is written once, in AbstractExecModel.pre_predict (exec_models/base.py);
refer to it instead of restating it. Data-dependent first-predict work (torch.compile on the real
batch, kernel dispatch for the real shapes) is measured by design.
Shortened here. Read the whole file on GitHub.
Signals
- GitHub stars
- 336
- Forks
- 79
- Last commit
- Oct 2026
Ahel review
K1binfo
installs-packages
Automated review, not a security audit. Ruleset v1+k2.
Advanced
- Item type
- skill
- Key
add-model- Source
- github.com/autogluon/tabarena
Related picks
Skill · wshobson
The pick for Pythonpython-pro
Skill · jeffallan
The pick for Pythonpptx
Skill · anthropics
More in Files & storagedocx
Skill · anthropics
More in Files & storageresearch
Skill · mattpocock
More in Files & storageto-tickets
Skill · mattpocock
More in Files & storage