Add Model to TabArena

SkillFiles & storage

Lets your agent add a Claude skill-style scaffold that integrates a new ML model into the TabArena benchmark.

Use Add Model to TabArena in Claude, ChatGPT or Ahel Desktop

Free. Sign in, add Add Model to TabArena and connect your AI. About a minute.

Also: Claude Code · Cursor · Codex

Then ask your AI: use the Add Model to TabArena skill

Details

Instructions available. Your AI can read the instructions. Execution depends on the setup they require.

Add Ahel to your AI once: Claude, ChatGPT, Cursor, Claude Code or Codex. Then ask it to use this.

Add Model to TabArenaStart free
About this skill

Add a new ML model to the TabArena benchmark system. Use this skill whenever the user wants to integrate a new tabular ML model into TabArena, even if they just say "add X model", "integrate X", "support X", or "wrap X for the benchmark". Creates all required files: the AutoGluon model wrapper, the

What this skill tells your AI

The instructions your AI receives, as published by autogluon/tabarena in .claude/skills/add-model/SKILL.md and read by Ahel’s review.

This skill integrates a new tabular ML model into the TabArena benchmark.

Every model lives in one folder at packages/tabarena/src/tabarena/models/<ModelKey>/. That folder contains the wrapper, the HPO generator, and the metadata — and is auto-discovered by tabarena.models._registry.discover_models(). There is no separate benchmark/models/ag/ layout anymore.

Per model, you create up to 5 source files, then edit three existing files. There is no per-model test file — the model is fit-tested automatically by the registry-driven tests/tabarena/models/test_all_models.py.

First: single model or external system?

TabArena has two integration paths; this skill's steps implement the single model path, which is the default:

  • Single model (default — everything below): one AutoGluon wrapper in models/<ModelKey>/, fit by TabArena's shared harness. It gets the shared preprocessing + validation protocol, the bagged / holdout / outer execution modes, an HPO search space, registry auto-discovery, and leaderboard integration.
  • External system: a self-contained ML system that does its own preprocessing, validation, HPO, and/or ensembling — AutoML frameworks, multi-model stacks, LLM/agent pipelines. Systems get their own folder and registry, packages/tabarena/src/tabarena/systems/<key>/ (system.py = the ExternalSystemModel subclass, hpo.py = the SystemConfigGenerator, info.py = the SystemInfo), and run through a bundle in system_experiments=True mode. That path is the add-system skill, not this one. Runnable references: examples/benchmarking/run_quickstart_tabarena_system.py, examples/beyondarena/run_quickstart_beyondarena_system.py and examples/advanced/run_async_tabarena_api_system.py (async/API-driven).

Ask instead of assuming: if what the user wants to add looks like a system — it ensembles or stacks multiple models, runs its own HPO or validation splits, or is described as a "framework", "AutoML tool", "pipeline", or "agent" — ask which integration they want before proceeding. When nothing suggests a system, take the single-model path and continue with Step 0.

Step 0: Gather inputs

Parse $ARGUMENTS for the model name. Then collect (ask only for what's missing or unclear):

InputExampleNotes
ModelName"TabPFN-2.6"Human-readable display name
ModelKey"tabpfnv26"Snake_case folder/file key (derive from ModelName)
ClassName"TabPFNv26"CamelCase class prefix (derive from ModelName)
ag_key"TA-TABPFN-2.6"AutoGluon registry key; prefix with "TA-"
ag_name"TA-TabPFN-2.6"AutoGluon display name; same as ag_key with proper casing
pip_package"tabpfn>=7.0.0"Pip install spec for pyproject.toml
doc_url"https://..."Documentation / GitHub / paper URL
model_typefoundationfoundation, torch, or sklearn
supports_gputrueWhether the model uses GPU
problem_typesbinary,multiclass,regressionSupported task types

Deriving keys: "TabPFN-2.6" → key "tabpfnv26", class prefix "TabPFNv26". "TabSTAR" → key "tabstar", class prefix "TabStar". Strip hyphens, lowercase for key; CamelCase for class.

Step 1: Understand the model API

If doc_url was provided, fetch it with WebFetch to understand:

  • Import path (e.g., from tabstar.tabstar_model import TabSTARClassifier)
  • Constructor parameters and their defaults
  • .fit(X, y, ...) signature
  • .predict() / .predict_proba() signature
  • Key hyperparameters to expose

Step 2: Pick the right base class and reference model

Choose the most similar existing model to read for detailed inspiration:

Model typeBase classRead this reference model
Torch-based foundation / pre-trained GPU (e.g. TabPFN, TabICL)AbstractTorchModel with a shared_weights declarationpackages/tabarena/src/tabarena/models/tabicl/model.py (the declaration of Step 3g); exaone_tabular/ for a library that loads inside its constructor, tabdpt/ for a loader the constructor calls
Torch NN trained from scratch (e.g. TabM, RealMLP)AbstractTorchModelpackages/tabarena/src/tabarena/models/tabm/model.py
Non-torch GPU model (e.g. JAX/Flax like TabFM, or any lib that manages its own device)AbstractModelpackages/tabarena/src/tabarena/models/tabstar/model.py
CPU / sklearn-like (e.g. KNN)AbstractModelpackages/tabarena/src/tabarena/models/knn/model.py

AbstractTorchModel is only for torch-based models. Its whole purpose is the torch device machinery — get_device() / _set_device() are abstract and the load path calls torch.cuda.is_available() (so a non-torch device string like "gpu" would crash it). If the model is not torch (JAX/Flax, or any library that manages device placement itself at the process level, e.g. via CUDA_VISIBLE_DEVICES / jax.devices()), inherit AbstractModel even though it runs on GPU, and just add the GPU resource attributes (default_num_gpus, minimum_num_gpus, _default_ag_args_ensemble_extra with sequential_local, plus _more_tags) — do not implement get_device/_set_device. tabstar/model.py (a GPU foundation model on AbstractModel) is the reference; tabfm/model.py is the JAX example.

Read the reference model file now (use the Read tool). Use it as a structural guide — you will adapt rather than copy.

Also read the annotated patterns in references/model_patterns.md — it contains templates for model.py, hpo.py, and info.py.

Step 3: Create new files

Create these files (paths relative to the repo root):

3a. packages/tabarena/src/tabarena/models/{ModelKey}/__init__.py

Re-export the public symbols so from tabarena.models.{ModelKey} import ... works:

from __future__ import annotations

from tabarena.models.{ModelKey}.hpo import gen_{ModelKey}
from tabarena.models.{ModelKey}.info import {ModelKey}_info, {ModelKey}_method_metadata

__all__ = ["gen_{ModelKey}", "{ModelKey}_info", "{ModelKey}_method_metadata"]

3b. packages/tabarena/src/tabarena/models/{ModelKey}/model.py

The AutoGluon wrapper class. Use the template in references/model_patterns.md section "Model wrapper template". Key points:

  • Start with from __future__ import annotations
  • Inherit from AbstractTorchModel (torch-based models) or AbstractModel (CPU models and non-torch GPU models — see Step 2: JAX/Flax etc. use AbstractModel)
  • Set ag_key, ag_name, ag_priority = 65, seed_name = "random_state", and _supported_problem_types = [...]
  • Implement _fit() and _set_default_params()
  • Declare config as class attributes, not override methods (AutoGluon 1.6). Read references/model_patterns.md → "Declare config as class attributes". Overriding supported_problem_types() is the one AutoGluon actively rejects: verify_model raises, so the model's smoke test fails. The others (_get_default_resources, get_minimum_resources, _get_default_ag_args_ensemble, _get_default_auxiliary_params) still work but are the old style. Never mutate self.params / self.params_aux after construction — it raises in 1.7.
  • Honor the _fit contract (read references/model_patterns.md → "The _fit contract"). The most common review findings on new wrappers are: ignoring the provided X_val/y_val (and instead auto-splitting a second holdout), ignoring time_limit, hardcoding the thread count instead of wiring num_cpus, and label-encoding + fillna(0) categoricals when the library handles them natively. models/realmlp/model.py is the reference for all of these. (In-context-learning foundation models have no train loop / no eval set, so they legitimately ignore time_limit + X_val — see sap_rpt_oss/tabstar/tabfm.)
  • For GPU models: also set default_resources_physical_cores_only = True, default_num_gpus = 1, minimum_num_gpus = 1, and _default_ag_args_ensemble_extra (with fold_fitting_strategy: sequential_local — and refit_folds: True for foundation/pre-trained TFMs; see the "Foundation models: set refit_folds=True" note in references/model_patterns.md. From-scratch NNs omit it), plus _more_tags() (with can_refit_full: True). Do not declare a can_estimate_memory_usage_static tag: AutoGluon derives it from whether you implement _estimate_memory_usage_static. Only torch models (AbstractTorchModel) additionally implement get_device() / _set_device(); non-torch GPU models on AbstractModel must NOT (they have no .to(device)).
  • Docstring must include: description, paper title, authors, codebase URL, license
  • Keep optional third-party imports (the wrapped library itself) inside _fit / per-method scope so importing this module never requires the optional dep at top-level
  • Decide the model's untimed warm-up (Step 3g) while you have the library docs in hand
  • Foundation model with Hugging Face weights: always pin revision= on every hf_hub_download / snapshot_download call the wrapper or its prefetch_weights makes, and pass the pinned file to the library where it takes a path; never resolve against the repo's moving default branch. See "Foundation-model weights: always pin the HF checkpoint revision" in references/model_patterns.md for how to resolve the commit.

3c. packages/tabarena/src/tabarena/models/{ModelKey}/hpo.py

The search-space generator. By default use an empty search space (like TabPFN-2.6) — only add hyperparameters if the user explicitly asks or if the model has obvious tunable knobs. See template in references/model_patterns.md section "hpo.py template".

3d. packages/tabarena/src/tabarena/models/{ModelKey}/info.py

Defines {ModelKey}_method_metadata: MethodMetadata and {ModelKey}_info: ModelInfo. info.py is the single source the auto-discovery registry walks — populating it correctly is how the model becomes visible to discover_models(). See template in references/model_patterns.md section "info.py template".

When you set ag_key/model_key here, also classify the model in get_model_family (Step 4d) — those keys decide the model's leaderboard family, and skipping this makes it show as ❓ Other on the website.

3e. Multi-file support code (optional)

If the wrapper needs helper modules (preprocessors, vendored upstream code, large internal classes), put them in a private subfolder of packages/tabarena/src/tabarena/models/{ModelKey}/:

  • _internal/ — for hand-written helpers (preprocessors, internal classes, adapters)
  • _vendor/ — only for code copied verbatim from an upstream project; keep the original layout/license alongside

Both subfolders need their own empty __init__.py. Import them from model.py via absolute paths, e.g. from tabarena.models.{ModelKey}._internal.preprocessing import Preprocessor.

3f. Test config (no per-model test file)

There is no per-model test file. tests/tabarena/models/test_all_models.py is parametrized over the model registry, so it fits the new model automatically once its info.py is discoverable. It skips on ImportError (optional dep missing) and for GPU-only models without CUDA.

Only touch tests/tabarena/models/smoke_configs.py if the model's toy fit needs a speed-up: add one entry to SMOKE_OVERRIDES, keyed by the model's MethodMetadata.method (the registry key), e.g. "{ModelName}": ModelSmokeTest({"max_epochs": 1}), or ModelSmokeTest(problem_types=("regression",)) for a regression-only model. If the model fits fine with default hyperparameters on all problem types, add nothing. A wrapper that declares its cheapness knobs as the cheap_hyperparameters ClassVar (Step 3g) needs no entry either: smoke_for merges them into the smoke config and the warm-up dummy fit uses the same dict.

3g. Warm-up (untimed environment warm-up): decide, don't skip

TabArena runs an untimed warm-up before every timed fit (AbstractExecModel.warmup_fn, dispatched per model class by tabarena.models.warmup.warmup_model_cls), so one-time per-environment costs (library imports, JIT and kernel compilation, the CUDA context, pretrained weights) stay out of time_train_s / time_infer_s and the fit time limit. The layers are additive and run in this order for every model; a new model declares only what the generic layers cannot see:

  1. An optional warmup classmethod, warmup(cls, *, problem_type=None, num_cpus=None, num_gpus=None, hyperparameters=None, **kwargs) -> None, for work the declarative layers cannot express (a library's own kernel pre-compilation, allocator settings that must precede the CUDA context). References: chimeraboost, mitra_v2.
  2. Torch import plus CUDA context for AbstractTorchModel subclasses (automatic).
  3. warmup_modules: ClassVar[tuple[str, ...]] = ("yourlib", "yourlib.submodule"): the modules _fit and _predict import lazily, merged over the MRO. A "torch" entry also creates the CUDA context for a torch-backed model on plain AbstractModel.
  4. The ag_key map in warmup.py for AutoGluon built-ins (LightGBM, CatBoost, XGBoost, ...).
  5. A dummy fit and predict on a small synthetic dataset (make_synthetic_frames, fixed seed, no task data), which triggers the lazy imports, library caches and kernel loads a first fit pays. For a class that declares shared_weights (below) it also builds and registers the network the timed fit reuses. Opt-outs on the class: warmup_dummy_fit = False, warmup_dummy_fit_kwargs (n_rows, n_features, n_categorical, time_limit) and cheap_hyperparameters (cheapness knobs such as n_estimators=1, merged over the config; never an input of the network loader, since the dummy fit must build the network the real fit looks up).
SituationDeclare
sklearn-like or lightweightNothing; the dummy fit covers it.
Torch model on AbstractTorchModelwarmup_modules with the heavy extra imports (transformers, ...).
Torch-backed model on AbstractModelwarmup_modules = ("torch", "yourlib"). References: modernnca, xrfm, tabstar.
Library JIT-compiles kernels (numba, JAX, custom CUDA)A warmup classmethod calling the library's pre-compile entry point when one exists (chimeraboost, whose warmup() needs chimeraboost>=0.14.1). Ask the user for the entry point and minimum version when the docs do not say.
Foundation model with pretrained weightsA shared_weights declaration (below) plus warmup_modules for the library.
The library itself blocks the warm-up or the sharing (import-time global side effect, lazy load at the first predict, load inside __init__, leaked global state)A developer fix (below): the smallest workaround, headed Developer fix, with the upstream ask written down.

Fairness contract (the tabarena/models/warmup.py module docstring is the reference): data-independent work only; never task data, never task- or data-specific state carried into the fit, never a global random number generator advanced.

When the library gets in the way: the developer-fix pattern. The warm-up and the shared weights assume a library that imports without side effects, loads its network in one separable call and does not touch process-global state. Several do not, and the fix belongs upstream. Do not skip the warm-up for such a model (warmup_modules = (), warmup_dummy_fit = False were the interim answer for iLTM and are gone); write the smallest workaround in the wrapper instead, and mark it so it can be found and removed later. The four shapes seen so far, with the in-tree reference for each:

The library...Developer fixReference
flips a global torch flag or the root logger when imported or fitted (TF32, cuDNN, logging.basicConfig)pin the flag explicitly in _fit right after the import (one policy for every fit, whatever imported first) and save/restore the rest in a context manager around the fitiltm/model.py (_isolate_iltm_global_state, allow_tf32 = False)
builds its network only at the first predictbuild it at the end of _fit, so the load is shared and the timed predict starts warmnori/model.py (the eager _get_predictor() call)
loads the checkpoint inside __init__, with no _load_model and no network= argumenta load_network(...) function in <model>/_estimators.py replicating the loading half of the constructor (declared as the shared_weights loader), plus a constructor replica when the constructor cannot take a prebuilt network, guarded against library bumpsexaone_tabular/_estimators.py (TabDPT had one until layer6ai-labs/TabDPT-inference#79 merged)
declares no logger until __init__ ran, breaks after an unpickle, or has another bug the fit path hitsthe narrowest patch at the call site, idempotent, applied where the code path enters the libraryiltm/model.py (_ensure_iltm_logger_patched)

The header is the contract: the module docstring, function docstring or comment opens with Developer fix (or Developer fix: inline), then says what the library does, what it should offer instead, the version the fix was written against, and the upstream issue or PR once filed. Before writing one, ask the library's maintainers for the seam (a separable loader, a network= argument, an import without side effects); file the issue or PR, link it from the header, and remove the fix when the library ships it. TabDPT is the precedent: its constructor replica linked layer6ai-labs/TabDPT-inference#79, and once that merged tabdpt/model.py declared the upstream _load_model like TabICL, with the extra moved to the first release that carries it (tabdpt>=1.3.1). grep -rn "Developer fix" packages/tabarena/src/tabarena/models lists what is outstanding. Report every developer fix you add in Step 8.

Foundation models share one network. A bagged fit of an in-context model would otherwise build the same frozen network once per fold child and once more for the refit child, inside the timed fit. Every such library builds its network inside one call its fit makes: a method on the estimator (tabpfn _initialize_model_variables, tabicl _load_model), a module function (causilo load_pretrained_model) or a classmethod (Tab2D.from_pretrained). The wrapper names that call and the inputs that decide which network it builds; AutoGluon's AbstractTorchModel does the rest:

from autogluon.core.models.abstract import SharedWeights

class {ClassName}Model(AbstractTorchModel):
    shared_weights: ClassVar[SharedWeights] = SharedWeights(
        loader=("somelib:SomeClassifier._load_model", "somelib:SomeRegressor._load_model"),
        key=("checkpoint_version", "model_path"),  # the loader's inputs that pick the network
        disabled_by=("kv_cache",),                  # inputs under which the library writes into it
    )
    cheap_hyperparameters: ClassVar[dict] = {"n_estimators": 1}

loader is "package.module:Class.method" for a method the estimator's fit calls, or "package.module:function" (also Class.classmethod) for a call that returns the network; several when the library has one class per task. key names the estimator attributes (method) or call arguments (function) that decide which network is built; the device is part of the key by itself, from the loader's device input or, for a loader without one, from self.device when _fit sets it before the fit, else the fit's num_gpus. disabled_by lists inputs (names, or a predicate over the inputs) under which the library writes into or casts the network, so that fit builds its own. copy_per_fit=True is for a fine-tuning model: every fit trains a deep copy and pickles it. Nothing in _fit changes: the library's own code runs on the first call in the process and registers what it produced; every later call with the same inputs and device reuses it, the fitted model pickles without the network and takes it back on load, set_device swaps registry entries, and get_info()["shared_weights"] records it. Sharing is off for a fit with ag_args_fit={"share_pretrained_weights": False}, per class through model_class_settings={"<ag_key>": {"share_weights": False}}, and for a configuration named in disabled_by; fold fitting and refit_folds stay the wrapper's ordinary ensemble arguments.

The library contract. What a library has to offer is one loader call whose inputs are estimator parameters (checkpoint name or path, device, and any flag that shapes the network) and whose effect is to produce the network, as a return value or as attributes it sets. Check this before integrating: if the library loads inside __init__ (EXAONE Tabular; TabDPT until its maintainers merged a separable _load_model), ask the maintainers for a separable loader or a network= constructor argument first, and only then fall back to an adapter in <model>/_estimators.py: a load_network(...) function that replicates the loading half of the constructor (the declared loader) and, when the constructor cannot take a prebuilt network, a constructor replica. Mark the module as a developer fix in its docstring, say what the library should offer instead, and guard the replica against library bumps (exaone_tabular/_estimators.py). A library whose __setstate__ re-runs its loader (tabicl) reloads shared for free; one that pickles its network (tabpfn) is covered by the generic weightless pickle. A loader the constructor calls (tabdpt _load_model) works too, as long as the keyed attributes are set before the call; an input the constructor resolves after it (tabdpt compile) cannot go in disabled_by, so the wrapper overrides _shares_weights for it. Wrappers that do not care about sharing simply declare nothing (TabDPT v1.1 and TabDPT-Turbo, whose releases have no separable loader, TabPFN-Wide, SAP-RPT-OSS, iLTM whose library caches for itself).

references/model_patterns.md under "Shared pretrained weights" has the field-by-field reference; tabicl/model.py is the plain in-tree case, causilo/model.py a function loader without a device input, mitra_v2/model.py a fine-tuning model, tabdpt/model.py a loader the constructor calls and exaone_tabular/ the adapter.

Tests: tests/tabarena/models/test_shared_weights_models.py discovers every registered class that declares shared_weights and checks the declaration (well-formed, the loader resolves when the library is installed, the keyed inputs are the loader's parameters, the named disablers switch sharing off); nothing to register per wrapper. pytest -m models tests/tabarena/models/test_shared_weights_models.py -k <Model> fits the real library twice, asserts one network, a weightless pickle that reloads and predicts the same, and equal predictions with sharing switched off.

Inference side. The exec model persists the fitted model in memory around the predict timer (AGWrapper.persist, through persist_inference.persist_for_inference) and calls an optional instance method prepare_for_inference(self) -> None on every persisted object, bagged children included. Its contract is written once, in AbstractExecModel.pre_predict (exec_models/base.py); refer to it instead of restating it. Data-dependent first-predict work (torch.compile on the real batch, kernel dispatch for the real shapes) is measured by design.

Shortened here. Read the whole file on GitHub.

Signals

GitHub stars
336
Forks
79
Last commit
Oct 2026

Ahel review

  • K1binfo
    installs-packages

Automated review, not a security audit. Ruleset v1+k2.

Advanced
Item type
skill
Key
add-model
Source
github.com/autogluon/tabarena