NVIDIA has released Kumo Tabular, an open-weight foundation model family that does the work enterprise machine learning has handed to gradient-boosted trees for two decades — but without a training run per dataset. The model takes a table containing labeled rows, treats those rows as context, and predicts labels for new rows in a single forward pass. NVIDIA published the weights on Hugging Face on September 29 as part of its Kumo Structured collection.
The point is what disappears. Instead of collecting labels, engineering features, selecting an algorithm, tuning hyperparameters and validating a separate model for every business question, teams hand Kumo Tabular labeled examples and query rows. Classification returns class probabilities; regression returns 999 quantiles, which yield both a point prediction and an uncertainty estimate.
Three sizes ship, from 28 million to 215 million parameters, with classifiers and regressors offered as separate models. The weights are released under the OpenMDW-1.1 license, which permits commercial use, while the GPU-native library that runs them, structured-data-models, is Apache 2.0. There is no hosted endpoint on Hugging Face: the model runs on the operator's own CUDA GPU, with NVIDIA recommending cuDF so that dataframe operations stay on the device.
The architecture is a Transformer built around the structure of a table rather than around a sequence. Column attention learns what a value means within its column, row attention learns how features interact, and a final stage lets query rows attend to the labeled context rows — an approach NVIDIA says is inspired by TabICL and TabPFN. Because the context never looks at the queries, its keys and values are computed once and reused for follow-up predictions. A component called Length-aware Attention Temperature scales attention as tables grow. A single pass covers up to 10 classes, and the library extends that with error-correcting output codes. Context can reach 60,000 rows by 100 columns.
Training is the unusual part: NVIDIA used no real-world enterprise datasets. Kumo Tabular was pretrained entirely on synthetic tables generated from structural causal models, in three stages that grew context from 1,024 rows to 10,240 and then to 60,000. The Small, Medium and Large models saw roughly 35 million, 71 million and 137 million synthetic tables respectively. The generator deliberately injects the imperfections that break models trained on tidy data — missing values, correlated features, outliers, high-cardinality categorical columns, heavy-tailed regression targets, and coarsened labels where otherwise identical rows disagree.
The reported numbers are strong, and they are NVIDIA's own. On TabArena the model places first with an ELO of 1,950, ahead of tuned gradient boosting, AutoGluon and other tabular foundation models; on a single RTX 6000 Pro under NVIDIA's uniform evaluation setup, it runs 17 times faster than LimiX-2. BeyondArena shows an ELO of 1,418 with an Improvability score of 7.78%, also first. On TALENT it takes the top overall ranking, with average ranks of 6.67 for classification accuracy, 3.98 for classification log-loss and 4.22 for regression RMSE. On ScoringBench, which tests predictive distributions, the Large and Medium models place first and second. No independent replication was public when the weights appeared.
The limits are worth reading before deployment. Kumo Tabular accepts numeric and categorical columns only; text, images and timestamps must be converted into features first. A single forward pass is capped at 10 classes unless error-correcting output codes are used. NVIDIA warns that accuracy can degrade on tables far outside the training ranges, or when query rows come from a different distribution than the context rows, and asks users to validate on held-out data. The training recipe and the synthetic-data generators are promised for later, with no date attached.
What this changes is the shape of an AI project rather than the scoreboard. Tabular prediction — churn, default, demand, price — sits inside most enterprise workflows, and its economics have been dominated by labor: the feature pipeline, the tuning sweep, the per-problem model to retrain and monitor. Collapsing that into "supply labeled rows, get predictions" moves the scarce resource from engineering time to data curation and evaluation. It also shrinks the attack surface slightly, because a model that never ingests free text cannot be prompt-injected. The caveat is the familiar one: a vendor leaderboard is a reason to benchmark on your own tables, not evidence that the gap will reproduce there.
Comments (0)
Log in to join the discussion
Log InNo comments yet