NVIDIA Kumo Tabular is a publicly downloadable single-table model worth piloting, but current evidence does not justify replacing CatBoost, LightGBM, or TabPFN in production.
The practical verdict: pilot Kumo Tabular, do not switch production traffic yet
Kumo Tabular is worth testing when your data already lives in one table, labels are available, and you want to measure whether in-context prediction saves modeling work. Treat strict latency, CPU-only serving, managed hosting, or multi-table data as test gates rather than assuming Kumo Tabular will satisfy them.
Official weights and API documentation are available, but public Kumo-specific comparisons remain limited. Validate it on your own holdout.
What NVIDIA actually released
NVIDIA’s Kumo-Tabular model card describes a pretrained model for classification and regression. The official KumoTabular API documentation shows an in-context workflow: labeled rows provide context, and query rows receive predictions without fitting new weights for each dataset.
The public package exposes small, medium, and large variants. NVIDIA’s structured-data model catalog lists roughly 27 million to 216 million parameters across those variants. The model card lists the OpenMDW-1.1 license and provides a Python installation and inference path. Review the license text for commercial deployment and redistribution conditions rather than treating “public weights” as unrestricted use.
Do not confuse this release with KumoRFM. Kumo Tabular targets a single table. NVIDIA’s Kumo Relational overview describes KumoRFM as a separate model for relational data. The Kumo Tabular documentation explicitly marks related-table support as unavailable.
| Question | Kumo Tabular answer |
|---|---|
| Main tasks | Classification and regression |
| Input shape | One table with context and query rows |
| Per-dataset training | Not required for the in-context prediction path |
| Related tables | Not supported in the public KumoTabular API |
| Weights | Public on Hugging Face |
| License shown in model card | OpenMDW-1.1 |
| Hosted inference | No Hugging Face Inference Provider is listed |
The official model card and API page reviewed for this article do not publish a simple maximum-row or maximum-feature figure. They do expose model configuration and task restrictions, so row count, feature count, context/query composition, memory use, and accepted class count should be recorded during the pilot rather than inferred from another tabular model.
The constraints that decide whether it fits
Kumo Tabular’s fit depends on three operational constraints:
- Data shape: Kumo Tabular does not consume related tables natively. If predictive fields come from customers, orders, products, and support events, define and validate the flattening or aggregation policy before comparing models.
- Inference scale: Public documentation describes model variants and task support, but it does not provide an independently validated table of production throughput, GPU memory, or cost per prediction. Measure how context size changes latency and memory.
- Deployment access: The weights and Python interface are public, but that is different from a managed API with an SLA. Teams needing regional routing, quota guarantees, or vendor-hosted operations should make those requirements explicit pilot gates.
How Kumo Tabular compares with TabPFN and trained baselines
There is no reliable public apples-to-apples result in the reviewed sources proving Kumo Tabular beats current TabPFN, CatBoost, LightGBM, or XGBoost. The cited third-party benchmarks inform pilot design, not Kumo performance claims.
| Workload | First comparison to run | Evaluation purpose |
|---|---|---|
| Clean single table, small or medium labeled sample | Kumo Tabular versus TabPFN | Compare table-native in-context prediction on the same split |
| Categorical-heavy table | Kumo Tabular versus CatBoost | Use CatBoost as the trained categorical-data baseline |
| Stable, repeated batch scoring | Kumo Tabular versus LightGBM or XGBoost | Compare an in-context model with trained-model serving baselines |
| Signal spread across linked tables | Flat-table baseline versus relational approach | Measure whether the chosen aggregation loses useful structure |
| Rapid schema exploration | Kumo Tabular pilot | Test whether avoiding a task-specific fit loop saves useful time |
The AnoFox comparison measured several tabular foundation models on one 8-core CPU and found that they beat the tested scikit-learn models by about 2–7% on five small datasets, while runtime could be much higher. In one churn test, Mitra’s warm inference took 19.3 seconds versus 0.035 seconds for logistic regression.
A separate AIMultiple benchmark found TabFM won 15 of 19 datasets but required roughly 40 times the compute of TabPFN 3 or TabICLv2. Its reported full TabFM run cost about $27 on B200 GPUs, compared with about $0.65 for TabPFN 3. Again, these figures define measurements to collect; they are not Kumo results.
A defensible first test
Use a fixed holdout and make Kumo Tabular earn a place beside a trained baseline.
- Freeze one time-aware or stratified train/test split before trying models. Keep the test labels unavailable during prediction.
- Validate the table schema: target column, missing values, categorical columns, duplicate entities, and any fields created after the prediction timestamp.
- Run Kumo Tabular on the raw supported table, recording model variant, context and query rows, feature count, accepted class count, batch size, hardware, cold-start time, warm latency, and peak memory.
- Run CatBoost or LightGBM on the same rows and target. Record preprocessing time separately from fit and prediction time.
- Compare a task-appropriate metric: AUROC and calibration for binary classification, macro-F1 for imbalanced multiclass work, and MAE or RMSE for regression.
- Repeat with a smaller and larger context sample. If quality is stable but latency grows sharply, the model may be better for exploration than serving.
- Predefine three pass criteria: minimum metric lift, maximum p95 latency, and maximum infrastructure cost per batch. Advance only if Kumo Tabular clears all three.
Do not use a vendor benchmark as a substitute for leakage checks. “No feature engineering” does not mean “no data validation.”