Nvidia's Kumo Tabular Reads a Spreadsheet Cold and Beats Every Tuned Model Doing It
A new open AI model from Nvidia makes predictions on business data tables without any training on your data, no setup, and no specialist required, and it just topped every major leaderboard.

Key points
- Nvidia's Kumo Tabular tops four tabular-data benchmarks: TabArena, BeyondArena, TALENT and ScoringBench.
- The model makes predictions from a spreadsheet in a single pass with zero training, zero tuning and zero feature engineering required from the user.
- Kumo Tabular was trained entirely on artificial data, not on any real customer or business records, and comes in three sizes ranging from 28 million to 215 million parameters.
- The model is free to download from Hugging Face and is licensed for commercial use under OpenMDW-1.1.
- Kumo Tabular runs 17 times faster than its closest benchmark rival.
Most of the AI that quietly runs business life lives not in a chatbot but in a spreadsheet. Customer churn predictions, loan defaults, demand forecasts: all of it comes from tables of rows and columns. For two decades, those predictions required a data scientist to hand-build a model from scratch each time a new question came up. Nvidia just released something that skips that entire process.
Kumo Tabular, announced this month and published on Hugging Face, is a foundation model, a general-purpose AI pretrained so thoroughly it can be applied to new tasks without further training. Hand it a table where some rows already have labels (say, "churned" or "stayed") and it immediately predicts labels for the unlabelled rows. No feature engineering (the laborious manual process of deciding which columns to feed a model), no hyperparameter tuning (fiddling with a model's internal settings to make it learn properly), no waiting. One pass and you have answers.
That's a serious shift. The standard approach, gradient-boosted trees (a family of prediction algorithms that has dominated business AI for years), works well but demands a full training run for every new problem. Kumo Tabular borrows the idea that made large language models so useful: in-context learning, where the model reads labelled examples as its prompt and generalises from them on the spot, without updating its internal weights.
How does it actually work?
Kumo Tabular treats a table the way a language model treats a sentence: it reads the structure, not just the numbers.
The model uses three layers of attention, a mechanism that lets AI weigh which pieces of information matter most relative to each other. Column attention checks where a value sits within its own column's range (is 42 typical here, or an outlier?). Row attention works across a single row to catch how different columns interact. A final pass relates the labelled context rows to the unlabelled query rows to produce a prediction. The model code is fully open on GitHub.
One clever fix addresses a real-world headache: attention tends to get blurry when tables get very large. Kumo Tabular scales the sharpness of its attention automatically as table size grows, by adjusting a temperature value tied to the logarithm of the number of rows, keeping predictions reliable whether you hand it a modest table or up to 60,000 rows.
Was any real data used to train it?
No, and that matters. Nvidia trained Kumo Tabular entirely on synthetic tables, generated by a procedural algorithm rather than pulled from any real database. The generator builds random causal graphs (mathematical structures that mimic cause-and-effect relationships in real data) and samples millions of artificial tables from them. Kumo Tabular's small, medium and large variants saw roughly 35 million, 71 million and 137 million such tables during training respectively.
Training on fake-but-realistic data sidesteps the privacy concerns that dog models trained on real business records. We've covered synthetic data's growing role in AI development since our first story on the topic in July, and Kumo Tabular is one of the cleaner examples of the approach paying off in benchmark results.
This is a genuinely useful tool for any organisation sitting on labelled spreadsheet data with no ML team to act on it. The limitation worth flagging: Kumo Tabular handles only numbers and categories. Text columns, image data and timestamps aren't supported yet, which rules it out for plenty of real-world tables. Watch for the training recipe and synthetic data generator, which Nvidia says it will release soon; that's where researchers will find the real building blocks.
Common questions
Do I need a data scientist to use Kumo Tabular?
Not to run basic predictions. You provide a table with some labelled rows, and the model returns predictions for the rest in one step. Knowing whether those predictions are trustworthy still benefits from someone who understands your data.
Is my business data used to train the model?
No. Kumo Tabular was trained entirely on synthetic, artificially generated tables. When you run it on your own data, that data is not fed back into any training process.
Does it cost anything?
The model weights and code are free to download from Hugging Face and GitHub under a commercial-use licence. You do need computing hardware to run it, typically a GPU (a specialised chip built for AI workloads).



