<?xml version="1.0" encoding="UTF-8"?>
<rss  xmlns:atom="http://www.w3.org/2005/Atom" 
      xmlns:media="http://search.yahoo.com/mrss/" 
      xmlns:content="http://purl.org/rss/1.0/modules/content/" 
      xmlns:dc="http://purl.org/dc/elements/1.1/" 
      version="2.0">
<channel>
<title>Josh&#39;s AI/ML Experiments</title>
<link>https://CJosh88.github.io/ai-portfolio/</link>
<atom:link href="https://CJosh88.github.io/ai-portfolio/index.xml" rel="self" type="application/rss+xml"/>
<description>Small, reproducible experiments with new ML/AI releases.</description>
<generator>quarto-1.10.18</generator>
<lastBuildDate>Fri, 02 Oct 2026 00:00:00 GMT</lastBuildDate>
<item>
  <title>Can a pretrained model beat XGBoost without training?</title>
  <dc:creator>Josh Chettiar</dc:creator>
  <link>https://CJosh88.github.io/ai-portfolio/experiments/2026-10-kumo-tabular/</link>
  <description><![CDATA[ 





<p><a href="https://colab.research.google.com/github/CJosh88/ML/blob/main/experiments/2026-10-kumo-tabular/notebook.ipynb"><img src="https://colab.research.google.com/assets/colab-badge.svg" class="img-fluid" alt="Open in Colab"></a> <a href="../../experiments/2026-10-kumo-tabular/notebook.html">Full notebook with outputs</a> · <a href="https://github.com/CJosh88/ML/tree/main/experiments/2026-10-kumo-tabular">Code and data on GitHub</a></p>
<div class="callout callout-style-simple callout-note callout-titled">
<div class="callout-header d-flex align-content-center">
<div class="callout-icon-container">
<i class="callout-icon"></i>
</div>
<div class="callout-title-container flex-fill">
<span class="screen-reader-only">Note</span>TL;DR
</div>
</div>
<div class="callout-body-container callout-body">
<ul>
<li><strong>With zero tuning, NVIDIA Kumo Tabular beat default XGBoost and LightGBM on 25 of 25 train/test splits</strong> across five OpenML datasets, by +0.034 ROC-AUC on average.</li>
<li><strong>The advantage is largest with little data.</strong> On adult it was +0.05 ROC-AUC at 100 training rows and +0.01 at 10,000.</li>
<li><strong>On Bank Marketing it edged out my hand-tuned 2020 models.</strong> It had the best PR-AUC both with the leaky <code>duration</code> feature (0.566 vs 0.542 for tuned XGBoost) and without it (0.373 vs 0.349).</li>
<li><strong>Rebuilding the 2020 models exposed a bug in my 2020 notebook.</strong> Its random forest scores look inflated by train/test leakage.</li>
<li><strong>The cost is speed.</strong> On a 10k-row table Kumo large took about 80 s on a T4 GPU, while XGBoost took 0.3 s on CPU. Kumo small was nearly as accurate at about 11 s.</li>
</ul>
</div>
</div>
<section id="the-question" class="level2">
<h2 class="anchored" data-anchor-id="the-question">The question</h2>
<p>Gradient-boosted trees have been the default answer for tabular data for a decade. Tabular foundation models take a different route, and NVIDIA’s <a href="https://huggingface.co/blog/nvidia/kumo-tabular">Kumo Tabular</a> reports first place on the TabArena and TALENT benchmarks.</p>
<p>I wanted to check two things on a free Colab GPU:</p>
<ol type="1">
<li>With zero tuning, does it beat default XGBoost and LightGBM on small and medium classification tables?</li>
<li>How does it compare with the hand-tuned models from a <a href="https://github.com/CJosh88/scriptz/blob/master/Bank%20Marketing%20ML.ipynb">Bank Marketing notebook I wrote in 2020</a>?</li>
</ol>
</section>
<section id="how-in-context-learning-works" class="level2">
<h2 class="anchored" data-anchor-id="how-in-context-learning-works">How in-context learning works</h2>
<p>A gradient-boosted tree model is <strong>trained</strong>. You give it your training rows, it builds hundreds of trees to fit them, and you then use those trees to predict new rows. Every dataset gets its own freshly fitted model, usually with hyperparameter tuning on top.</p>
<p>Kumo Tabular is <strong>never trained on your data</strong>. Its weights were fixed once, by NVIDIA, after pretraining on tens of millions of <em>synthetic</em> tables. The tables were generated from random causal graphs, so the model has seen a huge variety of “features cause a target” patterns, but no real dataset. To make a prediction, you pass it two things in a single forward pass:</p>
<ol type="1">
<li>your labelled training rows (the <strong>context</strong>), and</li>
<li>the unlabelled rows you want predictions for (the <strong>queries</strong>).</li>
</ol>
<p>Inside the transformer, each query row attends to the context rows and their labels and infers the relationship on the fly, much as an LLM picks up a task from worked examples in its prompt. Nothing is fitted, nothing is tuned and no weights change. The “learning” happens inside one forward pass, which is why it’s called <strong>in-context learning</strong> (ICL).</p>
<div class="cell" data-layout-align="default">
<div class="cell-output-display">
<div>
<p></p><figure class="figure"><p></p>
<div>
<pre class="mermaid mermaid-js">flowchart TB
  subgraph T["XGBoost / LightGBM: train, then predict"]
    direction LR
    A["Training rows&lt;br/&gt;+ labels"] --&gt; B["Fit hundreds of trees&lt;br/&gt;(+ tuning)"]
    B --&gt; C["New model for&lt;br/&gt;this dataset"]
    D["Test rows"] --&gt; C
    C --&gt; E["Predictions"]
  end
  subgraph K["Kumo Tabular: in-context learning"]
    direction LR
    F["Training rows&lt;br/&gt;+ labels"] --&gt; H["Frozen pretrained transformer&lt;br/&gt;one forward pass, no weight updates"]
    G["Test rows"] --&gt; H
    H --&gt; I["Predictions"]
  end
  T ~~~ K
</pre>
</div>
<p></p><figcaption> Same training rows, two ways of using them.</figcaption> </figure><p></p>
</div>
</div>
</div>
<table class="caption-top table">
<colgroup>
<col style="width: 33%">
<col style="width: 33%">
<col style="width: 33%">
</colgroup>
<thead>
<tr class="header">
<th></th>
<th>XGBoost / LightGBM</th>
<th>Kumo Tabular (ICL)</th>
</tr>
</thead>
<tbody>
<tr class="odd">
<td>Uses your training rows</td>
<td>to fit a new model</td>
<td>as context in a forward pass</td>
</tr>
<tr class="even">
<td>Weights change?</td>
<td>yes, built from scratch per dataset</td>
<td>no, frozen after pretraining</td>
</tr>
<tr class="odd">
<td>Hyperparameter tuning</td>
<td>usually needed</td>
<td>none</td>
</tr>
<tr class="even">
<td>What it learned beforehand</td>
<td>nothing</td>
<td>priors from ~35–137M synthetic tables</td>
</tr>
<tr class="odd">
<td>Compute</td>
<td>cheap CPU</td>
<td>GPU; cost grows with context size</td>
</tr>
<tr class="even">
<td>Practical limit</td>
<td>none really</td>
<td>context size (here 10k rows per ensemble member)</td>
</tr>
</tbody>
</table>
<p>So this isn’t zero-shot. The model needs labelled examples, and in every comparison here it gets exactly the same training rows as XGBoost and LightGBM. It’s zero-<em>training</em>. The pretrained prior is what lets it do well from very few rows.</p>
</section>
<section id="setup" class="level2">
<h2 class="anchored" data-anchor-id="setup">Setup</h2>
<table class="caption-top table">
<colgroup>
<col style="width: 33%">
<col style="width: 33%">
<col style="width: 33%">
</colgroup>
<thead>
<tr class="header">
<th>Part</th>
<th>What</th>
<th>Models</th>
</tr>
</thead>
<tbody>
<tr class="odd">
<td>A</td>
<td>5 OpenML-CC18 binary datasets × 5 stratified 75/25 splits (train capped at 10k rows, test at 5k)</td>
<td>XGBoost (default), LightGBM (default), TabICLv2, Kumo Tabular small (28M params) and large (215M)</td>
</tr>
<tr class="even">
<td>Learning curve</td>
<td>adult and Bank Marketing (no <code>duration</code>), 100 to 10k training rows, 3 seeds</td>
<td>XGBoost, LightGBM, Kumo Tabular large</td>
</tr>
<tr class="odd">
<td>B</td>
<td>UCI Bank Marketing on the 2020 split and encoding, with and without the leaky <code>duration</code> feature</td>
<td>The four 2020 models, rebuilt with their 2020 tuned hyperparameters, plus all Part A models</td>
</tr>
<tr class="even">
<td>Controls</td>
<td>Leak check for the 2020 numbers; shuffled-label check for the in-context models</td>
<td>as above</td>
</tr>
</tbody>
</table>
<ul>
<li><strong>No tuning anywhere</strong> except the 2020 models, which keep their 2020 tuned settings. The GBDTs use library defaults with native categorical handling.</li>
<li><strong>Kumo settings:</strong> 8 ensemble members, fp16, at most 10k context rows per member.</li>
<li><strong>Hardware:</strong> Colab T4 GPU for the in-context models, Colab’s 2 vCPUs for the trees.</li>
<li><strong>Reproducibility:</strong> a full rerun reproduced every Part A ROC-AUC to 4 decimal places.</li>
</ul>
<div class="callout callout-style-default callout-tip callout-titled">
<div class="callout-header d-flex align-content-center collapsed" data-bs-toggle="collapse" data-bs-target=".callout-2-contents" aria-controls="callout-2" aria-expanded="false" aria-label="Toggle callout">
<div class="callout-icon-container">
<i class="callout-icon"></i>
</div>
<div class="callout-title-container flex-fill">
<span class="screen-reader-only">Tip</span>Getting it installed
</div>
<div class="callout-btn-toggle d-inline-block border-0 py-1 ps-1 pe-0 float-end"><i class="callout-toggle"></i></div>
</div>
<div id="callout-2" class="callout-2-contents callout-collapse collapse">
<div class="callout-body-container callout-body">
<p>Kumo Tabular ships in NVIDIA’s <a href="https://github.com/NVIDIA/structured-data-models"><code>structured-data-models</code></a> library, which is still alpha. Two traps:</p>
<ul>
<li><strong><code>pip install structured-data-models</code> installs an empty <code>0.0.0a0</code> placeholder</strong>, even though the Hugging Face model card suggests it. Install from GitHub instead. The notebook pins commit <code>ce95710</code>.</li>
<li><strong>The code snippet in NVIDIA’s blog post has typos</strong> (<code>pd.load_csv</code>, <code>drop_column</code>). Follow the repo’s own examples instead.</li>
</ul>
<p>The weights download from Hugging Face without a token, and everything runs on a free Colab T4.</p>
</div>
</div>
</div>
</section>
<section id="results" class="level2">
<h2 class="anchored" data-anchor-id="results">Results</h2>
<section id="part-a-five-openml-datasets-5-splits-each" class="level3">
<h3 class="anchored" data-anchor-id="part-a-five-openml-datasets-5-splits-each">Part A: five OpenML datasets, 5 splits each</h3>
<p>Mean ROC-AUC over 5 splits:</p>
<table class="caption-top table">
<colgroup>
<col style="width: 14%">
<col style="width: 14%">
<col style="width: 14%">
<col style="width: 14%">
<col style="width: 14%">
<col style="width: 14%">
<col style="width: 14%">
</colgroup>
<thead>
<tr class="header">
<th>Dataset</th>
<th>Rows</th>
<th>XGBoost</th>
<th>LightGBM</th>
<th>TabICLv2</th>
<th>Kumo small</th>
<th>Kumo large</th>
</tr>
</thead>
<tbody>
<tr class="odd">
<td>credit-g</td>
<td>1,000</td>
<td>0.767</td>
<td>0.770</td>
<td>0.802</td>
<td>0.805</td>
<td><strong>0.809</strong></td>
</tr>
<tr class="even">
<td>diabetes</td>
<td>768</td>
<td>0.794</td>
<td>0.803</td>
<td>0.838</td>
<td>0.839</td>
<td><strong>0.843</strong></td>
</tr>
<tr class="odd">
<td>kc1</td>
<td>2,109</td>
<td>0.809</td>
<td>0.793</td>
<td>0.867</td>
<td>0.866</td>
<td><strong>0.872</strong></td>
</tr>
<tr class="even">
<td>phoneme</td>
<td>5,404</td>
<td>0.952</td>
<td>0.951</td>
<td>0.974</td>
<td>0.976</td>
<td><strong>0.979</strong></td>
</tr>
<tr class="odd">
<td>adult</td>
<td>48,842</td>
<td>0.914</td>
<td>0.922</td>
<td>0.923</td>
<td>0.929</td>
<td><strong>0.930</strong></td>
</tr>
<tr class="even">
<td><strong>Mean rank</strong></td>
<td></td>
<td>4.52</td>
<td>4.44</td>
<td>2.72</td>
<td>2.24</td>
<td><strong>1.08</strong></td>
</tr>
</tbody>
</table>
<ul>
<li><strong>Kumo large</strong> beat the better of the two GBDTs on <strong>25 of 25</strong> splits, by +0.034 ROC-AUC on average. Its smallest margin was +0.007.</li>
<li><strong>Kumo small</strong> also won 25/25 (+0.030). <strong>TabICLv2</strong> won 24/25, with one tie (+0.028).</li>
<li><strong>The gap was largest on the smaller tables</strong> (kc1 +0.06, diabetes +0.04, credit-g +0.04) and smallest on adult (+0.008).</li>
<li><strong>The in-context models were also better calibrated.</strong> Default XGBoost’s log loss was about 1.5× Kumo large’s on credit-g (0.70 vs 0.47) and on diabetes (0.75 vs 0.46).</li>
</ul>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://CJosh88.github.io/ai-portfolio/experiments/2026-10-kumo-tabular/figures/auc_by_dataset.png" class="img-fluid figure-img"></p>
<figcaption>Mean ROC-AUC per dataset; error bars are the standard deviation over 5 splits.</figcaption>
</figure>
</div>
</section>
<section id="learning-curve-the-less-data-the-bigger-the-win" class="level3">
<h3 class="anchored" data-anchor-id="learning-curve-the-less-data-the-bigger-the-win">Learning curve: the less data, the bigger the win</h3>
<p>Mean ROC-AUC over 3 seeds:</p>
<table class="caption-top table">
<thead>
<tr class="header">
<th>Dataset</th>
<th>Training rows</th>
<th>XGBoost</th>
<th>LightGBM</th>
<th>Kumo large</th>
</tr>
</thead>
<tbody>
<tr class="odd">
<td>adult</td>
<td>100</td>
<td>0.807</td>
<td>0.740</td>
<td><strong>0.859</strong></td>
</tr>
<tr class="even">
<td>adult</td>
<td>1,000</td>
<td>0.877</td>
<td>0.882</td>
<td><strong>0.912</strong></td>
</tr>
<tr class="odd">
<td>adult</td>
<td>10,000</td>
<td>0.911</td>
<td>0.917</td>
<td><strong>0.927</strong></td>
</tr>
<tr class="even">
<td>Bank Marketing (no <code>duration</code>)</td>
<td>100</td>
<td>0.588</td>
<td>0.584</td>
<td><strong>0.648</strong></td>
</tr>
<tr class="odd">
<td>Bank Marketing (no <code>duration</code>)</td>
<td>1,000</td>
<td>0.653</td>
<td>0.662</td>
<td><strong>0.720</strong></td>
</tr>
<tr class="even">
<td>Bank Marketing (no <code>duration</code>)</td>
<td>10,000</td>
<td>0.693</td>
<td>0.719</td>
<td><strong>0.740</strong></td>
</tr>
</tbody>
</table>
<ul>
<li><strong>Kumo large led at every training size, and its lead shrank as data grew.</strong></li>
<li><strong>It needed far fewer rows to match the trees.</strong> On Bank Marketing, Kumo with 1,000 rows matched the best GBDT with 10,000 rows (0.720 vs 0.719). On adult it needed 3,000 rows (0.922 vs 0.917).</li>
</ul>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://CJosh88.github.io/ai-portfolio/experiments/2026-10-kumo-tabular/figures/learning_curve.png" class="img-fluid figure-img"></p>
<figcaption>ROC-AUC vs training-set size (log scale), mean ± standard deviation over 3 seeds.</figcaption>
</figure>
</div>
</section>
<section id="part-b-a-rematch-with-my-2020-bank-marketing-models" class="level3">
<h3 class="anchored" data-anchor-id="part-b-a-rematch-with-my-2020-bank-marketing-models">Part B: a rematch with my 2020 Bank Marketing models</h3>
<p>Test set: 9,043 rows, 11.7% positive.</p>
<p>The 2020 notebook kept <code>duration</code> (call length). UCI warns against using it for a realistic model, because you only know it once the call is over. So Part B runs twice: once with the 2020 feature set (like for like) and once without <code>duration</code> (the honest task).</p>
<table class="caption-top table">
<colgroup>
<col style="width: 20%">
<col style="width: 20%">
<col style="width: 20%">
<col style="width: 20%">
<col style="width: 20%">
</colgroup>
<thead>
<tr class="header">
<th>Model</th>
<th>2020 features: ROC-AUC</th>
<th>2020 features: PR-AUC</th>
<th>No <code>duration</code>: ROC-AUC</th>
<th>No <code>duration</code>: PR-AUC</th>
</tr>
</thead>
<tbody>
<tr class="odd">
<td>Logistic regression (2020)</td>
<td>0.860</td>
<td>0.466</td>
<td>0.690</td>
<td>0.268</td>
</tr>
<tr class="even">
<td>Logistic regression + SMOTE (2020)</td>
<td>0.864</td>
<td>0.468</td>
<td>0.691</td>
<td>0.269</td>
</tr>
<tr class="odd">
<td>Random forest + SMOTE (2020)</td>
<td>0.885</td>
<td>0.492</td>
<td>0.715</td>
<td>0.310</td>
</tr>
<tr class="even">
<td>XGBoost, tuned and cost-sensitive (2020)</td>
<td>0.895</td>
<td>0.542</td>
<td>0.733</td>
<td>0.349</td>
</tr>
<tr class="odd">
<td>XGBoost (default)</td>
<td>0.888</td>
<td>0.508</td>
<td>0.722</td>
<td>0.331</td>
</tr>
<tr class="even">
<td>LightGBM (default)</td>
<td>0.897</td>
<td>0.540</td>
<td><strong>0.740</strong></td>
<td>0.351</td>
</tr>
<tr class="odd">
<td>TabICLv2</td>
<td>0.898</td>
<td>0.551</td>
<td>0.737</td>
<td>0.344</td>
</tr>
<tr class="even">
<td>Kumo small</td>
<td>0.900</td>
<td>0.564</td>
<td>0.737</td>
<td>0.356</td>
</tr>
<tr class="odd">
<td>Kumo large</td>
<td><strong>0.901</strong></td>
<td><strong>0.566</strong></td>
<td>0.739</td>
<td><strong>0.373</strong></td>
</tr>
</tbody>
</table>
<ul>
<li><strong>Kumo large had the highest PR-AUC in both variants, with zero tuning.</strong> It beat the 2020 XGBoost, which took about 8 minutes of random search, by +0.024 in both variants. Without <code>duration</code>, ROC-AUC is effectively a tie with default LightGBM (0.739 vs 0.740).</li>
<li><strong><code>duration</code> was doing most of the work.</strong> Removing it cost every model 0.16–0.17 ROC-AUC, so much of my 2020 notebook’s apparent skill came from a feature that isn’t available at prediction time.</li>
<li><strong>At a 0.5 threshold the picture differs.</strong> Kumo large with <code>duration</code> gets precision 0.64, recall 0.40 and F1 0.49. The 2020 random forest (SMOTE) gets 0.44, 0.71 and 0.54. Nothing tells the in-context models about the 88/12 class imbalance, so they predict “yes” conservatively, whereas SMOTE and <code>scale_pos_weight</code> deliberately trade precision for recall. Choosing the threshold from the precision-recall curve closes that gap.</li>
</ul>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://CJosh88.github.io/ai-portfolio/experiments/2026-10-kumo-tabular/figures/bank_pr_curves.png" class="img-fluid figure-img"></p>
<figcaption>Precision-recall curves on the Bank Marketing test set. The dotted line is the 11.7% base rate.</figcaption>
</figure>
</div>
</section>
<section id="my-2020-numbers-dont-reproduce-and-why" class="level3">
<h3 class="anchored" data-anchor-id="my-2020-numbers-dont-reproduce-and-why">My 2020 numbers don’t reproduce, and why</h3>
<p>Rebuilt with the same split, encoding and tuned hyperparameters, the 2020 logistic regressions match what the 2020 notebook printed. The random forest and XGBoost don’t:</p>
<table class="caption-top table">
<thead>
<tr class="header">
<th>F1 at threshold 0.5</th>
<th>Printed in 2020</th>
<th>Rebuilt today</th>
</tr>
</thead>
<tbody>
<tr class="odd">
<td>Logistic regression</td>
<td>0.30</td>
<td>0.33</td>
</tr>
<tr class="even">
<td>Logistic regression + SMOTE</td>
<td>0.47</td>
<td>0.47</td>
</tr>
<tr class="odd">
<td>Random forest + SMOTE</td>
<td><strong>0.65</strong></td>
<td><strong>0.54</strong></td>
</tr>
<tr class="even">
<td>XGBoost</td>
<td><strong>0.64</strong></td>
<td><strong>0.51</strong></td>
</tr>
</tbody>
</table>
<p>The 2020 notebook shuffled the data <strong>without a seed</strong> and evaluated models <strong>loaded from <code>.pkl</code> files</strong>. If those pickles were trained in an earlier session, that session had a different shuffle. A large share of the 2020 “test” rows would then have been training rows for the pickled models.</p>
<p>Simulating exactly that (train on one shuffle’s split, score on another shuffle’s test split):</p>
<table class="caption-top table">
<thead>
<tr class="header">
<th>Model</th>
<th>F1, clean test</th>
<th>F1, reshuffled test</th>
<th>Printed in 2020</th>
</tr>
</thead>
<tbody>
<tr class="odd">
<td>Logistic regression</td>
<td>0.32</td>
<td>0.33</td>
<td>0.30</td>
</tr>
<tr class="even">
<td>Logistic regression + SMOTE</td>
<td>0.49</td>
<td>0.49</td>
<td>0.47</td>
</tr>
<tr class="odd">
<td>Random forest + SMOTE</td>
<td>0.54</td>
<td><strong>0.68</strong></td>
<td><strong>0.65</strong></td>
</tr>
<tr class="even">
<td>XGBoost</td>
<td>0.53</td>
<td>0.55</td>
<td>0.64</td>
</tr>
</tbody>
</table>
<ul>
<li><strong>Random forest:</strong> the leak reproduces its 2020 numbers almost exactly (precision 0.56 vs 0.54, recall 0.86 vs 0.84). A depth-16 forest memorises its training rows, so leaked rows inflate its score.</li>
<li><strong>Logistic regression:</strong> unaffected, as you’d expect from a model that can’t memorise.</li>
<li><strong>XGBoost:</strong> the leak explains only a small part of its gap; the rest is unexplained.</li>
</ul>
<p>My 2020 random forest was most likely never as good as the notebook reported.</p>
</section>
<section id="did-kumo-just-memorise-these-datasets" class="level3">
<h3 class="anchored" data-anchor-id="did-kumo-just-memorise-these-datasets">Did Kumo just memorise these datasets?</h3>
<p>All six datasets are well-known public benchmarks. Kumo Tabular is reportedly pretrained only on synthetic tables, but it’s fair to ask whether its edge comes from having effectively “seen” these datasets before.</p>
<p>To test that, each model got the same training rows with <strong>randomly permuted labels</strong>: same class balance, no real relationship to the features. A model that learns from its context should fall to chance (ROC-AUC ≈ 0.5). One that recognises the dataset would keep scoring well.</p>
<table class="caption-top table">
<thead>
<tr class="header">
<th>Model</th>
<th>Mean ROC-AUC, true labels</th>
<th>Mean ROC-AUC, shuffled labels</th>
</tr>
</thead>
<tbody>
<tr class="odd">
<td>Kumo large</td>
<td>0.858</td>
<td>0.475</td>
</tr>
<tr class="even">
<td>Kumo small</td>
<td>0.855</td>
<td>0.446</td>
</tr>
<tr class="odd">
<td>TabICLv2</td>
<td>0.849</td>
<td>0.455</td>
</tr>
<tr class="even">
<td>XGBoost (default)</td>
<td>0.814</td>
<td>0.474</td>
</tr>
</tbody>
</table>
<ul>
<li><strong>Every model collapsed to chance or below</strong>, and the in-context models collapsed as fully as XGBoost, which can’t have memorised anything.</li>
<li><strong>Individual runs are noisy</strong> (0.25–0.66), XGBoost included.</li>
<li><strong>No sign of memorisation.</strong> The performance comes from the labelled rows in context.</li>
</ul>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://CJosh88.github.io/ai-portfolio/experiments/2026-10-kumo-tabular/figures/shuffled_labels.png" class="img-fluid figure-img"></p>
<figcaption>Shuffled-label control: wide pale bars use the true labels, narrow dark bars the shuffled labels (mean of 3 permutations).</figcaption>
</figure>
</div>
</section>
<section id="what-it-costs" class="level3">
<h3 class="anchored" data-anchor-id="what-it-costs">What it costs</h3>
<p>Median fit-plus-predict time per split:</p>
<table class="caption-top table">
<colgroup>
<col style="width: 16%">
<col style="width: 16%">
<col style="width: 16%">
<col style="width: 16%">
<col style="width: 16%">
<col style="width: 16%">
</colgroup>
<thead>
<tr class="header">
<th>Dataset</th>
<th>XGBoost (CPU)</th>
<th>LightGBM (CPU)</th>
<th>TabICLv2 (T4)</th>
<th>Kumo small (T4)</th>
<th>Kumo large (T4)</th>
</tr>
</thead>
<tbody>
<tr class="odd">
<td>credit-g (750 train / 250 test)</td>
<td>0.09 s</td>
<td>0.06 s</td>
<td>0.39 s</td>
<td>0.55 s</td>
<td>1.2 s</td>
</tr>
<tr class="even">
<td>phoneme (4k / 1.4k)</td>
<td>0.12 s</td>
<td>0.12 s</td>
<td>0.97 s</td>
<td>0.85 s</td>
<td>7.3 s</td>
</tr>
<tr class="odd">
<td>adult (10k / 5k)</td>
<td>0.31 s</td>
<td>0.20 s</td>
<td>11.0 s</td>
<td>10.6 s</td>
<td>80.7 s</td>
</tr>
</tbody>
</table>
<p>On adult, Kumo large was about 250× slower than XGBoost and Kumo small about 35× slower. Kumo small gave up only 0.004 ROC-AUC on average for an 8× speed-up over large, so on this evidence it’s the better default.</p>
</section>
</section>
<section id="what-i-take-from-this" class="level2">
<h2 class="anchored" data-anchor-id="what-i-take-from-this">What I take from this</h2>
<ol type="1">
<li><strong>For small and medium tables, a pretrained in-context model is now a strong default.</strong> With no tuning, Kumo Tabular beat untuned GBDTs on every split I tested, often by a margin bigger than the split-to-split noise. On Bank Marketing it also beat my tuned 2020 XGBoost on PR-AUC.</li>
<li><strong>The less data you have, the bigger the win.</strong> Kumo needed roughly 3–10× fewer rows than the trees to reach the same ROC-AUC. As data grows towards the context limit, the GBDTs close in.</li>
<li><strong>Kumo small is the practical choice.</strong> It came within 0.004 ROC-AUC of large at about one-eighth of the cost.</li>
<li><strong>Thresholds still need thought.</strong> The in-context models aren’t told about class imbalance, so on imbalanced problems pick the operating point from the precision-recall curve rather than using 0.5.</li>
<li><strong>The rematch was most useful for checking my old work.</strong> It surfaced a leaky feature (<code>duration</code>) and a train/test leak that likely inflated my 2020 random forest results.</li>
</ol>
</section>
<section id="limitations" class="level2">
<h2 class="anchored" data-anchor-id="limitations">Limitations</h2>
<ul>
<li><strong>Small benchmark.</strong> Five binary datasets and one extra case study. NVIDIA’s claims rest on TabArena’s 51 datasets. These results are a sanity check on public data, not a ranking.</li>
<li><strong>Untuned baselines.</strong> Except for the 2020 models, the GBDTs ran at library defaults with no early stopping. Tuned GBDTs would close some of the gap.</li>
<li><strong>Context cap.</strong> Kumo saw at most 10k training rows per ensemble member, and used 8 members rather than the 16 in NVIDIA’s benchmark.</li>
<li><strong>Public datasets.</strong> The shuffled-label control rules out memorisation of these datasets, but it can’t rule out that the model’s design was tuned on public benchmarks like these. A dataset published after the model, or private data, would be a stronger test.</li>
<li><strong>Timing compares a GPU with a 2-vCPU CPU.</strong> It reflects practical cost on free Colab, not algorithmic efficiency.</li>
<li><strong>Fine-tuning not tested.</strong> Every Kumo number here is from the frozen pretrained model.</li>
</ul>
</section>
<section id="data-and-reproducibility" class="level2">
<h2 class="anchored" data-anchor-id="data-and-reproducibility">Data and reproducibility</h2>
<table class="caption-top table">
<thead>
<tr class="header">
<th>Dataset</th>
<th>Source</th>
<th>Licence</th>
</tr>
</thead>
<tbody>
<tr class="odd">
<td>credit-g, diabetes, kc1, phoneme, adult</td>
<td>OpenML <a href="https://www.openml.org/d/31">31</a>, <a href="https://www.openml.org/d/37">37</a>, <a href="https://www.openml.org/d/1067">1067</a>, <a href="https://www.openml.org/d/1489">1489</a>, <a href="https://www.openml.org/d/1590">1590</a></td>
<td>Public</td>
</tr>
<tr class="even">
<td>Bank Marketing (<code>bank-full.csv</code>)</td>
<td><a href="https://archive.ics.uci.edu/dataset/222/bank+marketing">UCI 222</a></td>
<td>CC BY 4.0</td>
</tr>
</tbody>
</table>
<p>Bank Marketing: S. Moro, P. Cortez and P. Rita (2014), <em>A Data-Driven Approach to Predict the Success of Bank Telemarketing</em>, Decision Support Systems 62:22–31.</p>
<p>Everything, including setup, every table, every chart and the controls, is in the <a href="../../experiments/2026-10-kumo-tabular/notebook.html">full notebook</a>. It runs top to bottom on a free Colab T4 with the <a href="https://colab.research.google.com/github/CJosh88/ML/blob/main/experiments/2026-10-kumo-tabular/notebook.ipynb">Open in Colab</a> button. The raw results CSVs are <a href="https://github.com/CJosh88/ML/tree/main/experiments/2026-10-kumo-tabular">on GitHub</a>.</p>


</section>

 ]]></description>
  <category>tabular</category>
  <category>foundation models</category>
  <category>in-context learning</category>
  <category>xgboost</category>
  <category>benchmarking</category>
  <guid>https://CJosh88.github.io/ai-portfolio/experiments/2026-10-kumo-tabular/</guid>
  <pubDate>Fri, 02 Oct 2026 00:00:00 GMT</pubDate>
  <media:content url="https://CJosh88.github.io/ai-portfolio/experiments/2026-10-kumo-tabular/figures/auc_by_dataset.png" medium="image" type="image/png" height="59" width="144"/>
</item>
</channel>
</rss>
