Hasty Briefsbeta

Bilingual

TabPFN and TabICL vs. tuned XGBoost: the model that doesn't train won 14/14

4 hours ago
  • TabPFN and TabICL, tabular foundation models that do not train on new data, outperformed tuned XGBoost on most of 14 datasets from the Grinsztajn benchmark, with the advantage holding up to 32,000 rows.
  • TabICL won 12 of 14 datasets in accuracy and all 14 in AUC, while TabPFN (v2.2.1) won 11 in accuracy and 13 in AUC against tuned XGBoost.
  • The key benefit is not large accuracy gains (mean ~0.01) but eliminating hyperparameter tuning, reducing compute time from minutes to under a second for most datasets.
  • The models struggle with wide tables (e.g., Bioresponse with 419 columns), where TabPFN degraded significantly, and TabICL was slower but still best.
  • A major finding: the most-cited model TabPFN (v8.3.0) now requires account registration and license acceptance to download, while TabICL remains openly available under a BSD license.
  • Technical issues included PyTorch silently using CPU instead of GPU due to version mismatch, and OpenML API outages, highlighting reproducibility challenges.
  • Limitations: only binary classification tested, 3,000-row ceiling for main table, hyperparameter search limited to 25 combinations, and categoricals integer-encoded uniformly.