Researchers and commenters report that TabPFN and TabICL are evaluated against tuned XGBoost on a common benchmark setup. In the reported measurements, the methods “predict on a table” without training on the specific test table and still achieve better results on the benchmark’s tables.

The claim is based on tests across 14 datasets from the Grinsztajn benchmark, using the same data split and the same time budget for all compared methods. According to the sources, TabPFN and TabICL “win on 14 out of 14” tables versus tuned XGBoost under those conditions. The discussions also reference practical trade-offs, including reported differences in latency and VRAM usage, and suggest circumstances where tuned boosting could still be preferable. While outlets repeat the headline result, they focus less on methodological details beyond the shared benchmark and fairness constraints, and more on the performance-vs-resource implications.