MedBench: Benchmarking Machine Learning Methods for Medical Tabular Prediction
| dc.contributor.author | Höök, Vidar | |
| dc.contributor.author | Stöckel, Alva | |
| dc.contributor.department | Chalmers tekniska högskola / Institutionen för data och informationsteknik | sv |
| dc.contributor.department | Chalmers University of Technology / Department of Computer Science and Engineering | en |
| dc.contributor.examiner | Schneider, Gerardo | |
| dc.contributor.supervisor | Johansson, Fredrik | |
| dc.date.accessioned | 2026-09-23T14:22:00Z | |
| dc.date.issued | 2026 | |
| dc.date.submitted | ||
| dc.description.abstract | Tabular data is central to medical machine learning, yet recent tabular foundation models have mostly been evaluated on general-purpose benchmarks. We evaluate whether such models transfer to medical tabular prediction by constructing a benchmark from the ADNI, MIMIC-IV, and NHANES datasets, covering prognostic and diagnostic tasks across binary classification, multiclass classification, and regression. TabPFN, TabICL, and MITRAarecomparedwith classical baselines using modality specific metrics, aggregate rank analyses, and bootstrap uncertainty estimates. TabPFN and TabICL achieved the strongest overall performance on small- and medium-sized tasks, consistently ranking above classical baselines and tuned XG Boost. MITRA was competitive but less consistent. The advantage of TabPFN and TabICL was clearest on medium-sized tasks, while smaller tasks showed higher rank uncertainty due to limited sample sizes and fewer task instances. However, the largest tasks could not be evaluated with the full TabPFN and TabICL configurations under the available hardware constraints. On reduced-size variants of these tasks, TabPFN and TabICL remained competitive, but the results did not provide clear evidence that this advantage persisted as the number of samples increased. The results indicate that tabular foundation models can transfer effectively to heterogeneous medical prediction tasks, especially in small- and medium-data regimes. However, their advantages are not uniform across dataset scales, and aggregate benchmark results can overstate practical usefulness if failed runs and subsampling are not reported. Future medical tabular benchmarks should therefore evaluate both predictive performance and computational feasibility. This study used data obtained from the Alzheimer’s Disease Neuroimaging Initiative (ADNI) database. | |
| dc.identifier.coursecode | DATX05 | |
| dc.identifier.uri | https://hdl.handle.net/20.500.12380/312531 | |
| dc.language.iso | eng | |
| dc.setspec.uppsok | Technology | |
| dc.subject | Computer, science, computer science, engineering, project, thesis | |
| dc.title | MedBench: Benchmarking Machine Learning Methods for Medical Tabular Prediction | |
| dc.type.degree | Examensarbete för masterexamen | sv |
| dc.type.degree | Master's Thesis | en |
| dc.type.uppsok | H | |
| local.programme | Data science and AI (MPDSC), MSc |
