MedBench: Benchmarking Machine Learning Methods for Medical Tabular Prediction

dc.contributor.authorHöök, Vidar
dc.contributor.authorStöckel, Alva
dc.contributor.departmentChalmers tekniska högskola / Institutionen för data och informationstekniksv
dc.contributor.departmentChalmers University of Technology / Department of Computer Science and Engineeringen
dc.contributor.examinerSchneider, Gerardo
dc.contributor.supervisorJohansson, Fredrik
dc.date.accessioned2026-09-23T14:22:00Z
dc.date.issued2026
dc.date.submitted
dc.description.abstractTabular data is central to medical machine learning, yet recent tabular foundation models have mostly been evaluated on general-purpose benchmarks. We evaluate whether such models transfer to medical tabular prediction by constructing a benchmark from the ADNI, MIMIC-IV, and NHANES datasets, covering prognostic and diagnostic tasks across binary classification, multiclass classification, and regression. TabPFN, TabICL, and MITRAarecomparedwith classical baselines using modality specific metrics, aggregate rank analyses, and bootstrap uncertainty estimates. TabPFN and TabICL achieved the strongest overall performance on small- and medium-sized tasks, consistently ranking above classical baselines and tuned XG Boost. MITRA was competitive but less consistent. The advantage of TabPFN and TabICL was clearest on medium-sized tasks, while smaller tasks showed higher rank uncertainty due to limited sample sizes and fewer task instances. However, the largest tasks could not be evaluated with the full TabPFN and TabICL configurations under the available hardware constraints. On reduced-size variants of these tasks, TabPFN and TabICL remained competitive, but the results did not provide clear evidence that this advantage persisted as the number of samples increased. The results indicate that tabular foundation models can transfer effectively to heterogeneous medical prediction tasks, especially in small- and medium-data regimes. However, their advantages are not uniform across dataset scales, and aggregate benchmark results can overstate practical usefulness if failed runs and subsampling are not reported. Future medical tabular benchmarks should therefore evaluate both predictive performance and computational feasibility. This study used data obtained from the Alzheimer’s Disease Neuroimaging Initiative (ADNI) database.
dc.identifier.coursecodeDATX05
dc.identifier.urihttps://hdl.handle.net/20.500.12380/312531
dc.language.isoeng
dc.setspec.uppsokTechnology
dc.subjectComputer, science, computer science, engineering, project, thesis
dc.titleMedBench: Benchmarking Machine Learning Methods for Medical Tabular Prediction
dc.type.degreeExamensarbete för masterexamensv
dc.type.degreeMaster's Thesisen
dc.type.uppsokH
local.programmeData science and AI (MPDSC), MSc

Ladda ner

Original bundle

Visar 1 - 1 av 1
Hämtar...
Bild (thumbnail)
Namn:
CSE 26-156 AS VH.pdf
Size:
1.95 MB
Format:
Adobe Portable Document Format

License bundle

Visar 1 - 1 av 1
Hämtar...
Bild (thumbnail)
Namn:
license.txt
Size:
2.35 KB
Format:
Item-specific license agreed upon to submission
Description: