Pre-Transformer Models for Longevity Science and Deep Ageing Clocks: An Empirical Analysis

This article has 0 evaluations Published on
Read the full article Related papers
This article on Sciety

Abstract

The transformer now dominates artificial intelligence, and its diffusion into computational biology raises a concrete question for the ageing field: on the data that ageing research actually has, do attention-based models outperform the pre-transformer methods—penalized regression, kernel machines, tree ensembles, and the pre-2017 deep networks (multilayer perceptrons, convolutional and recurrent nets, autoencoders)— that built the ageing-clock paradigm? Comparative claims have rested largely on the structure of the data rather than on head-to-head experiments. We supply the experiment. We benchmark eight model classes for biological-age prediction across six public cohorts spanning five modalities: blood DNA-methylation (GSE40279), blood biochemistry (NHANES), whole-blood transcriptome (GTEx), gut microbiome (American Gut), wearable accelerometry (NHANES), and brain MRI (Cam-CAN). Every model is tuned and evaluated under one nested cross-validation protocol on identical splits, and we score not only accuracy but data efficiency, compute, interpretability, cross-cohort transfer, and—critically—the biological validity of each model’s age-acceleration residual against all-cause mortality. Pre-transformer models win outright on five of the six cohorts and essentially tie on the sixth. On tabular molecular omics, penalized regression and gradient boosting are best or tied-best (methylation MAE 3.4 yr, transcriptome 8.1 yr); on genuinely structured modalities the modality-native pre-transformer deep net wins (accelerometry ConvLSTM MAE 12.7 yr, brain-MRI 3D-CNN 4.7 yr). From-scratch transformers never win and often overfit; a pretrained transformer edges ahead only on the single large-𝑛 cohort (𝑛=34,000), by a practically negligible 0.10 yr, and only past a training size of ≈23,000 labelled samples—larger than almost any human ageing cohort—while costing three to four orders of magnitude more compute. Aggregate accuracy is, in fact, a near-tie; the pre-transformer case is decided elsewhere. First, minimizing chronological-age error does not maximize biological signal: the model with the lowest MAE (the pretrained transformer) carries the weakest mortality association, and both transformers rank at the bottom for risk stratification (HR ≈1.13 per SD versus ≈1.28 for elastic net), an empirical echo of “too much of a good thing.” Second, penalized regression transfers across cohorts best of all (+1.2 yr degradation), while the from-scratch and non-pretrained deep models degrade more than twice as much (+2.8–4.0 yr). For the low-sample, high-dimensional, interpretability-sensitive, and regulator-facing regime that defines ageing research, the pre-transformer toolkit is not merely adequate but the rational default; transformers are a targeted addition, not a replacement.

Related articles

Related articles are currently not available for this article.