CrcBiomeScreen: a reproducible workflow for class imbalance and cross-cohort generalisability in colorectal cancer microbiome prediction
Abstract
Gut microbiome profiles have shown promise for colorectal cancer screening, but differences in preprocessing, class-imbalance handling, modelling choices and validation design can lead to inconsistent estimates of predictive performance and cross-cohort generalisability. To address these challenges, we developed CrcBiomeScreen. This reproducible and modular R/Bioconductor workflow enables users to compare normalisation strategies, taxonomic levels, class-weighting approaches and machine-learning models, and to identify suitable analytical options for their own datasets. The workflow was developed and evaluated using 2,252 samples from a real-world colorectal cancer screening cohort, and its cross-cohort generalisability was assessed using nine independent international microbiome cohorts. Across these analyses, GMPR normalisation improved predictive performance across most modelling configurations, while genus-level profiles supported consistent feature representation across cohorts while retaining biological interpretability. Under the most extreme control-enriched imbalance scenario, class weighting maintained sensitivity to neoplasm cases while improving precision, F1-score, balanced accuracy and area under the receiver operating characteristic curve. Model-specific differences were also observed: XGBoost performed strongly within the primary cohort, whereas Random Forest showed greater stability across more heterogeneous external datasets. CrcBiomeScreen integrates preprocessing, model optimisation, held-out evaluation and external validation within a transparent analysis framework. Rather than prescribing a single universally optimal model, it helps researchers systematically evaluate analytical choices and select appropriate strategies for heterogeneous microbiome datasets.
Related articles
Related articles are currently not available for this article.