Structural and Evolutionary Predictors of VIM-2 Mutational Fitness Revealed by Leakage-Aware, Position-Aware Machine Learning

This article has 0 evaluations Published on
Read the full article Related papers
This article on Sciety

Abstract

Deep mutational scanning (DMS) provides a powerful experimental framework for resolving the functional consequences of protein mutations at scale, but predictive modeling of DMS landscapes is vulnerable to overly optimistic evaluation when variants from the same sequence position are distributed across training and test sets. Here, we developed a leakage-aware, position-aware machine-learning framework to identify generalizable determinants of mutational fitness in the VIM-2 metallo-β-lactamase. We analyzed 5,016 mutations spanning 266 sequence positions and nine experimentally defined phenotypes under different antibiotic concentrations and temperatures. Mutation-level predictors integrated physicochemical remodeling, evolutionary substitution scores, local sequence context, hydrogen-bond capacity, side-chain properties, and structural descriptors, including solvent accessibility and distances to the metal site, active-site pocket, and structural loops. Two complementary architectures were evaluated: a 65-predictor model excluding sequence position and a 66-predictor model incorporating position. Final architecture selection was performed exclusively using pooled position-aware out-of-fold performance from five-fold nested outer-test evaluation, while conventional random-split evaluation was retained as a secondary benchmark. Position-aware predictive performance varied substantially across phenotypes (R² = 0.169–0.506), with the position-inclusive architecture selected for six of nine phenotypes. Structural features consistently represented the dominant predictive feature group across all nine phenotypes, while individual predictors such as solvent accessibility, BLOSUM62, and metal-site proximity provided complementary information. Robustness analysis supported stable outer-fold and leave-one-fold-out performance across all nine phenotypes. Together, these results show that VIM-2 mutational fitness is partially predictable from integrated structural, biochemical, and evolutionary information, while revealing phenotype- and fitness-range-dependent predictive performance and underscoring the importance of position-disjoint validation for assessing transferable genotype–phenotype relationships in DMS datasets.

Related articles

Related articles are currently not available for this article.