Empirical substitution models of SARS-CoV-2 protein evolution for phylogenetic inference
Abstract
Probabilistic phylogenetic inference requires substitution models of molecular evolution, which should be as realistic as possible to yield accurate predictions. At the protein level, empirical substitution models are well established in phylogenetics. However, the number of available empirical substitution models remains limited, and only a few focus on rapidly evolving viruses. Indeed, no substitution models have yet been developed specifically for SARS-CoV-2 proteins, despite their potential to improve evolutionary predictions. Here, we present a set of empirical substitution models for SARS-CoV-2 proteins, including the main protease, papain-like protease, spike protein, and the complete proteome. We estimated these models using maximum likelihood from thousands of protein sequences and validated with independent datasets. Overall, the estimated empirical substitution models outperformed currently available empirical models in terms of phylogenetic likelihood and revealed distinct evolutionary patterns among SARS-CoV-2 proteins. Additionally, using forward-in-time evolutionary simulations, we found that SARS-CoV-2 proteins evolved under these empirical substitution models generated protein variants with folding stabilities more realistic than those produced under other empirical substitution models. We conclude that evolutionary inferences based on protein-specific substitution models are generally more accurate and biologically realistic than those based on generalist models, and recommend the empirical substitution models presented here for evolutionary predictions of SARS-CoV-2 proteins.
Related articles
Related articles are currently not available for this article.