MIDLPred: Ensemble Learning for Phenotype-Oriented Classification of Male Infertility Candidate Proteins from Primary Amino Acid Sequences

This article has 0 evaluations Published on
Read the full article Related papers
This article on Sciety

Abstract

Background: Male infertility is a complex and multifactorial condition resulting from the interplay of numerous genetic, molecular, physiological, and environmental factors. Advances in reproductive genetics have identified a growing number of genes and proteins involved in spermatogenesis, sperm motility, sperm morphology, chromatin organization, and testicular function. However, linking candidate proteins to clinically relevant infertility phenotypes remains a major challenge in molecular andrology and reproductive medicine. Objectives: The objective of this study was to develop a computational framework capable of identifying phenotype-associated molecular signatures directly from protein sequences and assisting in the prioritization of candidate proteins involved in male infertility. Materials and Methods: A dedicated dataset was generated through a bioinformatics workflow integrating information from the Male Infertility Knowledgebase (MIK), UniProt, and complementary biological resources. The initial dataset comprised 31,774 protein entries associated with male infertility. Following sequence quality assessment, duplicate removal, annotation harmonization, and filtering procedures, 10,678 curated protein sequences were retained. To facilitate biologically meaningful analysis, proteins were grouped into four phenotype-oriented categories: MOTILITY, MORPH, PHYSIO, and OTHER. A one-dimensional convolutional neural network (1D-CNN) combined with an ensemble learning strategy involving ten independently trained models was developed to learn phenotype-related sequence patterns. Model performance was assessed using stratified cross-validation, confusion matrix analysis, receiver operating characteristic (ROC) curves, and standard classification metrics. Results: MIDLPred achieved an overall classification accuracy of 96% and a weighted F1-score of 96% across the four-phenotype categories. Cross-validation analyses demonstrated stable and reproducible performance, while confusion matrix and ROC analyses confirmed the precision of the proposed framework. The model successfully identified sequence-derived patterns associated with distinct infertility phenotypes, suggesting that biologically relevant information is embedded within primary protein sequences. Discussion and conclusion: The results demonstrate the potential of deep learning approaches for exploring the molecular landscape of male infertility through protein sequence analysis. Rather than serving as a standalone diagnostic tool, MIDLPred provides a bioinformatics framework that may assist clinicians, reproductive biologists, geneticists, and andrologists in prioritizing candidate proteins and generating biologically relevant hypotheses for further investigation. These findings highlight the value of integrating artificial intelligence with reproductive genetics and molecular andrology and support future developments incorporating explainable artificial intelligence, protein language models, and multi-omics data to improve biological interpretation and clinical applicability. MIDLPred is publicly available at https://midlpred-amob.onrender.com/

Related articles

Related articles are currently not available for this article.