An Explainable Deep Learning Strategy Towards Rapid Bacterial DNA Classification for Pathogen Identification

This article has 0 evaluations Published on
Read the full article Related papers
This article on Sciety

Abstract

Accurate and scalable classification of genomic sequences is a critical task in bioinformatics, particularly for identifying bacterial species from raw DNA data. Traditional alignment-based techniques often suffer from high computational overhead and limited adaptability to short, noisy sequences. In this paper, we present an interpretable deep learning framework built on frequencynormalized k-mer embeddings to classify bacterial genomes with high accuracy. The proposed pipeline is alignment-free and processes raw DNA sequences into uniform chunks, which are then transformed into high-dimensional k-mer frequency vectors. These vectors serve as inputs to a deep neural network architecture enriched with batch normalization, dropout regularization, and ReLU activations to ensure generalization and robustness. Experimental results across multiple bacterial genomes Bacillus subtilis, E. coli, Staphylococcus aureus, and Pseudomonas aeruginosa demonstrate that our model achieves a classification accuracy exceeding 98%, with strong performance across precision, recall, and F1-score metrics. The frequency-based embedding scheme also keeps the features interpretable of learned features, making the system biologically meaningful and transparent. Our work highlights the effectiveness of deep learning models paired with statistical representations of genomic sequences, offering a scalable and interpretable solution for bacterial classification in large genomic datasets.

Related articles

Related articles are currently not available for this article.