Can Machine Learning Identify Meaningful Patterns in DNA Sequences? A Computational Study of DNA Motifs Associated with Transcription-Factor Binding
Abstract
This independent computational study examines whether machine-learning methods can distinguish CTCF-associated DNA sequences from a controlled background using short nucleotide patterns (k-mers). The analysis draws on a publicly available UniBind dataset for human CTCF in K562 cells, linked to ENCODE experiment ENCSR000EGM and JASPAR motif profile MA0139.1. The dataset comprises 46,846 PWM-derived sequences of 19 nucleotides each. Negative examples were generated through dinucleotide shuffling of the positive sequences, preserving local base-pair composition while removing higher-order sequence structure. Each sequence was represented using normalized 3-mer and 4-mer frequencies, yielding 320 numerical features. Logistic Regression and Random Forest classifiers were trained on an 80/20 held-out split and evaluated using accuracy, precision, recall, F1-score, ROC-AUC, and confusion-matrix analysis. Random Forest achieved the stronger performance (92.53% accuracy; ROC-AUC 0.979) compared with Logistic Regression (86.49% accuracy; ROC-AUC 0.938). These results indicate that short sequence-composition features carry substantial discriminative information relative to a dinucleotide-shuffled background. Because the negative class was derived from the positive sequences themselves, the findings should not be read as a general predictor of CTCF binding across arbitrary genomic DNA. Rather, the project's principal contribution is a transparent, reproducible workflow linking molecular biology, sequence analysis, and applied machine learning — the kind of end-to-end pipeline typically introduced at the undergraduate level.
Related articles
Related articles are currently not available for this article.