Development and validation of pLIN, a permanent lineage-numbering system for tracking antimicrobial resistance plasmids
Abstract
Plasmids are the main vehicle by which antimicrobial resistance genes move between bacteria, yet clinical microbiology still has no way to give a resistance-carrying plasmid a name that persists once it does. Replicon typing tells us a plasmid is, say, IncN; it does not tell us whether the IncN plasmid in this week's outbreak is the same plasmid that caused an outbreak two years ago in another hospital, or a different one that merely happens to share a replicon. We built pLIN (plasmid Lineage Identification Number) to close that gap. pLIN converts each plasmid sequence into a tetranucleotide composition vector, clusters plasmids within their replicon group at six cosine-distance thresholds calibrated against average nucleotide identity, and assigns a permanent six-level code that never changes as the reference database grows. Applied to a curated set of 8,077 complete plasmids spanning 28 replicon groups 20 Gram-negative, 4 Gram-positive, and 4 from the non-fermenters Acinetobacter baumannii and Pseudomonas aeruginosa pLIN resolved 3,073 distinct strain-level lineages, more than doubling the discriminatory power of replicon typing alone (Simpson's diversity index 0.985 versus 0.641). A distance-weighted k-nearest-neighbour classifier assigned replicon group membership with 91.1% accuracy, though this figure is inflated by a handful of very large classes; macro-averaged across all 28 groups, F1 was 0.666, reflecting genuinely weak performance on several low-frequency replicon groups with fewer than 20 training examples. Cosine distance correlated with FastANI identity across 4,970 pairwise comparisons (Spearman ρ = −0.348, P < 10–141), and plasmids sharing a full six-level code had a median ANI of 99.9%. Leave-one-out cross-validation on 57 plasmids drawn from published outbreaks reproduced the correct strain-level code in 94.7% of cases, and pLIN correctly flagged known high-risk lineages when tested retrospectively against 74 plasmids from 26 outbreak studies spanning 13 countries including a 90-member IncN lineage (pLIN 671) in which every member carries blaKPC-2 , and a five-replicon-group hub (pLIN 860) in which 44% of members carry mcr colistin resistance. Screening the full training set with AMRFinderPlus identified 1,635 carbapenemase genes, 1,838 extended-spectrum β-lactamase genes, and 204 mcr genes. Scaling the reference database to 79,305 plasmids resolved 57,886 unique codes in under half an hour on ordinary laboratory hardware, without altering any previously assigned code. We built seven additional modules around this core covering assembly quality, database coverage, recombination, novel-group discovery, evolutionary rate, cluster stability, and mobile element boundaries turning pLIN from a naming scheme into a working surveillance platform that clinical microbiology laboratories can run without dedicated bioinformatics support.
Related articles
Related articles are currently not available for this article.