CEDF-CS: Class-Balanced Prototype Condensationfor Resource-Efficient and Leak-Free Intrusion Detection in Industrial IoT and Enterprise Networks

This article has 0 evaluations Published on
Read the full article Related papers
This article on Sciety

Abstract

Modern intrusion-detection benchmarks contain millions of network flows, often dominated by a small number of attack families,making training large deep models prohibitive on industrial–IoT (IIoT) edge devices and biasing classical classifiers towardmajority traffic. Dataset condensation promises a remedy, but naive within-class record averaging—an attractive linear schemeused in recent “hierarchical fusion” frameworks—collapses minority structure and its reported gains appear strongly influencedby evaluation protocols in which condensation sources and evaluation data are not fully independent. We introduce CEDF-CS,a leak-free and prototype-based condensation framework that incorporates three key components: (i) a class-balancedk-means prototype allocation that re-allocates the compression budget toward rare attack classes; (ii) a per-class Gaussianmixture synthesis variant that preserves within-class covariance; and (iii) a CPU-only histogram gradient-boosting classifieralongside RBF-SVM and a shallow neural network. Under a rigorous protocol that splits before condensing, fits the scalerand feature scores on training data only, and evaluates on held-out original records over five seeds, balanced prototypesachieve full-data binary detection performance under a 32× training-set compression ratio: F1 = 0.996±0.003 and balancedaccuracy 0.9996±0.0002 on WUSTL-IIoT-2021 (1.19 M flows), indistinguishable from full-data training. On the eight-classCICIDS2017 benchmark (2.83 M flows), the same scheme substantially exceeds full-data balanced accuracy (0.843 vs. 0.659for HGB and 0.838 vs. 0.659 for NN); on macro-F1 the operationally-equivalent best variant is the unbalanced k-means prototype(E), which exceeds full-data training (0.699 vs. 0.671 for NN), while balanced prototypes (F) trade macro-F1 (0.597) for thelarge balanced-accuracy gain. Rare attack classes that full-data training effectively drops receive significantly improvedrepresentation in the condensed set. We further reproduce the inflated numbers reported by the original (leaky) protocol andmechanistically explain them via an association-amplification analysis.

Related articles

Related articles are currently not available for this article.