Large-Model-Assisted Analysis of Adversarial Vulnerability: Separating Evidence-Conflict Detection from Robust Classification

This article has 0 evaluations Published on
Read the full article Related papers
This article on Sciety

Abstract

Adversarial examples reveal a persistent mismatch between human visual invari-ance and the decision boundaries learned by deep neural networks. We study this mismatch as a mechanism problem rather than as a search for a larger backbone or a state-of-the-art robustness number. The proposed workflow uses a large model as an analysis assistant to generate hypotheses about fragile evidence. It then tests these hypotheses with paired clean–adversarial tensors, evidence-switching gates, and white-box attacks. We instantiate the idea on CIFAR-10 with compact ResNet-style backbones, an Aha-ResNet gate, a pairwise ranking variant called AhaV2, and GRPO-style policies that select among base, low-pass, edge, and semantic views. Across tested remote GPU configurations, AhaV2 and GRPO induce larger trigger or semantic-action shifts on adversar-ial inputs than on clean inputs. The best long-run AhaV2 model reaches 79.85% clean accuracy, 37.45% FGSM accuracy, and 6.65% PGD-50 two-restart accuracy. Human-aligned and joint semantic variants increase PGD-50 Aha shifts to 27.80 and 27.40 percentage points, respectively, but do not improve strong white-box accuracy. A PGD-5 plug-and-play study further shows that a 1,495-parameter AhaV2 trigger improves clean and FGSM accuracy on a frozen adversarially trained backbone, while only slightly improving PGD-50 accuracy. On TinyIma-geNet, the proposed Aha lens separates PGD-AT, TRADES, MART, and AWP into robust-boundary improvement and unresolved evidence-conflict detection. The main finding is diagnostic: detecting adversarial conflict is trainable and measurable, but it is separable from possessing a robust decision boundary in the alternative evidence path.

Related articles

Related articles are currently not available for this article.