Structured Evidence and Vision-Language Models for Interpretable Vision-Only UAV Behavior Analysis
Abstract
Vision-only counter-unmanned aerial vehicle (counter-UAV) systems offer low deployment cost, passive sensing, and compatibility with existing camera infrastructure, but most existing pipelines are limited to detection and tracking. As a result, operators receive target locations but little auditable evidence about the UAV's behavior or the urgency of the required response. This paper presents a four-layer framework for interpretable vision-only UAV behavior analysis. The framework first detects UAVs with a single-class YOLOv8 detector, associates detections with ByteTrack, repairs fragmented trajectories through greedy track merging, and converts each trajectory into structured motion evidence including speed statistics, linearity, scale-change ratio, curvature, and hovering ratio. The evidence is then combined with a compact key-frame mosaic and provided to a vision-language model, which produces behavior labels, response-urgency estimates, and natural-language rationales. Experiments on 20 DUT Anti-UAV sequences produced 52 valid tracks, of which 50 were manually annotated for behavior evaluation. The full multimodal configuration achieved the best primary-match accuracy (0.700) and any-match accuracy (0.760), while an image-only variant achieved only 0.220 primary-match accuracy. The results show that structured trajectory evidence is the dominant information source for UAV behavior recognition, while image evidence provides a smaller but useful gain in geometrically ambiguous cases.
Related articles
Related articles are currently not available for this article.