Learning what is normal: context and generation in video anomaly detection
Loading...
Authors
ORCID
https://orcid.org/0000-0001-6823-0388
Other Contributors
Issue Date
Type
Electronic thesis
Thesis
Thesis
Language
en_US
Keywords
Degree
PhD
Alternative Title
Abstract
Video anomaly detection (VAD) aims to automatically identify unusual or suspicious events in video streams, offering critical capabilities for public safety, surveillance, and industrial monitoring. While existing VAD methods can detect conspicuous anomalies using general priors, they often fail to capture the nuanced, context-dependent reasoning exhibited by experienced human operators. This limitation arises from challenges inherent to real-world environments, including time-varying normal behavior, spatially dependent definitions of normalcy, and anomalies that emerge only over extended temporal durations. This dissertation addresses these challenges under the one-class classification (OCC) setting, where only normal examples are available during training—a practical formulation given the rarity and high cost of labeling anomalous events. We develop a series of context-aware anomaly detection methods that explicitly model different dimensions of real-world behavior. First, we introduce temporal context-aware approaches that capture non-stationary normalcy patterns using memory-based modeling and contrastive alignment with temporal descriptors. Second, we propose a spatial context-dependent framework that discovers region-specific normal behavior by clustering object activities, enabling detection of anomalies defined purely by spatial violations. Third, we present a lightweight trajectory-based formulation for long-term anomaly detection, allowing duration-dependent behaviors such as loitering to be identified from arbitrarily long object trajectories. Through experiences with real-world surveillance data, we observe that operational environments often contain too few anomalous events to support meaningful evaluation or weak supervision. Motivated by this limitation, we further propose a generative supervision paradigm that leverages video generation models and multimodal large language models to synthesize plausible anomalous events from normal footage. These generated anomalies are refined using a semi-supervised teacher model and distilled into an efficient student detector, enabling scalable learning in data-scarce and open-world settings. Together, this work advances video anomaly detection by unifying temporal, spatial, and long-term contextual reasoning with scalable supervision strategies, bringing automated systems closer to the adaptive and context-aware understanding demonstrated by human security operators.
Description
May2026
School of Engineering
School of Engineering
Full Citation
Publisher
Rensselaer Polytechnic Institute, Troy, NY
