National Taiwan University, CSIE
2021TACnet: Weakly Supervised Video Anomaly Detection
Localized anomalous video segments from video-level labels by combining temporal attention, clustering, and a custom entropy loss in an end-to-end weakly supervised model.
- Tech stack
- PyTorch · C3D · Weakly supervised learning · Multiple-instance learning · Temporal attention · K-means clustering
- Operating context
- UCF-Crime and ShanghaiTech · Video-level labels to segment-level scores · 84.75% frame-level AUC on UCF-Crime
- Research artifacts
- Paper↗Code↗Slides↗Thesis record↗
Training and evaluation flow
Research contributions
End-to-end weakly supervised video pipeline
Designed and implemented a pipeline that trained on one normal or anomalous label per video while producing anomaly scores for individual temporal segments.
- Problem
- Videos had different lengths, anomalous events occupied only a small fraction of their frames, and frame-level training labels were too expensive to assume.
- Implementation
- Sampled every video into 32 segments of 16 frames, used a Sports-1M-pretrained C3D backbone, and added a 1D temporal convolution to capture longer-range context while reducing the feature dimension.
- Reasoning
- A uniform segment sequence made variable-length videos batchable, while retaining the video backbone inside the training path avoided treating precomputed clip averages as the final temporal representation.
- Outcome
- Created one end-to-end representation for video-level classification and segment-level localization across UCF-Crime and ShanghaiTech.
Attention-guided clustering for segment localization
Combined gated temporal attention with two-cluster K-means to infer which segments inside an anomalous video represented the event.
- Problem
- A video-level label identifies an anomalous bag but does not reveal which of its segments contain the anomaly.
- Implementation
- Learned attention weights for the bag prediction, concatenated each segment feature with its attention weight, clustered the segments with cosine distance, and treated the cluster with the higher mean segment score as the anomaly candidate.
- Reasoning
- Attention exposed each segment's contribution to the bag decision, while clustering converted that weak signal into a within-video separation without requiring frame labels. Cosine distance was more stable than Euclidean distance in the experiments.
- Outcome
- Removing the clustering path produced the largest top-down ablation drop, 1.97 percentage points on UCF-Crime.
Entropy loss and empirical evaluation
Proposed an entropy smoothness loss and evaluated the complete model through quantitative comparisons, qualitative traces, ablations, and clustering-distance experiments.
- Problem
- Penalizing differences between adjacent segment scores can blur the boundary between normal and anomalous behavior.
- Implementation
- Quantized segment scores into 20 intervals and minimized their distribution entropy, then analyzed component removals, attention weights, prediction curves, and feature clusters.
- Reasoning
- Concentrating predictions into fewer score ranges encouraged stable, separable outputs without assuming neighboring segments must always receive similar scores.
- Outcome
- The model achieved 84.75% frame-level AUC on UCF-Crime and 93.89% on ShanghaiTech in the paper's evaluation.
This page reports the 2021 research artifact and its original evaluation. The laboratory environment and trained artifacts are no longer available, so the public repository is preserved as frozen thesis code rather than presented as a currently reproducible production ML system.