← Work

National Taiwan University, CSIE

2021

TACnet: Weakly Supervised Video Anomaly Detection

Localized anomalous video segments from video-level labels by combining temporal attention, clustering, and a custom entropy loss in an end-to-end weakly supervised model.

Tech stack
PyTorch · C3D · Weakly supervised learning · Multiple-instance learning · Temporal attention · K-means clustering
Operating context
UCF-Crime and ShanghaiTech · Video-level labels to segment-level scores · 84.75% frame-level AUC on UCF-Crime
Research workflow

Training and evaluation flow

TACnet training and evaluation flowVariable-length videos are normalized into 32 temporal segments with 16 sampled frames each. During training, TACnet receives one video-level normal or anomalous label per video and produces 32 segment anomaly scores. During evaluation, those scores are expanded to the original frame timeline and compared with frame-level labels to calculate AUC.INPUT PREPARATIONInput videovariable lengthView the end-to-end weakly supervised video pipeline32 segmentsequal temporal spansView the end-to-end weakly supervised video pipeline16 framessampled per segmentOpen the full model architecture in the paperTACnetmodel architecturepaper ↗View attention-guided segment localization32 segment scorestemporal anomalypredictionsTRAINING SUPERVISIONVideo-level labelone per videonormal · anomalousView entropy loss and empirical evaluationJoint objectivesprediction · clusterranking · sparsityentropyFrame timelinescores restored to lengthFrame-level labelsanomaly intervalsevaluation onlyFrame-level AUC84.75% UCF-Crime93.89% ShanghaiTech
Training used one video-level label for each video. Frame-level labels entered only after inference to evaluate localization. The paper contains the full model architecture.

Research contributions

01

End-to-end weakly supervised video pipeline

Designed and implemented a pipeline that trained on one normal or anomalous label per video while producing anomaly scores for individual temporal segments.

Problem
Videos had different lengths, anomalous events occupied only a small fraction of their frames, and frame-level training labels were too expensive to assume.
Implementation
Sampled every video into 32 segments of 16 frames, used a Sports-1M-pretrained C3D backbone, and added a 1D temporal convolution to capture longer-range context while reducing the feature dimension.
Reasoning
A uniform segment sequence made variable-length videos batchable, while retaining the video backbone inside the training path avoided treating precomputed clip averages as the final temporal representation.
Outcome
Created one end-to-end representation for video-level classification and segment-level localization across UCF-Crime and ShanghaiTech.
02

Attention-guided clustering for segment localization

Combined gated temporal attention with two-cluster K-means to infer which segments inside an anomalous video represented the event.

Problem
A video-level label identifies an anomalous bag but does not reveal which of its segments contain the anomaly.
Implementation
Learned attention weights for the bag prediction, concatenated each segment feature with its attention weight, clustered the segments with cosine distance, and treated the cluster with the higher mean segment score as the anomaly candidate.
Reasoning
Attention exposed each segment's contribution to the bag decision, while clustering converted that weak signal into a within-video separation without requiring frame labels. Cosine distance was more stable than Euclidean distance in the experiments.
Outcome
Removing the clustering path produced the largest top-down ablation drop, 1.97 percentage points on UCF-Crime.
03

Entropy loss and empirical evaluation

Proposed an entropy smoothness loss and evaluated the complete model through quantitative comparisons, qualitative traces, ablations, and clustering-distance experiments.

Problem
Penalizing differences between adjacent segment scores can blur the boundary between normal and anomalous behavior.
Implementation
Quantized segment scores into 20 intervals and minimized their distribution entropy, then analyzed component removals, attention weights, prediction curves, and feature clusters.
Reasoning
Concentrating predictions into fewer score ranges encouraged stable, separable outputs without assuming neighboring segments must always receive similar scores.
Outcome
The model achieved 84.75% frame-level AUC on UCF-Crime and 93.89% on ShanghaiTech in the paper's evaluation.

This page reports the 2021 research artifact and its original evaluation. The laboratory environment and trained artifacts are no longer available, so the public repository is preserved as frozen thesis code rather than presented as a currently reproducible production ML system.