File size: 1,716 Bytes
55712ab
 
 
 
 
 
 
 
 
 
 
4f0fae5
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
---
license: mit
library_name: pytorch
pipeline_tag: video-classification
tags:
  - video-anomaly-detection
  - weakly-supervised-learning
  - transfer-learning
  - clip
---

# Model Card for PATT-Net

## Model Details
- **Architecture:** ProtectedTransferVAD (dual-branch NativeTemporalBackbone)
- **Base Feature Extractor:** OpenAI CLIP ViT-L/14 (768-dim)
- **Temporal Modeling:** Bi-directional LSTM with Masked Temporal Dilated Convolutions
- **Intended Use:** Video Anomaly Detection (VAD) research and controlled ablation studies.

## Training Data
- **Pre-training (Temporal Branch):** PreVAD dataset (35,279 videos, distinct from XD-Violence).
- **Target Fine-tuning:** UCF-Crime dataset.

## Evaluation and Validation
- **Metric:** Frame-level AUC and AP, evaluated via center-crop (stride-16) visual-only features.
- **Selection Criteria:** Best validation video AUC.
- **Scientific Limits:** The P-B contrast represents the whole system effect (branch mechanism + pretraining). Within this matched frozen-branch design, P-R isolates PreVAD rather than random source initialization. The P-R differences are small and not statistically significant across five seeds (AUC exact two-sided sign-flip `p = 0.1875`).

## Ethical Considerations and Licenses
- The repository code is provided under the MIT License. Third-party dataset and model terms continue to apply to data-derived artifacts.
- Users must comply with the licenses of the underlying third-party datasets (UCF-Crime and PreVAD) and the OpenAI CLIP feature extractor. The repository authors are not responsible for the misuse of these models in surveillance or real-world classification scenarios without appropriate safeguards.