File size: 1,716 Bytes
55712ab 4f0fae5 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 | ---
license: mit
library_name: pytorch
pipeline_tag: video-classification
tags:
- video-anomaly-detection
- weakly-supervised-learning
- transfer-learning
- clip
---
# Model Card for PATT-Net
## Model Details
- **Architecture:** ProtectedTransferVAD (dual-branch NativeTemporalBackbone)
- **Base Feature Extractor:** OpenAI CLIP ViT-L/14 (768-dim)
- **Temporal Modeling:** Bi-directional LSTM with Masked Temporal Dilated Convolutions
- **Intended Use:** Video Anomaly Detection (VAD) research and controlled ablation studies.
## Training Data
- **Pre-training (Temporal Branch):** PreVAD dataset (35,279 videos, distinct from XD-Violence).
- **Target Fine-tuning:** UCF-Crime dataset.
## Evaluation and Validation
- **Metric:** Frame-level AUC and AP, evaluated via center-crop (stride-16) visual-only features.
- **Selection Criteria:** Best validation video AUC.
- **Scientific Limits:** The P-B contrast represents the whole system effect (branch mechanism + pretraining). Within this matched frozen-branch design, P-R isolates PreVAD rather than random source initialization. The P-R differences are small and not statistically significant across five seeds (AUC exact two-sided sign-flip `p = 0.1875`).
## Ethical Considerations and Licenses
- The repository code is provided under the MIT License. Third-party dataset and model terms continue to apply to data-derived artifacts.
- Users must comply with the licenses of the underlying third-party datasets (UCF-Crime and PreVAD) and the OpenAI CLIP feature extractor. The repository authors are not responsible for the misuse of these models in surveillance or real-world classification scenarios without appropriate safeguards.
|