PATT-Net / README.md
mtastan's picture
Add PATT-Net model card metadata
55712ab verified
|
Raw History Blame Contribute Delete
1.72 kB
---
license: mit
library_name: pytorch
pipeline_tag: video-classification
tags:
- video-anomaly-detection
- weakly-supervised-learning
- transfer-learning
- clip
---
# Model Card for PATT-Net
## Model Details
- **Architecture:** ProtectedTransferVAD (dual-branch NativeTemporalBackbone)
- **Base Feature Extractor:** OpenAI CLIP ViT-L/14 (768-dim)
- **Temporal Modeling:** Bi-directional LSTM with Masked Temporal Dilated Convolutions
- **Intended Use:** Video Anomaly Detection (VAD) research and controlled ablation studies.
## Training Data
- **Pre-training (Temporal Branch):** PreVAD dataset (35,279 videos, distinct from XD-Violence).
- **Target Fine-tuning:** UCF-Crime dataset.
## Evaluation and Validation
- **Metric:** Frame-level AUC and AP, evaluated via center-crop (stride-16) visual-only features.
- **Selection Criteria:** Best validation video AUC.
- **Scientific Limits:** The P-B contrast represents the whole system effect (branch mechanism + pretraining). Within this matched frozen-branch design, P-R isolates PreVAD rather than random source initialization. The P-R differences are small and not statistically significant across five seeds (AUC exact two-sided sign-flip `p = 0.1875`).
## Ethical Considerations and Licenses
- The repository code is provided under the MIT License. Third-party dataset and model terms continue to apply to data-derived artifacts.
- Users must comply with the licenses of the underlying third-party datasets (UCF-Crime and PreVAD) and the OpenAI CLIP feature extractor. The repository authors are not responsible for the misuse of these models in surveillance or real-world classification scenarios without appropriate safeguards.