--- license: mit library_name: pytorch pipeline_tag: video-classification tags: - video-anomaly-detection - weakly-supervised-learning - transfer-learning - clip --- # Model Card for PATT-Net ## Model Details - **Architecture:** ProtectedTransferVAD (dual-branch NativeTemporalBackbone) - **Base Feature Extractor:** OpenAI CLIP ViT-L/14 (768-dim) - **Temporal Modeling:** Bi-directional LSTM with Masked Temporal Dilated Convolutions - **Intended Use:** Video Anomaly Detection (VAD) research and controlled ablation studies. ## Training Data - **Pre-training (Temporal Branch):** PreVAD dataset (35,279 videos, distinct from XD-Violence). - **Target Fine-tuning:** UCF-Crime dataset. ## Evaluation and Validation - **Metric:** Frame-level AUC and AP, evaluated via center-crop (stride-16) visual-only features. - **Selection Criteria:** Best validation video AUC. - **Scientific Limits:** The P-B contrast represents the whole system effect (branch mechanism + pretraining). Within this matched frozen-branch design, P-R isolates PreVAD rather than random source initialization. The P-R differences are small and not statistically significant across five seeds (AUC exact two-sided sign-flip `p = 0.1875`). ## Ethical Considerations and Licenses - The repository code is provided under the MIT License. Third-party dataset and model terms continue to apply to data-derived artifacts. - Users must comply with the licenses of the underlying third-party datasets (UCF-Crime and PreVAD) and the OpenAI CLIP feature extractor. The repository authors are not responsible for the misuse of these models in surveillance or real-world classification scenarios without appropriate safeguards.