Model Card for PATT-Net

Model Details

  • Architecture: ProtectedTransferVAD (dual-branch NativeTemporalBackbone)
  • Base Feature Extractor: OpenAI CLIP ViT-L/14 (768-dim)
  • Temporal Modeling: Bi-directional LSTM with Masked Temporal Dilated Convolutions
  • Intended Use: Video Anomaly Detection (VAD) research and controlled ablation studies.

Training Data

  • Pre-training (Temporal Branch): PreVAD dataset (35,279 videos, distinct from XD-Violence).
  • Target Fine-tuning: UCF-Crime dataset.

Evaluation and Validation

  • Metric: Frame-level AUC and AP, evaluated via center-crop (stride-16) visual-only features.
  • Selection Criteria: Best validation video AUC.
  • Scientific Limits: The P-B contrast represents the whole system effect (branch mechanism + pretraining). Within this matched frozen-branch design, P-R isolates PreVAD rather than random source initialization. The P-R differences are small and not statistically significant across five seeds (AUC exact two-sided sign-flip p = 0.1875).

Ethical Considerations and Licenses

  • The repository code is provided under the MIT License. Third-party dataset and model terms continue to apply to data-derived artifacts.
  • Users must comply with the licenses of the underlying third-party datasets (UCF-Crime and PreVAD) and the OpenAI CLIP feature extractor. The repository authors are not responsible for the misuse of these models in surveillance or real-world classification scenarios without appropriate safeguards.
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support