LaST-Net / README.md
JunXueTech's picture
Simplify model card and update author contact
7b18ff0 verified
|
Raw History Blame Contribute Delete
1.8 kB
---
tags:
- audio
- speech-deepfake-detection
- pytorch
- fairseq
base_model: facebook/wav2vec2-xls-r-300m
---
# LaST-Net
Model weights for **LaST-Net: Length-Aware Layer and Scale-Adaptive Temporal Network for Speech Deepfake Detection**.
**Training and inference code:** [JunXue-tech/LaST-Net](https://github.com/JunXue-tech/LaST-Net).
## Checkpoint
`best.pt` is the original epoch-52 checkpoint selected by the lowest mean development EER across 1, 2, 4 and 6 seconds on ASVspoof 2019 LA. It includes the fine-tuned XLS-R 300M frontend, LaST-Net backend, optimizer state and original training metadata.
## Usage
Follow the environment setup in the [code repository](https://github.com/JunXue-tech/LaST-Net). From that repository:
```bash
python download_model.py
python infer.py example.wav --checkpoint checkpoints/best.pt \
--ssl-path /path/to/xlsr2_300m.pt --seconds 6
```
The model constructor requires the original fairseq-format XLS-R 300M checkpoint, available from the [official XLS-R repository](https://github.com/facebookresearch/fairseq/tree/main/examples/wav2vec/xlsr), before loading the fine-tuned parameters from `best.pt`.
Input audio must be mono at 16 kHz. The supplied inference code evaluates 1–6 second inputs using prefix cropping and repetition of shorter recordings. Higher `bonafide_log_score` values favor bona fide speech. Scores are not calibrated probabilities.
## Evaluation
Duration-averaged EER (%) across 1–6 second inputs: 19LA **1.29**, 21LA **4.98**, 21DF **3.62**, and In-the-Wild **7.58**. Per-duration results and evaluation commands are provided in the code repository. Results depend on the evaluation protocol and preprocessing.
## Author
Jun Xue — [junxue@whu.edu.cn](mailto:junxue@whu.edu.cn)