File size: 6,083 Bytes
dbdbc56 3364668 dbdbc56 3364668 dbdbc56 59200ff dbdbc56 377b35b 7de0a45 59200ff dbdbc56 7de0a45 dbdbc56 7de0a45 3364668 00e4088 7de0a45 3364668 7de0a45 3364668 dbdbc56 3364668 dbdbc56 3364668 7de0a45 3364668 7de0a45 3364668 377b35b 3364668 7de0a45 3364668 7de0a45 3364668 7de0a45 dbdbc56 7de0a45 dbdbc56 7de0a45 dbdbc56 7de0a45 00e4088 7de0a45 dbdbc56 7de0a45 dbdbc56 7de0a45 dbdbc56 7de0a45 dbdbc56 7de0a45 dbdbc56 7de0a45 dbdbc56 7de0a45 3364668 7de0a45 dbdbc56 377b35b 3364668 7de0a45 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 | ---
license: mit
library_name: pytorch
pipeline_tag: audio-classification
tags:
- audio
- audio-classification
- deepfake-detection
- source-tracing
- open-set-recognition
- wavlm
- encodec
datasets:
- mueller91/MLAAD
metrics:
- accuracy
---
# sourcetrace β the fitted head of CoRA (open-set audio deepfake attribution)
The fitted head of [`sourcetrace`](https://github.com/pujariaditya/CoRA),
which names **which synthesis model produced a synthetic speech clip** β or reports
that the model is not one it has seen. It is the artifact behind the paper *CoRA:
Robust Open-Set Audio Deepfake Attribution with Neural Codec Residuals*.
**This is not a standalone model.** `mlaad_v5.pt` is a `torch.save` dict with
`format: "sourcetrace-method-checkpoint"`, loaded by `sourcetrace.method.Method.load`.
It holds the small trained head plus the fitted scoring stack (class anchors,
relative-Mahalanobis density, z-norm constants, conformal calibration). The
front-ends β `microsoft/wavlm-large` and `facebook/encodec_24khz` β are frozen, are
**not** included here, and are fetched separately by `./setup.sh`.
## Model
| | |
|---|---|
| Input | 2133-d feature vector, **not** audio |
| SSL front end | `microsoft/wavlm-large`, **frozen**, layers 1β4, mean\|std pooled β 2048-d |
| Spectral signature | 24 group-delay bands + 16 modulation bins + 12 codec-grid dims β 52-d |
| Codec residual | `facebook/encodec_24khz` reconstruction residual at 1.5 / 6 / 12 kbps β 33-d |
| Head | factorized gated: SSL 192 + spectral 64 + codec 32, tanh FiLM gates |
| Embedding | **one** 288-d vector `z = L2([z_ssl β 0.1 z_spec β 0.5 z_codec])` |
| Open-set score | anchor margin + relative Mahalanobis on `z`; label = nearest anchor |
| Training | MLAAD v5 only: Adam, lr 3e-4, no weight decay, 300 epochs, batch 512 |
The same `z` serves every task: the open-set score, the known-model label, and the
STOPA cross-corpus protocol (cosine to enrolment fingerprints), where the model is
applied **without any STOPA training**.
The codec residual is the contribution: re-encode each clip through EnCodec-24 kHz
and keep what it got *wrong*. Those 33 numbers come off the waveform the speech
model never sees. Removing the channel moves FPR95 from 0.50 % to 1.28 %, OOD-EER
from 2.45 % to 3.45 % and known-model accuracy from 99.86 % to 99.68 % on MLAAD v5,
and the STOPA held-out model EER from 6.80 % to 8.08 %.
The head in the public code is a reconstruction from the published method
description, not a recovered original; its construction order is load-bearing for
RNG reproducibility.
## Input
Not audio. A **2133-d** feature vector per clip, laid out as
`[ SSL 0:2048 | signature 2048:2100 | codec residual 2100:2133 ]`, produced by
`python -m sourcetrace.extract`. There is no way to run these weights without the
repository and an extracted feature cache.
## Files
| file | what it is |
|---|---|
| `mlaad_v5.pt` | the fitted head, 65 known synthesis models, split seed 42 / fit seed 0 |
It is sha256-verified on download against the digest compiled into
`scripts/download_weights.py`, not served from here. Earlier files of this
repository (an 800-d two-channel `mlaad_v5.pt` and a STOPA refit `stopa.pt`) are
retired: the current code does not load them.
## Use
```bash
git clone https://github.com/pujariaditya/CoRA && cd CoRA
pip install -e . && ./setup.sh
python scripts/download_weights.py # sha256-pinned
python -m sourcetrace.evaluate --task both --checkpoint checkpoints/mlaad_v5.pt
```
Downloading saves the fit and nothing else: evaluation still reads the feature
caches, so MLAAD v5 and STOPA must be downloaded and extracted first. See the
repository README for that step.
## Results
**MLAAD v5**, family-level open-set protocol: 65 known synthesis models, 9,620
evaluation utterances (4,375 known-model, 5,245 held-out-model), split seed 42, fit
seed 0.
| Metric | Value | Published |
|---|---|---|
| FPR95 β | **0.50 %** | 3.36 % (Neamtu et al.) |
| OOD-EER β | 2.45 % | β |
| Known-model accuracy β | 99.86 % | β |
**STOPA**, released split and cosine-scoring protocol, the MLAAD-trained model
applied without retraining (33,200 enrolment and 629,800 probe utterances):
| Metric | Value |
|---|---|
| Held-out synthesis models, EER β | **6.80 %** |
| Known synthesis models, EER β | 9.24 % |
| Vocoder attribution, held-out / known, EER β | 7.81 % / 3.86 % |
STOPA's own trained baselines report 35.34 % (ASVspoof-trained AASIST), 47.75 %
(STOPA-trained AASIST) and 49.55 % (ResNet-34) for held-out models; the 16.43 %
zero-shot EER of Chhibber et al. uses a different data split, so no margin is claimed
over it.
**Reproducible, not just reported.** Refitting from the public code at these seeds
reproduces `results/ablation/full.json` exactly, and this file is that fit.
**Single-seed.** One split seed, one fit seed. These are point estimates; the paper
reports the spread over five training seeds.
## Limitations
- Trained on MLAAD v5 only. Attribution across other corpora, languages, codecs or
recording conditions is tested only on STOPA.
- For research on open-set attribution. **Not validated for forensic, legal or
moderation use**, and the abstention rule is calibrated on this protocol β its
coverage guarantee does not transfer off it.
- The harm from an attribution model is a confident wrong name, not a refusal. On an
unseen synthesis model the calibrated answer is *unknown*, and that answer is the
point of the system; do not deploy it anywhere the abstention is discarded, and do
not present an attribution as evidence about a person.
## Licence
MIT, matching the code repository. MLAAD and STOPA carry their own terms; no audio is
redistributed here.
## Citation
See [`CITATION.cff`](https://github.com/pujariaditya/CoRA/blob/master/CITATION.cff)
in the code repository, which is the single source for how to cite this.
Please also cite the benchmarks (MLAAD, STOPA) and the baselines this is compared
against (Neamtu et al.; Chhibber et al., Odyssey 2026).
|