File size: 6,083 Bytes
dbdbc56
 
3364668
 
dbdbc56
 
 
 
 
 
 
 
3364668
 
 
 
dbdbc56
 
59200ff
dbdbc56
377b35b
7de0a45
59200ff
 
dbdbc56
7de0a45
dbdbc56
7de0a45
 
 
 
3364668
 
 
 
 
 
00e4088
7de0a45
 
 
 
 
 
 
 
 
 
 
 
3364668
7de0a45
 
 
3364668
 
 
 
dbdbc56
 
 
 
 
3364668
dbdbc56
 
3364668
 
7de0a45
 
 
3364668
7de0a45
 
 
 
3364668
 
 
 
377b35b
3364668
7de0a45
 
3364668
 
7de0a45
 
3364668
 
7de0a45
dbdbc56
7de0a45
 
 
dbdbc56
7de0a45
dbdbc56
7de0a45
00e4088
7de0a45
dbdbc56
7de0a45
 
dbdbc56
7de0a45
 
 
 
 
dbdbc56
7de0a45
 
 
 
dbdbc56
7de0a45
 
dbdbc56
7de0a45
 
dbdbc56
 
 
 
7de0a45
3364668
 
 
 
7de0a45
 
 
dbdbc56
 
 
 
 
 
 
 
377b35b
3364668
 
 
7de0a45
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
---
license: mit
library_name: pytorch
pipeline_tag: audio-classification
tags:
  - audio
  - audio-classification
  - deepfake-detection
  - source-tracing
  - open-set-recognition
  - wavlm
  - encodec
datasets:
  - mueller91/MLAAD
metrics:
  - accuracy
---

# sourcetrace β€” the fitted head of CoRA (open-set audio deepfake attribution)

The fitted head of [`sourcetrace`](https://github.com/pujariaditya/CoRA),
which names **which synthesis model produced a synthetic speech clip** β€” or reports
that the model is not one it has seen. It is the artifact behind the paper *CoRA:
Robust Open-Set Audio Deepfake Attribution with Neural Codec Residuals*.

**This is not a standalone model.** `mlaad_v5.pt` is a `torch.save` dict with
`format: "sourcetrace-method-checkpoint"`, loaded by `sourcetrace.method.Method.load`.
It holds the small trained head plus the fitted scoring stack (class anchors,
relative-Mahalanobis density, z-norm constants, conformal calibration). The
front-ends β€” `microsoft/wavlm-large` and `facebook/encodec_24khz` β€” are frozen, are
**not** included here, and are fetched separately by `./setup.sh`.

## Model

| | |
|---|---|
| Input | 2133-d feature vector, **not** audio |
| SSL front end | `microsoft/wavlm-large`, **frozen**, layers 1–4, mean\|std pooled β†’ 2048-d |
| Spectral signature | 24 group-delay bands + 16 modulation bins + 12 codec-grid dims β†’ 52-d |
| Codec residual | `facebook/encodec_24khz` reconstruction residual at 1.5 / 6 / 12 kbps β†’ 33-d |
| Head | factorized gated: SSL 192 + spectral 64 + codec 32, tanh FiLM gates |
| Embedding | **one** 288-d vector `z = L2([z_ssl β€– 0.1 z_spec β€– 0.5 z_codec])` |
| Open-set score | anchor margin + relative Mahalanobis on `z`; label = nearest anchor |
| Training | MLAAD v5 only: Adam, lr 3e-4, no weight decay, 300 epochs, batch 512 |

The same `z` serves every task: the open-set score, the known-model label, and the
STOPA cross-corpus protocol (cosine to enrolment fingerprints), where the model is
applied **without any STOPA training**.

The codec residual is the contribution: re-encode each clip through EnCodec-24 kHz
and keep what it got *wrong*. Those 33 numbers come off the waveform the speech
model never sees. Removing the channel moves FPR95 from 0.50 % to 1.28 %, OOD-EER
from 2.45 % to 3.45 % and known-model accuracy from 99.86 % to 99.68 % on MLAAD v5,
and the STOPA held-out model EER from 6.80 % to 8.08 %.

The head in the public code is a reconstruction from the published method
description, not a recovered original; its construction order is load-bearing for
RNG reproducibility.

## Input

Not audio. A **2133-d** feature vector per clip, laid out as
`[ SSL 0:2048 | signature 2048:2100 | codec residual 2100:2133 ]`, produced by
`python -m sourcetrace.extract`. There is no way to run these weights without the
repository and an extracted feature cache.

## Files

| file | what it is |
|---|---|
| `mlaad_v5.pt` | the fitted head, 65 known synthesis models, split seed 42 / fit seed 0 |

It is sha256-verified on download against the digest compiled into
`scripts/download_weights.py`, not served from here. Earlier files of this
repository (an 800-d two-channel `mlaad_v5.pt` and a STOPA refit `stopa.pt`) are
retired: the current code does not load them.

## Use

```bash
git clone https://github.com/pujariaditya/CoRA && cd CoRA
pip install -e . && ./setup.sh
python scripts/download_weights.py      # sha256-pinned
python -m sourcetrace.evaluate --task both --checkpoint checkpoints/mlaad_v5.pt
```

Downloading saves the fit and nothing else: evaluation still reads the feature
caches, so MLAAD v5 and STOPA must be downloaded and extracted first. See the
repository README for that step.

## Results

**MLAAD v5**, family-level open-set protocol: 65 known synthesis models, 9,620
evaluation utterances (4,375 known-model, 5,245 held-out-model), split seed 42, fit
seed 0.

| Metric | Value | Published |
|---|---|---|
| FPR95 ↓ | **0.50 %** | 3.36 % (Neamtu et al.) |
| OOD-EER ↓ | 2.45 % | β€” |
| Known-model accuracy ↑ | 99.86 % | β€” |

**STOPA**, released split and cosine-scoring protocol, the MLAAD-trained model
applied without retraining (33,200 enrolment and 629,800 probe utterances):

| Metric | Value |
|---|---|
| Held-out synthesis models, EER ↓ | **6.80 %** |
| Known synthesis models, EER ↓ | 9.24 % |
| Vocoder attribution, held-out / known, EER ↓ | 7.81 % / 3.86 % |

STOPA's own trained baselines report 35.34 % (ASVspoof-trained AASIST), 47.75 %
(STOPA-trained AASIST) and 49.55 % (ResNet-34) for held-out models; the 16.43 %
zero-shot EER of Chhibber et al. uses a different data split, so no margin is claimed
over it.

**Reproducible, not just reported.** Refitting from the public code at these seeds
reproduces `results/ablation/full.json` exactly, and this file is that fit.

**Single-seed.** One split seed, one fit seed. These are point estimates; the paper
reports the spread over five training seeds.

## Limitations

- Trained on MLAAD v5 only. Attribution across other corpora, languages, codecs or
  recording conditions is tested only on STOPA.
- For research on open-set attribution. **Not validated for forensic, legal or
  moderation use**, and the abstention rule is calibrated on this protocol β€” its
  coverage guarantee does not transfer off it.
- The harm from an attribution model is a confident wrong name, not a refusal. On an
  unseen synthesis model the calibrated answer is *unknown*, and that answer is the
  point of the system; do not deploy it anywhere the abstention is discarded, and do
  not present an attribution as evidence about a person.

## Licence

MIT, matching the code repository. MLAAD and STOPA carry their own terms; no audio is
redistributed here.

## Citation

See [`CITATION.cff`](https://github.com/pujariaditya/CoRA/blob/master/CITATION.cff)
in the code repository, which is the single source for how to cite this.

Please also cite the benchmarks (MLAAD, STOPA) and the baselines this is compared
against (Neamtu et al.; Chhibber et al., Odyssey 2026).