deepsafe commited on
Commit
fb92c8e
·
verified ·
1 Parent(s): 6be09d5

Add ensemble meta-learners, scalers and calibrators

Browse files
README.md ADDED
@@ -0,0 +1,59 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ license: other
3
+ license_name: polyform-noncommercial-1.0.0
4
+ license_link: https://polyformproject.org/licenses/noncommercial/1.0.0
5
+ tags:
6
+ - deepfake-detection
7
+ - ensemble
8
+ ---
9
+
10
+ # DeepSafe Ensemble Artifacts
11
+
12
+ The meta-learners that turn 19 individual detector scores into one calibrated
13
+ verdict, for [DeepSafe](https://github.com/deepsafehq/deepsafe-bench).
14
+
15
+ Unlike the model code and weights in the other DeepSafe repositories, **these
16
+ are first-party**: trained by us, licensed PolyForm Noncommercial 1.0.0, same
17
+ as the project.
18
+
19
+ ## Contents
20
+
21
+ | Modality | Meta-learner | Held-out AUC |
22
+ |---|---|---|
23
+ | Image | LightGBM over 7 models | 0.9466 |
24
+ | Audio | Random Forest over 3 models | 0.8290 |
25
+ | Video | XGBoost over 9 models | 0.6694 |
26
+
27
+ Each modality ships four files:
28
+
29
+ - `<modality>_meta_learner.pkl` — the trained model
30
+ - `<modality>_scaler.pkl` — feature scaling
31
+ - `<modality>_calibrator.pkl` — Platt calibration, so scores read as probabilities
32
+ - `<modality>_config.json` — feature order and fallback weights
33
+
34
+ Trained on the 15,499-sample medium evaluation tier. Without these, the
35
+ inference server produces per-model scores but no ensemble verdict.
36
+
37
+ ## Usage
38
+
39
+ `setup.sh` fetches these automatically. Manually:
40
+
41
+ ```python
42
+ from huggingface_hub import snapshot_download
43
+ snapshot_download("deepsafe/ensemble", local_dir="models/ensemble/artifacts")
44
+ ```
45
+
46
+ ## The numbers are the point
47
+
48
+ Held-out video AUC is **0.6694**. That is barely above chance on generators the
49
+ models were not trained for, and it is lower than the cross-validated figure
50
+ produced during training. We publish the held-out number because the gap
51
+ between the two is the finding. See
52
+ [BENCHMARK.md](https://github.com/deepsafehq/deepsafe-bench/blob/main/BENCHMARK.md).
53
+
54
+ ## Security note
55
+
56
+ These are Python pickles, which execute code on load. Only load them from a
57
+ source you trust. DeepSafe's loader refuses any pickle outside its configured
58
+ artifacts directory, but that is a guardrail, not a guarantee. If you are
59
+ security-sensitive, retrain your own with `deepsafe fit --tier 1`.
audio_calibrator.pkl ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:07c382f323a548522bb3f153c10211ff98af56ef571b84ccd8083fb61a27a9af
3
+ size 713
audio_config.json ADDED
@@ -0,0 +1,34 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "model_order": [
3
+ "shiftyspeech",
4
+ "safeear",
5
+ "nes2net"
6
+ ],
7
+ "algorithm": "RandomForestClassifier",
8
+ "hyperparams": {
9
+ "max_depth": 8,
10
+ "n_estimators": 500,
11
+ "random_state": 42
12
+ },
13
+ "cv_auc": 0.8748,
14
+ "cv_f1_at_05": 0.8749,
15
+ "optimal_threshold": 0.47,
16
+ "optimal_f1": 0.8787,
17
+ "ece_before_calibration": 0.0132,
18
+ "ece_after_calibration": 0.0338,
19
+ "calibrated_auc": 0.8748,
20
+ "calibrated_optimal_threshold": 0.39,
21
+ "calibrated_optimal_f1": 0.8789,
22
+ "calibrator_saved": true,
23
+ "n_samples": 3500,
24
+ "n_real": 1000,
25
+ "n_fake": 2500,
26
+ "measured_model_aucs": {
27
+ "shiftyspeech": 0.7347,
28
+ "safeear": 0.6035,
29
+ "nes2net": 0.8407
30
+ },
31
+ "training_date": "2026-04-12",
32
+ "training_data": "monolith_eval_medium_15499_patched",
33
+ "note": "3-model RandomForestClassifier trained on 15,499-sample medium eval. Replaces small-eval (50 sample) artifacts."
34
+ }
audio_meta_learner.pkl ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:ea2389cd393b42880d6f1672a9281b1ff7789193fd36bbd6512027c97afe6d88
3
+ size 7928136
audio_scaler.pkl ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:11b6d94e752a1247301e6790f8b50ed3d1a9a7c355132172d925ab4e028fd296
3
+ size 488
image_calibrator.pkl ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:e3083c3e12fd399aa4d3ec67f9cf9b78a2c29ff72895a63b04680b07e8951d56
3
+ size 713
image_config.json ADDED
@@ -0,0 +1,48 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "model_order": [
3
+ "aide",
4
+ "cospy",
5
+ "effort",
6
+ "fsd",
7
+ "npr",
8
+ "universal",
9
+ "yermandy"
10
+ ],
11
+ "algorithm": "LGBMClassifier",
12
+ "hyperparams": {
13
+ "colsample_bytree": 1.0,
14
+ "learning_rate": 0.05,
15
+ "max_depth": 4,
16
+ "min_child_weight": 0.001,
17
+ "n_estimators": 200,
18
+ "random_state": 42,
19
+ "reg_alpha": 0.0,
20
+ "reg_lambda": 1.0,
21
+ "subsample": 1.0
22
+ },
23
+ "cv_auc": 0.9705,
24
+ "cv_f1_at_05": 0.9033,
25
+ "optimal_threshold": 0.41,
26
+ "optimal_f1": 0.9065,
27
+ "ece_before_calibration": 0.0096,
28
+ "ece_after_calibration": 0.0281,
29
+ "calibrated_auc": 0.9705,
30
+ "calibrated_optimal_threshold": 0.35,
31
+ "calibrated_optimal_f1": 0.9065,
32
+ "calibrator_saved": true,
33
+ "n_samples": 10000,
34
+ "n_real": 5000,
35
+ "n_fake": 5000,
36
+ "measured_model_aucs": {
37
+ "aide": 0.7478,
38
+ "cospy": 0.8679,
39
+ "effort": 0.8353,
40
+ "fsd": 0.8961,
41
+ "npr": 0.6373,
42
+ "universal": 0.6711,
43
+ "yermandy": 0.5912
44
+ },
45
+ "training_date": "2026-04-12",
46
+ "training_data": "monolith_eval_medium_15499_patched",
47
+ "note": "7-model LGBMClassifier trained on 15,499-sample medium eval. Replaces small-eval (50 sample) artifacts."
48
+ }
image_meta_learner.pkl ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:1536f01547326da6436d4462411205cf333a35abbf128c6cbf3b13e913c44cd3
3
+ size 318970
image_scaler.pkl ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:71e2c7ba983fdd80a3b9abb41276af90003d92b396a446bc5bc226d120dfdade
3
+ size 584
video_calibrator.pkl ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:5bfd34dd98bf48a11337d88dca299577d03f9d1d2e8232a717808714e323084f
3
+ size 713
video_config.json ADDED
@@ -0,0 +1,48 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "model_order": [
3
+ "fakestormer",
4
+ "sbi",
5
+ "dfd_fcg",
6
+ "pwtf_dvd",
7
+ "lipfd",
8
+ "recce",
9
+ "mintime",
10
+ "npr_video",
11
+ "univfd_video"
12
+ ],
13
+ "algorithm": "XGBClassifier",
14
+ "hyperparams": {
15
+ "learning_rate": 0.05,
16
+ "max_depth": 3,
17
+ "n_estimators": 300,
18
+ "random_state": 42,
19
+ "reg_lambda": 10.0
20
+ },
21
+ "cv_auc": 0.8898,
22
+ "cv_f1_at_05": 0.8133,
23
+ "optimal_threshold": 0.48,
24
+ "optimal_f1": 0.8147,
25
+ "ece_before_calibration": 0.0219,
26
+ "ece_after_calibration": 0.0221,
27
+ "calibrated_auc": 0.8898,
28
+ "calibrated_optimal_threshold": 0.48,
29
+ "calibrated_optimal_f1": 0.8141,
30
+ "calibrator_saved": true,
31
+ "n_samples": 1999,
32
+ "n_real": 999,
33
+ "n_fake": 1000,
34
+ "measured_model_aucs": {
35
+ "fakestormer": 0.6618,
36
+ "sbi": 0.6781,
37
+ "dfd_fcg": 0.5531,
38
+ "pwtf_dvd": 0.585,
39
+ "lipfd": 0.6312,
40
+ "recce": 0.6625,
41
+ "mintime": 0.5813,
42
+ "npr_video": 0.8022,
43
+ "univfd_video": 0.6798
44
+ },
45
+ "training_date": "2026-04-12",
46
+ "training_data": "monolith_eval_medium_15499_patched",
47
+ "note": "9-model XGBClassifier trained on 15,499-sample medium eval. Replaces small-eval (50 sample) artifacts."
48
+ }
video_meta_learner.pkl ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:b68f468744cfee759ce18641c5622b15d5877e9046db366312a0db9e5ab92750
3
+ size 346122
video_scaler.pkl ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:c489f49490a0fafa115c7ec3cc1bd49fb196c5881e4cae04516e219847f110db
3
+ size 632