zeechimp commited on
Commit
164140e
Β·
verified Β·
1 Parent(s): 7652f71

Update README.md

Browse files
Files changed (1) hide show
  1. README.md +170 -23
README.md CHANGED
@@ -1,39 +1,186 @@
1
- # LiarDetector v4 (HuggingFace)
 
 
 
 
 
 
 
 
 
 
 
 
 
 
2
 
3
- Self-Consistent-Liar Detector: given two aligned 1-D streams `A` and `B`,
4
- detects whether either is "lying" (distorted), which one, and what kind of
5
- distortion.
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
6
 
7
  ## Architecture
8
 
9
- - Input: 44 hand-crafted features from `(A, B)` β€” raw moments, difference
10
- moments, residual statistics, correlations, reference-free signatures.
11
- - Trunk: `Linear(44β†’96) β†’ ReLU β†’ Linear(96β†’64) β†’ ReLU`
12
- - Three heads: binary (which stream lies), family (distortion type),
13
- presence (is there a lie at all).
14
- - **Masked loss**: binary + family heads only update on single-lie pairs;
15
- presence head trains on all pairs. This fixes the v2/v3 loss instability.
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
16
 
17
  ## Usage
18
 
19
  ```python
20
- from liar_detector import (
21
- LiarDetectorForLieDetection,
22
- LiarDetectorFeatureExtractor,
23
- synth_true, apply_lie,
24
- )
25
  import numpy as np
 
 
 
26
 
27
- model = LiarDetectorForLieDetection.from_pretrained("user/liar-detector-v4").eval()
28
- fe = LiarDetectorFeatureExtractor.from_pretrained("user/liar-detector-v4")
 
 
 
29
 
30
- rng = np.random.default_rng(0)
31
- A = synth_true(rng)
32
- B = apply_lie(A, rng, "offset")
 
 
 
 
 
 
 
 
 
33
 
34
- import torch
35
  with torch.no_grad():
36
- out = model(features=torch.from_numpy(fe(A, B)).unsqueeze(0))
37
 
38
  print("which lies:", out.binary_probs.argmax().item()) # 0=A, 1=B
39
  print("family :", out.family_probs.argmax().item())
 
1
+ ---
2
+ language:
3
+ - en
4
+ license: apache-2.0
5
+ library_name: transformers
6
+ tags:
7
+ - liar-detection
8
+ - signal-processing
9
+ - anomaly-detection
10
+ - pytorch
11
+ - numpy
12
+ - audio
13
+ - time-series
14
+ pipeline_tag: audio-classification
15
+ ---
16
 
17
+ # Liar Detector v4
18
+
19
+ ## Model Summary
20
+
21
+ `liar-detector-v4` is a lightweight, fully-connected neural network that detects
22
+ whether one of two aligned 1-D signals is "lying" β€” i.e., distorted by a known
23
+ manipulation family. Given two streams `A` and `B`, it simultaneously predicts:
24
+
25
+ 1. **Which stream lies** (binary head)
26
+ 2. **What kind of lie it is** (family head, 6 classes)
27
+ 3. **Whether any lie is present at all** (presence head)
28
+
29
+ The model is designed for research on self-consistent signal verification,
30
+ sensor fusion, and reference-free anomaly detection. It operates on 44
31
+ hand-crafted features extracted from the `(A, B)` pair and requires no external
32
+ reference signal.
33
+
34
+ ## Model Details
35
+
36
+ - **Developed by:** Sylv Q ([zeechimp](https://huggingface.co/zeechimp))
37
+ - **Model type:** Multi-head MLP (3 heads, shared trunk)
38
+ - **Language(s):** N/A (numeric signal input)
39
+ - **License:** Apache 2.0
40
+ - **Finetuned from:** Not applicable (trained from scratch)
41
+ - **Repository:** [zeechimp/liar-detector-v4](https://huggingface.co/zeechimp/liar-detector-v4)
42
 
43
  ## Architecture
44
 
45
+ | Component | Detail |
46
+ |-----------|--------|
47
+ | Input | 44-dim feature vector extracted from `(A, B)` |
48
+ | Trunk | `Linear(44β†’96) β†’ ReLU β†’ Linear(96β†’64) β†’ ReLU` |
49
+ | Binary head | `Linear(64β†’2)` β€” which stream lies (0=A, 1=B) |
50
+ | Family head | `Linear(64β†’6)` β€” distortion type |
51
+ | Presence head | `Linear(64β†’2)` β€” is any lie present |
52
+ | Parameters | ~11K |
53
+ | Framework | PyTorch + πŸ€— Transformers |
54
+
55
+ ### Feature Groups (44 total)
56
+
57
+ - **Raw moments** β€” mean, log-std, skew, kurtosis for both streams (8)
58
+ - **Canonical reference** β€” `abs_mean_A - abs_mean_B`, `log_std_ratio_canonical` (2)
59
+ - **Difference features** β€” paired difference moments (6)
60
+ - **Residual statistics** β€” moments, autocorrelation, cross-correlation (12)
61
+ - **Regression** β€” residual slope / explained variance (2)
62
+ - **Reference-free signatures** β€” quantization, stair-step, inversion, smoothness (10)
63
+ - **Uniqueness** β€” unique-fraction ratios (4)
64
+
65
+ ## Intended Uses & Limitations
66
+
67
+ ### Direct Use
68
+
69
+ - Research on signal integrity and self-consistency checking
70
+ - Benchmarking reference-free anomaly detectors
71
+ - Educational demonstrations of multi-head classification
72
+
73
+ ### Downstream Use
74
+
75
+ - Sensor-fusion pipelines where one stream may be tampered
76
+ - Audio/telemetry verification (with domain-specific retraining)
77
+
78
+ ### Out-of-Scope Use
79
+
80
+ - **Not** a general-purpose "lie detector" for human speech or text
81
+ - **Not** validated on real-world sensor data (trained on synthetic signals)
82
+ - **Not** suitable for high-stakes decisions without extensive domain adaptation
83
+
84
+ ## Training Details
85
+
86
+ ### Training Data
87
+
88
+ Synthetic 1-D signals generated as sums of three sinusoids (frequencies
89
+ 0.01–0.15 Hz, amplitudes 0.5–2.0), with additive Gaussian noise. Distortions
90
+ are injected via the following families:
91
+
92
+ | Family | Type | Description |
93
+ |--------|------|-------------|
94
+ | `offset` | ID | Additive constant bias |
95
+ | `scale` | ID | Multiplicative gain (0.85–1.15Γ—) |
96
+ | `saturation` | ID | Hard clipping at Β±M |
97
+ | `quantization` | ID | Rounding to a grid |
98
+ | `lag` | ID | Integer time shift (1–4 samples) |
99
+ | `harmonic` | ID | Quadratic self-interaction term |
100
+ | `deadzone` | OOD | Zeroing of near-zero values |
101
+ | `dropout` | OOD | Random sample-and-hold |
102
+ | `drift` | OOD | Cumulative Gaussian random walk |
103
+ | `inversion` | OOD | Sign flip for small values |
104
+
105
+ ### Training Procedure
106
+
107
+ - **Epochs:** 200
108
+ - **Batch size:** 64
109
+ - **Optimizer:** AdamW
110
+ - **Learning rate:** 3e-3
111
+ - **Loss:** Masked cross-entropy (binary + family heads only train on single-lie
112
+ pairs; presence head trains on all pairs)
113
+ - **Calibration:** Temperature scaling fitted per head
114
+
115
+ ## Evaluation Results
116
+
117
+ ### In-Distribution (ID)
118
+
119
+ | Condition | Accuracy | ECE |
120
+ |-----------|----------|-----|
121
+ | Binary (which stream lies) | 86.1% | 0.116 |
122
+ | Family (distortion type) | 91.2% | 0.309 |
123
+ | Presence (any lie) | 72.9% | 0.043 |
124
+
125
+ ### Per-Family Binary Accuracy (ID)
126
+
127
+ | Family | Accuracy |
128
+ |--------|----------|
129
+ | offset | 96.3% |
130
+ | scale | 57.5% |
131
+ | saturation | 98.9% |
132
+ | quantization | 68.6% |
133
+ | lag | 94.6% |
134
+ | harmonic | 98.8% |
135
+
136
+ ### Out-of-Distribution (OOD)
137
+
138
+ | Family | Binary Accuracy |
139
+ |--------|-----------------|
140
+ | deadzone | 87.2% |
141
+ | dropout | 96.4% |
142
+ | drift | 81.2% |
143
+ | inversion | 90.0% |
144
+
145
+ **Overall OOD binary accuracy:** 88.7% (ECE 0.137)
146
+
147
+ ### Sanity Checks
148
+
149
+ | Scenario | Presence Accuracy |
150
+ |----------|-------------------|
151
+ | Both honest | 61.5% |
152
+ | Both lying | 56.3% |
153
 
154
  ## Usage
155
 
156
  ```python
157
+ import torch
 
 
 
 
158
  import numpy as np
159
+ from transformers import AutoModel, AutoConfig
160
+ from huggingface_hub import hf_hub_download
161
+ import json
162
 
163
+ # Load model
164
+ model = AutoModel.from_pretrained(
165
+ "zeechimp/liar-detector-v4",
166
+ trust_remote_code=True
167
+ ).eval()
168
 
169
+ # Load feature extractor normalization stats
170
+ fe_path = hf_hub_download("zeechimp/liar-detector-v4", "feature_extractor.json")
171
+ with open(fe_path) as f:
172
+ fe_stats = json.load(f)
173
+ mu = np.array(fe_stats["mean"], dtype=np.float32)
174
+ sigma = np.array(fe_stats["std"], dtype=np.float32)
175
+
176
+ # Extract features (see feature_extractor.py in the repo)
177
+ # A, B are 1-D numpy arrays of equal length
178
+ features = extract_features(A, B) # returns 44-dim vector
179
+ x = (features - mu) / sigma
180
+ x = torch.from_numpy(x).unsqueeze(0) # (1, 44)
181
 
 
182
  with torch.no_grad():
183
+ out = model(features=x)
184
 
185
  print("which lies:", out.binary_probs.argmax().item()) # 0=A, 1=B
186
  print("family :", out.family_probs.argmax().item())