Liar Detector v4
Model Summary
liar-detector-v4 is a lightweight, fully-connected neural network that detects
whether one of two aligned 1-D signals is "lying" β i.e., distorted by a known
manipulation family. Given two streams A and B, it simultaneously predicts:
- Which stream lies (binary head)
- What kind of lie it is (family head, 6 classes)
- Whether any lie is present at all (presence head)
The model is designed for research on self-consistent signal verification,
sensor fusion, and reference-free anomaly detection. It operates on 44
hand-crafted features extracted from the (A, B) pair and requires no external
reference signal.
Model Details
- Developed by: Sylv Q (zeechimp)
- Model type: Multi-head MLP (3 heads, shared trunk)
- Language(s): N/A (numeric signal input)
- License: Apache 2.0
- Finetuned from: Not applicable (trained from scratch)
- Repository: zeechimp/liar-detector-v4
Architecture
| Component |
Detail |
| Input |
44-dim feature vector extracted from (A, B) |
| Trunk |
Linear(44β96) β ReLU β Linear(96β64) β ReLU |
| Binary head |
Linear(64β2) β which stream lies (0=A, 1=B) |
| Family head |
Linear(64β6) β distortion type |
| Presence head |
Linear(64β2) β is any lie present |
| Parameters |
~11K |
| Framework |
PyTorch + π€ Transformers |
Feature Groups (44 total)
- Raw moments β mean, log-std, skew, kurtosis for both streams (8)
- Canonical reference β
abs_mean_A - abs_mean_B, log_std_ratio_canonical (2)
- Difference features β paired difference moments (6)
- Residual statistics β moments, autocorrelation, cross-correlation (12)
- Regression β residual slope / explained variance (2)
- Reference-free signatures β quantization, stair-step, inversion, smoothness (10)
- Uniqueness β unique-fraction ratios (4)
Intended Uses & Limitations
Direct Use
- Research on signal integrity and self-consistency checking
- Benchmarking reference-free anomaly detectors
- Educational demonstrations of multi-head classification
Downstream Use
- Sensor-fusion pipelines where one stream may be tampered
- Audio/telemetry verification (with domain-specific retraining)
Out-of-Scope Use
- Not a general-purpose "lie detector" for human speech or text
- Not validated on real-world sensor data (trained on synthetic signals)
- Not suitable for high-stakes decisions without extensive domain adaptation
Training Details
Training Data
Synthetic 1-D signals generated as sums of three sinusoids (frequencies
0.01β0.15 Hz, amplitudes 0.5β2.0), with additive Gaussian noise. Distortions
are injected via the following families:
| Family |
Type |
Description |
offset |
ID |
Additive constant bias |
scale |
ID |
Multiplicative gain (0.85β1.15Γ) |
saturation |
ID |
Hard clipping at Β±M |
quantization |
ID |
Rounding to a grid |
lag |
ID |
Integer time shift (1β4 samples) |
harmonic |
ID |
Quadratic self-interaction term |
deadzone |
OOD |
Zeroing of near-zero values |
dropout |
OOD |
Random sample-and-hold |
drift |
OOD |
Cumulative Gaussian random walk |
inversion |
OOD |
Sign flip for small values |
Training Procedure
- Epochs: 200
- Batch size: 64
- Optimizer: AdamW
- Learning rate: 3e-3
- Loss: Masked cross-entropy (binary + family heads only train on single-lie
pairs; presence head trains on all pairs)
- Calibration: Temperature scaling fitted per head
Evaluation Results
In-Distribution (ID)
| Condition |
Accuracy |
ECE |
| Binary (which stream lies) |
86.1% |
0.116 |
| Family (distortion type) |
91.2% |
0.309 |
| Presence (any lie) |
72.9% |
0.043 |
Per-Family Binary Accuracy (ID)
| Family |
Accuracy |
| offset |
96.3% |
| scale |
57.5% |
| saturation |
98.9% |
| quantization |
68.6% |
| lag |
94.6% |
| harmonic |
98.8% |
Out-of-Distribution (OOD)
| Family |
Binary Accuracy |
| deadzone |
87.2% |
| dropout |
96.4% |
| drift |
81.2% |
| inversion |
90.0% |
Overall OOD binary accuracy: 88.7% (ECE 0.137)
Sanity Checks
| Scenario |
Presence Accuracy |
| Both honest |
61.5% |
| Both lying |
56.3% |
Usage
import torch
import numpy as np
from transformers import AutoModel, AutoConfig
from huggingface_hub import hf_hub_download
import json
model = AutoModel.from_pretrained(
"zeechimp/liar-detector-v4",
trust_remote_code=True
).eval()
fe_path = hf_hub_download("zeechimp/liar-detector-v4", "feature_extractor.json")
with open(fe_path) as f:
fe_stats = json.load(f)
mu = np.array(fe_stats["mean"], dtype=np.float32)
sigma = np.array(fe_stats["std"], dtype=np.float32)
features = extract_features(A, B)
x = (features - mu) / sigma
x = torch.from_numpy(x).unsqueeze(0)
with torch.no_grad():
out = model(features=x)
print("which lies:", out.binary_probs.argmax().item())
print("family :", out.family_probs.argmax().item())
print("presence :", out.presence_probs.argmax().item())