Liar Detector v5

Model Summary

liar-detector-v5 is a lightweight, fully-connected neural network that detects whether one of two aligned 1-D signals is "lying" — i.e., distorted by a known manipulation family. Given two streams A and B, it simultaneously predicts:

  1. Which stream lies (binary head: 0 = A lies, 1 = B lies)
  2. What kind of lie it is (family head, 6 classes)
  3. Whether any lie is present at all (presence head)

The model is designed for research on self-consistent signal verification, sensor fusion, and reference-free anomaly detection. It operates on 44 hand-crafted features extracted from the (A, B) pair and requires no external reference signal.

Model Details

  • Developed by: Sylv Q (zeechimp)
  • Model type: Multi-head MLP (3 heads, shared trunk, presence adapter)
  • Language(s): N/A (numeric signal input)
  • License: Apache 2.0
  • Finetuned from: Not applicable (trained from scratch)
  • Repository: zeechimp/liar-detector-v5

Architecture

Component Detail
Input 44-dim feature vector extracted from (A, B)
Trunk Linear(44→96) → ReLU → Linear(96→64) → ReLU
Binary head Linear(64→2) — which stream lies (0=A, 1=B)
Family head Linear(64→6) — distortion type
Presence adapter Linear(64→64) → ReLU → Linear(64→64)
Presence head Linear(64→2) — is any lie present
Parameters ~19,500
Framework PyTorch + 🤗 Transformers

The presence head has its own adapter between the trunk and the classifier. The trunk features that discriminate "which stream differs" are not identical to the features that discriminate "does any stream differ", so the presence head needs a learned remapping rather than competing with the binary and family heads for the same representation.

Feature Groups (44 total)

  • Raw moments — mean, log-std, skew, kurtosis for both streams (8)
  • Canonical reference — abs_mean_A - abs_mean_B, log_std_ratio_canonical (2)
  • Difference features — paired difference moments (6)
  • Residual statistics — moments, autocorrelation, cross-correlation (12)
  • Regression — residual slope / explained variance (2)
  • Reference-free signatures — quantization, stair-step, inversion, smoothness (10)
  • Uniqueness — unique-fraction ratios (4)

Intended Uses & Limitations

Direct Use

  • Research on signal integrity and self-consistency checking
  • Benchmarking reference-free anomaly detectors
  • Educational demonstrations of multi-head classification with calibration

Downstream Use

  • Sensor-fusion pipelines where one stream may be tampered
  • Audio/telemetry verification (with domain-specific retraining)

Out-of-Scope Use

  • Not a general-purpose "lie detector" for human speech or text
  • Not validated on real-world sensor data (trained on synthetic signals)
  • Not suitable for high-stakes decisions without extensive domain adaptation

Training Details

Training Data

Synthetic 1-D signals generated as sums of three sinusoids (frequencies 0.01–0.15 Hz, amplitudes 0.5–2.0), with additive Gaussian noise (noise_std ~ U(0.05, 0.15) * std(x)). Distortions are injected via the following families:

Family Type Description
offset ID Additive constant bias
scale ID Multiplicative gain (0.85–1.15×)
saturation ID Hard clipping at ±M
quantization ID Rounding to a grid
lag ID Integer time shift (1–4 samples)
harmonic ID Quadratic self-interaction term
deadzone OOD Zeroing of near-zero values
dropout OOD Random sample-and-hold
drift OOD Cumulative Gaussian random walk
inversion OOD Sign flip for small values

Training Procedure

  • Epochs: 200
  • Batch size: 64
  • Optimizer: AdamW (lr = 3e-3, weight_decay = 1e-5)
  • Loss: Masked cross-entropy. Binary and family heads train only on single-lie pairs; presence head trains on all pairs with class weights [1.47, 0.76] to compensate for the 66/34 scenario split.
  • Calibration: Temperature scaling fitted per head on a held-out 20% split.

Evaluation Results

In-Distribution (ID)

Condition Accuracy ECE
Binary (which stream lies) 88.2% 0.091
Family (distortion type) 92.3% 0.224
Presence (any lie) 71.5% 0.143

Per-Family Binary Accuracy (ID)

Family Accuracy
offset 95.8%
scale 57.4%
saturation 100.0%
quantization 76.0%
lag 99.4%
harmonic 99.4%

Out-of-Distribution (OOD)

Family Binary Accuracy
deadzone 93.6%
dropout 96.8%
drift 90.4%
inversion 90.4%

Overall OOD binary accuracy: 92.8% (ECE 0.148)

Known Limitations

  • scale family is weak (57.4%). When two streams differ only by a multiplicative factor, the difference is proportional to the signal itself. Without a reference, this is hard to distinguish from amplitude noise.
  • Presence head is 5.5 points above the majority baseline (66%). The head has learned to detect some distortions but not all. Distinguishing "signal-difference" from "noise-difference" at small magnitudes requires noise-normalized features that this version does not include.

Usage

Load from the Hub

import json
import torch
import numpy as np
from huggingface_hub import hf_hub_download
from liar_detector_v5 import (
    LiarDetectorModel,
    FeatureScaler,
    synth_true,
    apply_lie,
    extract_features,
)

REPO = "zeechimp/liar-detector-v5"

# 1. Load model and feature scaler
model = LiarDetectorModel.from_pretrained(REPO).eval()
scaler = FeatureScaler.load(hf_hub_download(REPO, "feature_extractor.json"))

# 2. Build a synthetic pair (A honest, B distorted)
rng = np.random.default_rng(0)
A_clean = synth_true(rng)
sx = float(np.std(A_clean)) + 1e-9
noise_std = rng.uniform(0.05, 0.15) * sx
A = A_clean + noise_std * rng.standard_normal(len(A_clean))
B = apply_lie(A_clean, rng, "offset") + noise_std * rng.standard_normal(len(A_clean))

# 3. Extract features, normalize, predict
x = scaler(A, B)
with torch.no_grad():
    out = model(features=torch.from_numpy(x).unsqueeze(0))

print("binary  :", out.binary_probs[0].tolist())    # [P(A lies), P(B lies)]
print("family  :", out.family_probs[0].tolist())    # 6-class distribution
print("presence:", out.presence_probs[0].tolist())  # [P(both honest), P(somebody lies)]

### Run the full CLI

The training, evaluation, and demo scripts live in the
[`example_usage.py`](example_usage.py) companion file. Download it from this
repo and run:

```bash
# Demo over 20 trials, reports per-head vote
python liar_detector_v5.py demo --model . --family offset

# Full ID/OOD evaluation
python liar_detector_v5.py eval --model .

Calibration

Temperatures fitted on the held-out validation split are stored in config.json and applied automatically inside forward():

Head Temperature
Binary 32.43
Family 19.09
Presence 0.53

The *_probs fields in LiarDetectorOutput are already calibrated. The raw *_logits are available if you want to apply your own scaling.

Citation

@misc{zeechimp2026liardetector,
  author       = {Sylv Q},
  title        = {Liar Detector v5: Self-Consistent-Liar Detection
                  with Multi-Head Classification and Per-Head Calibration},
  year         = {2026},
  publisher    = {Hugging Face},
  howpublished = {\url{https://huggingface.co/zeechimp/liar-detector-v5}}
}

Model Card Contact

For questions or issues, please open a discussion on the model page or contact zeechimp.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Evaluation results

  • Binary Accuracy (ID) on Synthetic Sinusoid Distortion Pairs
    self-reported
    0.882
  • Family Accuracy (ID) on Synthetic Sinusoid Distortion Pairs
    self-reported
    0.923
  • Presence Accuracy (ID) on Synthetic Sinusoid Distortion Pairs
    self-reported
    0.715
  • Binary Accuracy (OOD) on Synthetic Sinusoid Distortion Pairs
    self-reported
    0.928
  • Binary ECE (ID) on Synthetic Sinusoid Distortion Pairs
    self-reported
    0.091
  • Family ECE (ID) on Synthetic Sinusoid Distortion Pairs
    self-reported
    0.224
  • Presence ECE (ID) on Synthetic Sinusoid Distortion Pairs
    self-reported
    0.143