Liar Detector v4

Model Summary

liar-detector-v4 is a lightweight, fully-connected neural network that detects whether one of two aligned 1-D signals is "lying" β€” i.e., distorted by a known manipulation family. Given two streams A and B, it simultaneously predicts:

  1. Which stream lies (binary head)
  2. What kind of lie it is (family head, 6 classes)
  3. Whether any lie is present at all (presence head)

The model is designed for research on self-consistent signal verification, sensor fusion, and reference-free anomaly detection. It operates on 44 hand-crafted features extracted from the (A, B) pair and requires no external reference signal.

Model Details

  • Developed by: Sylv Q (zeechimp)
  • Model type: Multi-head MLP (3 heads, shared trunk)
  • Language(s): N/A (numeric signal input)
  • License: Apache 2.0
  • Finetuned from: Not applicable (trained from scratch)
  • Repository: zeechimp/liar-detector-v4

Architecture

Component Detail
Input 44-dim feature vector extracted from (A, B)
Trunk Linear(44β†’96) β†’ ReLU β†’ Linear(96β†’64) β†’ ReLU
Binary head Linear(64β†’2) β€” which stream lies (0=A, 1=B)
Family head Linear(64β†’6) β€” distortion type
Presence head Linear(64β†’2) β€” is any lie present
Parameters ~11K
Framework PyTorch + πŸ€— Transformers

Feature Groups (44 total)

  • Raw moments β€” mean, log-std, skew, kurtosis for both streams (8)
  • Canonical reference β€” abs_mean_A - abs_mean_B, log_std_ratio_canonical (2)
  • Difference features β€” paired difference moments (6)
  • Residual statistics β€” moments, autocorrelation, cross-correlation (12)
  • Regression β€” residual slope / explained variance (2)
  • Reference-free signatures β€” quantization, stair-step, inversion, smoothness (10)
  • Uniqueness β€” unique-fraction ratios (4)

Intended Uses & Limitations

Direct Use

  • Research on signal integrity and self-consistency checking
  • Benchmarking reference-free anomaly detectors
  • Educational demonstrations of multi-head classification

Downstream Use

  • Sensor-fusion pipelines where one stream may be tampered
  • Audio/telemetry verification (with domain-specific retraining)

Out-of-Scope Use

  • Not a general-purpose "lie detector" for human speech or text
  • Not validated on real-world sensor data (trained on synthetic signals)
  • Not suitable for high-stakes decisions without extensive domain adaptation

Training Details

Training Data

Synthetic 1-D signals generated as sums of three sinusoids (frequencies 0.01–0.15 Hz, amplitudes 0.5–2.0), with additive Gaussian noise. Distortions are injected via the following families:

Family Type Description
offset ID Additive constant bias
scale ID Multiplicative gain (0.85–1.15Γ—)
saturation ID Hard clipping at Β±M
quantization ID Rounding to a grid
lag ID Integer time shift (1–4 samples)
harmonic ID Quadratic self-interaction term
deadzone OOD Zeroing of near-zero values
dropout OOD Random sample-and-hold
drift OOD Cumulative Gaussian random walk
inversion OOD Sign flip for small values

Training Procedure

  • Epochs: 200
  • Batch size: 64
  • Optimizer: AdamW
  • Learning rate: 3e-3
  • Loss: Masked cross-entropy (binary + family heads only train on single-lie pairs; presence head trains on all pairs)
  • Calibration: Temperature scaling fitted per head

Evaluation Results

In-Distribution (ID)

Condition Accuracy ECE
Binary (which stream lies) 86.1% 0.116
Family (distortion type) 91.2% 0.309
Presence (any lie) 72.9% 0.043

Per-Family Binary Accuracy (ID)

Family Accuracy
offset 96.3%
scale 57.5%
saturation 98.9%
quantization 68.6%
lag 94.6%
harmonic 98.8%

Out-of-Distribution (OOD)

Family Binary Accuracy
deadzone 87.2%
dropout 96.4%
drift 81.2%
inversion 90.0%

Overall OOD binary accuracy: 88.7% (ECE 0.137)

Sanity Checks

Scenario Presence Accuracy
Both honest 61.5%
Both lying 56.3%

Usage

import torch
import numpy as np
from transformers import AutoModel, AutoConfig
from huggingface_hub import hf_hub_download
import json

# Load model
model = AutoModel.from_pretrained(
    "zeechimp/liar-detector-v4",
    trust_remote_code=True
).eval()

# Load feature extractor normalization stats
fe_path = hf_hub_download("zeechimp/liar-detector-v4", "feature_extractor.json")
with open(fe_path) as f:
    fe_stats = json.load(f)
mu = np.array(fe_stats["mean"], dtype=np.float32)
sigma = np.array(fe_stats["std"], dtype=np.float32)

# Extract features (see feature_extractor.py in the repo)
# A, B are 1-D numpy arrays of equal length
features = extract_features(A, B)  # returns 44-dim vector
x = (features - mu) / sigma
x = torch.from_numpy(x).unsqueeze(0)  # (1, 44)

with torch.no_grad():
    out = model(features=x)

print("which lies:", out.binary_probs.argmax().item())   # 0=A, 1=B
print("family     :", out.family_probs.argmax().item())
print("presence   :", out.presence_probs.argmax().item()) # 1 = somebody lies
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support