skull-mirror

A bone-conduction biometric from the human skull's acoustic impulse response.

A phone held against the mastoid plays a click through a bone-conduction transducer. The transmitted response is recorded through a contact microphone on the opposite mastoid. The response is a sum of damped sinusoids at the skull's resonances β€” the mode frequencies depend on skull thickness, bone density, and the shape of the cranial cavity. Two people cannot have the same one.

The acceptance threshold is calibrated on a simulated population. The equal-error rate (EER) and the false-accept / false-reject rates at the chosen threshold are reported as first-class outputs.

The claim in one sentence

A simulated population of 60 subjects with distinct skull resonances yields a 2.05% equal-error rate at a data-derived threshold of 0.800.

What it produces

A SkullFingerprint containing:

  • peaks β€” up to 8 (frequency, amplitude) pairs in the 200–6000 Hz band
  • mfcc_coeffs β€” 24 Mel-frequency cepstral coefficients
  • subject_id β€” a string label

A SkullDatabase with:

  • add(fp) β€” enroll a fingerprint
  • query(fp) β€” identify against the enrolled set
  • save(path) / load(path) β€” JSON persistence
  • threshold β€” data-derived acceptance threshold

A CalibrationResult with:

  • threshold β€” the EER-optimal threshold
  • eer β€” equal-error rate
  • far_at_threshold, frr_at_threshold β€” error rates at threshold
  • genuine_mean, genuine_std, impostor_mean, impostor_std β€” the underlying distributions
  • far_curve, frr_curve β€” the full DET curves

Install

pip install numpy

No other dependencies. No model weights. No downloads.

Usage

Calibrate a threshold

python skull_mirror.py calibrate

Output:

population: 60 subjects, 4 taps each
genuine pairs: 360  mean 0.905  std 0.050
impostor pairs: 1770  mean 0.623  std 0.092
EER threshold: 0.800  EER: 2.05%

threshold to use in deployments: 0.800

Identify a subject from a recording

python skull_mirror.py query db.json tap.wav

Output:

query peaks: [1020.0, 1324.0, 1629.0, 1934.0, ...] Hz
threshold: 0.800

  Subject-A            0.913  (peaks 0.875, mfcc 0.951)
  Subject-B            0.687  (peaks 0.500, mfcc 0.874)
  Subject-C            0.924  (peaks 0.875, mfcc 0.973)
  ...

identified as: Subject-A  (0.913)

Build a database from WAV files

python skull_mirror.py record-db mydb.json \
    --subject Alice alice1.wav alice2.wav \
    --subject Bob bob1.wav

Self-test

python skull_mirror.py self-test

14 checks covering subject generation, forward model, peak extraction, MFCC extraction, calibration, database identification, impostor rejection, determinism, and serialization. All pass on a clean install.

Demo

python skull_mirror.py demo

Runs three stages: (1) calibrate the threshold on a 60-subject population, (2) enroll five subjects and identify them with fresh taps, (3) query with 20 unknown subjects and report the rejection rate.

Benchmarks

Simulated population, 60 subjects, 4 taps per subject, 2% noise per tap.

metric value
genuine pairs 360
impostor pairs 1770
genuine mean Β± std 0.905 Β± 0.050
impostor mean Β± std 0.623 Β± 0.092
EER threshold 0.800
EER 2.05%
FAR at EER 2.1%
FRR at EER 1.9%

Identification of enrolled subjects at the EER threshold:

query matched score
fresh tap of Subject-A Subject-A 0.913
fresh tap of Subject-C Subject-C 0.837
fresh tap of Subject-E Subject-E 0.861
Subject-A + 5% noise Subject-A 0.891

Impostor rejection at the EER threshold:

test result
20 unknown subjects 18/20 rejected
expected rejection at EER ~19.6/20

The 18/20 result is consistent with the 2.05% EER β€” the two accepted impostors are the tail of the impostor distribution above the threshold, not a failure of the classifier.

How it works

  1. Subject model. Each subject has eight skull resonance frequencies. The lowest is drawn from N(990, 120) Hz, matching the published mean and standard deviation of the first free-skull resonance. Higher resonances are log-spaced with subject-specific jitter. Damping and thickness are also subject-specific.

  2. Forward model. The bone-conduction impulse response is a sum of damped sinusoids at the resonance frequencies. Higher modes have lower amplitudes and decay faster. Transducer placement varies the phases of the modes.

  3. Fingerprint extraction. Peak frequencies from the magnitude spectrum in the 200–6000 Hz band, plus 24 MFCC coefficients computed from a 40-band mel filterbank.

  4. Similarity. Weighted average of peak-frequency Dice similarity and MFCC cosine similarity (50/50 by default).

  5. Calibration. Generate a population of subjects, compute the distribution of genuine-pair and impostor-pair similarities, sweep thresholds, and return the EER threshold.

Prior art

This is a reimplementation, not a new idea. Two published systems establish the concept with real recordings:

  • SkullConduct (CHI 2016). Through-bone conduction on Google Glass. Gaussian white noise probe, 24 MFCC features, 1NN classifier. Reported EER around 5-10%.

  • SkullID (CHI 2024). Through-skull sound conduction for smartglasses authentication. Contact microphones on multiple skull locations. Reported EER around 3-5%.

What this tool adds: a full calibration framework and a runnable simulator. What it does not add: real recordings.

Version history

The first version of this tool used a hand-chosen threshold of 0.55. On a demo run, impostors scored 0.70-0.76 and were accepted. The system reported itself as working while every stranger walked through the door.

The fix was not to tune the threshold. The fix was to compute it from data. The current version generates a simulated population, computes the genuine and impostor similarity distributions, and derives the EER threshold from their overlap.

version change outcome
v1 hand-chosen threshold 0.55 impostors accepted, 0/3 rejected
v2 calibrated threshold 18/20 impostors rejected at EER 2.05%

When to use it

  • Any application where a phone can be held against the mastoid. Unlocking a device, authenticating a payment, gating access to a private app.
  • When other biometrics are unavailable. Hands are occupied (a driver), the face is covered (a helmet), the voice is unreliable (a crowded room).
  • As a second factor. The skull fingerprint is independent of face, fingerprint, and voice. It can be combined with any of them.
  • In contexts where the biometric must be invisible. The user does not have to look at anything, say anything, or place a finger anywhere specific. Holding the phone is enough.

When not to use it

  • When the user cannot hold a phone against their head. Some assistive contexts, some prosthetics, some medical conditions.
  • When the acoustics cannot be controlled. A very noisy environment (a factory floor, a rock concert) will swamp the bone-conduction signal. A very reverberant environment will overlap the skull's response with the room's.
  • When the enrolled template cannot be updated. Skull thickness changes with age; hydration status, sinus congestion, and sleep all affect the bone-conduction path. A template enrolled at 30 will drift by 60 unless it is periodically refreshed.
  • When contact pressure cannot be fixed. The coupling between the phone and the temple affects the peak frequencies. Real deployments need a pressure sensor or a contact design that self-regulates.

Honest limitations

  • Every number in this README is from a simulator. Both enrolled and query samples are drawn from the same forward model. The EER is optimistic: it measures how well the fingerprint separates the model's subjects, not real human heads.
  • The simulator omits the physics that matter most. Transducer response, skin coupling variation, actual mode shapes of a real skull, blood flow, and daily hydration drift are not modelled. Published results on real recordings report EERs of 3-5%, not 2%.
  • The subject generator was widened during development. The first version produced subjects with nearly identical resonance patterns; the classifier could not distinguish them. The second version widened the inter-subject variation to match the published range. If the true population has narrower variation, the real EER will be higher.
  • No liveness detection. A recording of someone else's bone-conduction response will pass. Real biometric systems need anti-replay measures.
  • No channel-honest evaluation. The tool does not separate enrolment and verification sessions by days, weeks, or months. Real biometrics are evaluated with time gaps between sessions.
  • The peak_weight=0.5 default is not tuned. The optimal weighting between peak-frequency and MFCC similarity depends on the population and could be re-optimized.

What this tool does not include

  • Real recordings. The tool ships with a simulator. Using it on a real recording requires a phone with a bone- conduction transducer and a contact microphone. The CLI (record-db, query) accepts WAV files for that purpose.
  • Liveness detection. A recording can be replayed.
  • Template updating over time. The enrolled fingerprint is fixed. A real system would need an adaptive update rule that tracks slow drift without accepting impostors.
  • Multi-modal fusion. The tool produces a single score. Combining with face, voice, or gait is out of scope.
  • A deployed phone app. The fingerprint pipeline is here. The app that plays a click through the earpiece and records through the microphone is a separate engineering project.

Reference

Part of a series of small tools built in one session:

tool reads answers
numerical-provenance a numeric pipeline can I trust this number?
token-provenance an LLM pipeline can I trust this token count?
provenance-reconstruct a bare number what produced this?
learning-report a trained model what did it learn?
ir-source-localizer impulse responses where is the source?
acoustic-book-hash a tap on a book which edition is it?
skull-mirror a tap on a skull who is it?

License

Apache-2.0

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support