skull-mirror
A bone-conduction biometric from the human skull's acoustic impulse response.
A phone held against the mastoid plays a click through a bone-conduction transducer. The transmitted response is recorded through a contact microphone on the opposite mastoid. The response is a sum of damped sinusoids at the skull's resonances β the mode frequencies depend on skull thickness, bone density, and the shape of the cranial cavity. Two people cannot have the same one.
The acceptance threshold is calibrated on a simulated population. The equal-error rate (EER) and the false-accept / false-reject rates at the chosen threshold are reported as first-class outputs.
The claim in one sentence
A simulated population of 60 subjects with distinct skull resonances yields a 2.05% equal-error rate at a data-derived threshold of 0.800.
What it produces
A SkullFingerprint containing:
peaksβ up to 8 (frequency, amplitude) pairs in the 200β6000 Hz bandmfcc_coeffsβ 24 Mel-frequency cepstral coefficientssubject_idβ a string label
A SkullDatabase with:
add(fp)β enroll a fingerprintquery(fp)β identify against the enrolled setsave(path)/load(path)β JSON persistencethresholdβ data-derived acceptance threshold
A CalibrationResult with:
thresholdβ the EER-optimal thresholdeerβ equal-error ratefar_at_threshold,frr_at_thresholdβ error rates at thresholdgenuine_mean,genuine_std,impostor_mean,impostor_stdβ the underlying distributionsfar_curve,frr_curveβ the full DET curves
Install
pip install numpy
No other dependencies. No model weights. No downloads.
Usage
Calibrate a threshold
python skull_mirror.py calibrate
Output:
population: 60 subjects, 4 taps each
genuine pairs: 360 mean 0.905 std 0.050
impostor pairs: 1770 mean 0.623 std 0.092
EER threshold: 0.800 EER: 2.05%
threshold to use in deployments: 0.800
Identify a subject from a recording
python skull_mirror.py query db.json tap.wav
Output:
query peaks: [1020.0, 1324.0, 1629.0, 1934.0, ...] Hz
threshold: 0.800
Subject-A 0.913 (peaks 0.875, mfcc 0.951)
Subject-B 0.687 (peaks 0.500, mfcc 0.874)
Subject-C 0.924 (peaks 0.875, mfcc 0.973)
...
identified as: Subject-A (0.913)
Build a database from WAV files
python skull_mirror.py record-db mydb.json \
--subject Alice alice1.wav alice2.wav \
--subject Bob bob1.wav
Self-test
python skull_mirror.py self-test
14 checks covering subject generation, forward model, peak extraction, MFCC extraction, calibration, database identification, impostor rejection, determinism, and serialization. All pass on a clean install.
Demo
python skull_mirror.py demo
Runs three stages: (1) calibrate the threshold on a 60-subject population, (2) enroll five subjects and identify them with fresh taps, (3) query with 20 unknown subjects and report the rejection rate.
Benchmarks
Simulated population, 60 subjects, 4 taps per subject, 2% noise per tap.
| metric | value |
|---|---|
| genuine pairs | 360 |
| impostor pairs | 1770 |
| genuine mean Β± std | 0.905 Β± 0.050 |
| impostor mean Β± std | 0.623 Β± 0.092 |
| EER threshold | 0.800 |
| EER | 2.05% |
| FAR at EER | 2.1% |
| FRR at EER | 1.9% |
Identification of enrolled subjects at the EER threshold:
| query | matched | score |
|---|---|---|
| fresh tap of Subject-A | Subject-A | 0.913 |
| fresh tap of Subject-C | Subject-C | 0.837 |
| fresh tap of Subject-E | Subject-E | 0.861 |
| Subject-A + 5% noise | Subject-A | 0.891 |
Impostor rejection at the EER threshold:
| test | result |
|---|---|
| 20 unknown subjects | 18/20 rejected |
| expected rejection at EER | ~19.6/20 |
The 18/20 result is consistent with the 2.05% EER β the two accepted impostors are the tail of the impostor distribution above the threshold, not a failure of the classifier.
How it works
Subject model. Each subject has eight skull resonance frequencies. The lowest is drawn from N(990, 120) Hz, matching the published mean and standard deviation of the first free-skull resonance. Higher resonances are log-spaced with subject-specific jitter. Damping and thickness are also subject-specific.
Forward model. The bone-conduction impulse response is a sum of damped sinusoids at the resonance frequencies. Higher modes have lower amplitudes and decay faster. Transducer placement varies the phases of the modes.
Fingerprint extraction. Peak frequencies from the magnitude spectrum in the 200β6000 Hz band, plus 24 MFCC coefficients computed from a 40-band mel filterbank.
Similarity. Weighted average of peak-frequency Dice similarity and MFCC cosine similarity (50/50 by default).
Calibration. Generate a population of subjects, compute the distribution of genuine-pair and impostor-pair similarities, sweep thresholds, and return the EER threshold.
Prior art
This is a reimplementation, not a new idea. Two published systems establish the concept with real recordings:
SkullConduct (CHI 2016). Through-bone conduction on Google Glass. Gaussian white noise probe, 24 MFCC features, 1NN classifier. Reported EER around 5-10%.
SkullID (CHI 2024). Through-skull sound conduction for smartglasses authentication. Contact microphones on multiple skull locations. Reported EER around 3-5%.
What this tool adds: a full calibration framework and a runnable simulator. What it does not add: real recordings.
Version history
The first version of this tool used a hand-chosen threshold of 0.55. On a demo run, impostors scored 0.70-0.76 and were accepted. The system reported itself as working while every stranger walked through the door.
The fix was not to tune the threshold. The fix was to compute it from data. The current version generates a simulated population, computes the genuine and impostor similarity distributions, and derives the EER threshold from their overlap.
| version | change | outcome |
|---|---|---|
| v1 | hand-chosen threshold 0.55 | impostors accepted, 0/3 rejected |
| v2 | calibrated threshold | 18/20 impostors rejected at EER 2.05% |
When to use it
- Any application where a phone can be held against the mastoid. Unlocking a device, authenticating a payment, gating access to a private app.
- When other biometrics are unavailable. Hands are occupied (a driver), the face is covered (a helmet), the voice is unreliable (a crowded room).
- As a second factor. The skull fingerprint is independent of face, fingerprint, and voice. It can be combined with any of them.
- In contexts where the biometric must be invisible. The user does not have to look at anything, say anything, or place a finger anywhere specific. Holding the phone is enough.
When not to use it
- When the user cannot hold a phone against their head. Some assistive contexts, some prosthetics, some medical conditions.
- When the acoustics cannot be controlled. A very noisy environment (a factory floor, a rock concert) will swamp the bone-conduction signal. A very reverberant environment will overlap the skull's response with the room's.
- When the enrolled template cannot be updated. Skull thickness changes with age; hydration status, sinus congestion, and sleep all affect the bone-conduction path. A template enrolled at 30 will drift by 60 unless it is periodically refreshed.
- When contact pressure cannot be fixed. The coupling between the phone and the temple affects the peak frequencies. Real deployments need a pressure sensor or a contact design that self-regulates.
Honest limitations
- Every number in this README is from a simulator. Both enrolled and query samples are drawn from the same forward model. The EER is optimistic: it measures how well the fingerprint separates the model's subjects, not real human heads.
- The simulator omits the physics that matter most. Transducer response, skin coupling variation, actual mode shapes of a real skull, blood flow, and daily hydration drift are not modelled. Published results on real recordings report EERs of 3-5%, not 2%.
- The subject generator was widened during development. The first version produced subjects with nearly identical resonance patterns; the classifier could not distinguish them. The second version widened the inter-subject variation to match the published range. If the true population has narrower variation, the real EER will be higher.
- No liveness detection. A recording of someone else's bone-conduction response will pass. Real biometric systems need anti-replay measures.
- No channel-honest evaluation. The tool does not separate enrolment and verification sessions by days, weeks, or months. Real biometrics are evaluated with time gaps between sessions.
- The
peak_weight=0.5default is not tuned. The optimal weighting between peak-frequency and MFCC similarity depends on the population and could be re-optimized.
What this tool does not include
- Real recordings. The tool ships with a simulator.
Using it on a real recording requires a phone with a bone-
conduction transducer and a contact microphone. The CLI
(
record-db,query) accepts WAV files for that purpose. - Liveness detection. A recording can be replayed.
- Template updating over time. The enrolled fingerprint is fixed. A real system would need an adaptive update rule that tracks slow drift without accepting impostors.
- Multi-modal fusion. The tool produces a single score. Combining with face, voice, or gait is out of scope.
- A deployed phone app. The fingerprint pipeline is here. The app that plays a click through the earpiece and records through the microphone is a separate engineering project.
Reference
Part of a series of small tools built in one session:
| tool | reads | answers |
|---|---|---|
numerical-provenance |
a numeric pipeline | can I trust this number? |
token-provenance |
an LLM pipeline | can I trust this token count? |
provenance-reconstruct |
a bare number | what produced this? |
learning-report |
a trained model | what did it learn? |
ir-source-localizer |
impulse responses | where is the source? |
acoustic-book-hash |
a tap on a book | which edition is it? |
skull-mirror |
a tap on a skull | who is it? |
License
Apache-2.0