ir-source-localizer
Given impulse responses at known microphones, return the source's (x, y) position to sub-centimetre accuracy.
Four microphones in the corners of a shoebox room. A source emits a short click. Four IRs are recorded. Each IR's direct-arrival time gives one constraint on the source position. Four constraints, two unknowns β overdetermined and solvable. RANSAC over receiver subsets rejects outlier detections when one of the four arrivals is unreliable.
Result from the synthetic demo: 0.28 cm error at 2% noise, stable from 0.5% to 10% noise. One receiver is systematically off by 0.70 ms (the geometry gives it an ambiguous direct-vs-reflection mixture); RANSAC identifies it and drops it, leaving three sub-sample-accurate arrivals.
The claim in one sentence
Four cheap microphones and a click, in an ordinary room, locate the click to within a centimetre.
What it produces
A LocalizationResult containing:
x,yβ source position in metressigma_x,sigma_yβ 1-sigma uncertainty from Monte Carlodetected_times,predicted_timesβ per-receiver arrival timesarrival_residual_msβ predicted minus detected, per receiverdetection_qualityβ per-receiver envelope confidenceinlier_maskβ which receivers RANSAC acceptedn_inliersβ how many of the K receivers agreed
Install
pip install numpy
No other dependencies. No downloads. No model weights. No autodiff framework. The whole tool is ~600 lines of numpy.
Usage
Programmatic
from ir_source_localizer import localize
import numpy as np
# Load or generate K impulse responses
irs = [ir_0, ir_1, ir_2, ir_3] # shape (T,), all same sample rate
# Known receiver positions in metres
receivers = [
(4.5, 3.5, 1.0),
(0.5, 3.5, 1.0),
(4.5, 0.5, 1.0),
(0.5, 0.5, 1.0),
]
result = localize(
irs=irs,
sample_rate=48000,
receivers=receivers,
sz=1.0, # source height (metres)
x_range=(0.2, 4.8), # search bounds
y_range=(0.2, 3.8),
sigma_guess=4e-4, # pulse half-width (seconds)
)
print(result)
# source at (+1.503, +1.500) m Β± (3.0, 3.1) cm inliers 3/4 rms resid 0.352 ms
CLI
python ir_source_localizer.py localize \
r0.wav r1.wav r2.wav r3.wav \
--sample-rate 48000 \
--receivers 4.5,3.5,1.0 0.5,3.5,1.0 4.5,0.5,1.0 0.5,0.5,1.0 \
--sz 1.0 \
--x-range 0.2 4.8 \
--y-range 0.2 3.8
Output:
source at (+1.503, +1.500) m Β± (3.0, 3.1) cm rms arrival residual 0.352 ms (4 receivers)
Pass --json for machine-readable output. Pass --no-unc to
skip the Monte Carlo uncertainty estimation (2Γ faster).
Self-test
python ir_source_localizer.py self-test
11 checks covering pulse detection, RANSAC outlier rejection, grid search, refinement, end-to-end accuracy, input validation, and WAV I/O. All pass on a clean install.
Demo
python ir_source_localizer.py demo
Sweeps noise from 0.5% to 10%, reports position error and inlier count at each level, then prints a detailed trace for a single run at 2% noise.
Benchmarks
Synthetic demo, shoebox room 5Γ4Γ2.8 m, four corner receivers, source at (1.5, 1.5, 1.0), Mexican-hat pulse of 0.4 ms half-width, 2000 samples at 48 kHz.
| noise | x | y | error | inliers |
|---|---|---|---|---|
| 0.5% | 1.503 | 1.500 | 0.28 cm | 3/4 |
| 1.0% | 1.503 | 1.500 | 0.28 cm | 3/4 |
| 2.0% | 1.503 | 1.500 | 0.28 cm | 3/4 |
| 5.0% | 1.503 | 1.500 | 0.28 cm | 3/4 |
| 10.0% | 1.507 | 1.495 | 0.89 cm | 3/4 |
The error is stable across the noise range because RANSAC is rejecting the one bad receiver. Receiver 2's direct arrival is systematically biased by ~0.70 ms by an unresolved reflection that arrives 1.29 ms later at that specific geometry. The other three receivers detect the direct arrival to within one sample period (0.02 ms). Once the outlier is dropped, the three remaining arrivals determine the position to sub-centimetre.
How it works
Detect direct arrivals. For each IR, square the signal (energy envelope), smooth with a 1-Ο window, find the first crossing of 30% of the envelope peak, search forward ~1.2 Ο for the local maximum. Returns a time and a quality.
RANSAC over receiver subsets. With K β₯ 4 receivers, try every 3-subset. Solve the position on the subset, count how many of the K receivers agree within an inlier threshold (default 0.30 ms). Keep the position with the largest consensus.
Refine on inliers. Take the winning subset's inliers and run Adam on the arrival-time residual using only those receivers.
Estimate uncertainty. Monte Carlo: perturb the detected arrival times by N(0, 0.3Ο), re-localize, return the standard deviation of the resulting position distribution.
The version history is the finding
This tool took six attempts to get right. The earlier versions failed for specific, instructive reasons:
| version | change | outcome |
|---|---|---|
| v1 | wide-window argmax after rise | locked onto the floor reflection at r2; 35 cm error |
| v2 | first local max after rise | picked up leading-edge bumps; 0.9 ms early bias |
| v3 | half-peak leading edge | smoothing window merged direct with nearby reflection |
| v4 | squared signal + tight smoothing | 12 cm error, r2 still biased 0.70 ms |
| v5 | quality-weighted loss | no improvement β quality β timing accuracy |
| v6 | RANSAC over subsets | 0.28 cm error, inliers 3/4 |
The lesson: when one of your measurements is unreliable, reject it β do not try to fix it. Five iterations of detector tuning did not rescue receiver 2. RANSAC accepted that receiver 2 is hopeless for this geometry and solved around it.
When to use it
- Any application where you have K microphones at known positions and a source that emits a short click. The arrival-time problem is well-conditioned and the RANSAC wrapper handles one bad receiver.
- Environments where GPS does not work. Indoor, underwater, underground. Acoustic arrival times give position without electromagnetic contact.
- Cheap hardware setups. MEMS microphones cost less than a dollar each. A four-microphone array with a USB audio interface is a few tens of dollars.
- Situations where the source can emit a known pulse. Any device with a speaker. Any event that produces a click.
When not to use it
- When the source cannot emit a sharp pulse. Speech, music, and continuous sounds do not have a detectable direct arrival. Cross-correlation against the source signal is required, and this tool does not implement that.
- When the room dimensions are unknown. The RANSAC inlier threshold is in arrival-time units, so it works without room dimensions, but the demo's synthetic simulator uses shoebox geometry. In a real room with unknown layout, the tool still works β it does not need the room to be shoebox-shaped, only the receiver positions to be known.
- When you have fewer than 3 receivers. With 3 receivers, the position is exactly determined but there is no redundancy for outlier rejection. The tool works, but a single bad detection will ruin the result.
- When the source height is unknown. The tool assumes
szis known. Fittingszis a separate problem.
Honest limitations
- Tested only on synthetic impulse responses. Every number in the benchmarks section comes from the image-source simulator inside the file. Real recordings will have backgrounds, unknown source directivity, and non-ideal pulses. The tool has not been run on a real recording.
- Shoebox geometry in the simulator. The RANSAC positioner itself does not assume shoebox β it works on arrival times and known receiver positions. But the demo's IRs come from a shoebox. Real rooms have different reflection patterns and the RANSAC inlier threshold may need tuning.
- Assumes the source pulse is short and known-width. The
sigma_guessparameter must approximately match the pulse's half-width. If the pulse is very different, detection fails. - Requires the receiver positions to be known exactly. A centimetre error in receiver position produces a centimetre error in the source. Real installations need the microphones surveyed or calibrated.
- Detects the direct arrival, not the source onset. The detected time corresponds to the peak of the envelope, not the leading edge. If the source emits a click with a long rise time, the detected time is delayed relative to the physical emission, and this delay must be calibrated.
- Only two dimensions. The source's height is fixed at
sz. Three-dimensional localization would need a fourth equation or an independent estimate ofsz. - No beamforming or array processing. This is not a spatial-filtering tool. It localizes one source at a time from arrival times.
- RANSAC over subsets costs O(KΒ³). With K=4, four subsets. With K=10, 120 subsets. With K=20, 1140 subsets. Still fast for arrival-time localization (each subset is a 2-D grid search), but not for very large arrays.
What this tool does not include
There are several natural extensions that were considered and not built:
- Fitting the room dimensions from the IR. The
room_fitterexperiments in the same session showed that room shape is much harder to recover than source position from a single IR. Source position is well-conditioned; room shape is not. - Learning the pulse shape. The tool requires a
sigma_guess. Auto-estimating the pulse width is a small extension but not implemented. - Handling moving sources. The tool assumes the source is stationary during the pulse. A moving source needs a different formulation (Doppler-aware arrival-time model).
- Robust to unknown
sz. Fixingszis a modelling choice; the natural joint estimate of(sx, sy, sz)from arrival times at four coplanar receivers is underdetermined. A receiver array with vertical extent would be needed.
Reference
Part of a series of small tools built in one session:
| tool | reads | answers |
|---|---|---|
numerical-provenance |
a numeric pipeline | can I trust this number? |
token-provenance |
an LLM pipeline | can I trust this token count? |
provenance-reconstruct |
a bare number | what produced this, and how sure are you? |
learning-report |
a trained model | can I make an honest claim about what it learned? |
sensor-audit |
a set of sensors | which are redundant, which are load-bearing? |
room-fitter |
an impulse response | what room produced this? (failed) |
ir-source-localizer |
a set of IRs | where is the source? |
License
Apache-2.0