VisionState wheel reader
The model VisionState โ a Home Assistant app that turns camera images into sensors, locally โ reads mechanical counters with (water and gas meters with rolling digit wheels).
It looks at one wheel and tells how far it has turned: a distribution over 100 positions around the wheel (0.0, 0.1, โฆ 9.9), so a wheel half way between two digits is read as such. VisionState splits the counter into one cell per wheel and picks the most likely value whose wheel positions fit together (a wheel only turns while the wheel to its right goes from 9 to 0).
| File | Size | Input | Output |
|---|---|---|---|
wheels-v1.onnx |
2.2 MB, 565 k parameters | cells: N ร 1 ร 48 ร 32, float, 0โฆ1 |
logits: N ร 100 |
Preprocessing of a cell: grey from each pixel's darkest colour channel (coloured wheels look like black ones), resized to 32 ร 48 (bilinear), autocontrast with 1 % cut-off, divided by 255. About 2 ms for eight wheels on one CPU thread.
Training data
All of it may be used for anything:
- Drawn wheels (600 000 cells) with fonts from Google Fonts (SIL Open Font License): many colours, lighting, blur, noise and JPEG artefacts, with exactly known positions.
- Word-Wheel Water Meter Dataset (Scientific Data, 2026), recognition crops โ CC0.
- Checked readings shared by VisionState users in GitHub issues #32 and #40 โ CC0.
Training code, data sources and how to make the model again:
tools/wheelreader.
Results
Whole reading exact (last wheel rounded), read with VisionState's decoding:
- A water meter with red decimal wheels seen by an ESP32 camera with its flash, held out from training: 17 of 17.
- Two water meters never seen in training: 30โ33 of 35 and 15โ18 of 18, no value too high.
- Word-Wheel test set (2,400 photos): 97.7 % when the photo is the right way up.
Not for pointer dials. Licence: Apache-2.0.