lyte-codes commited on
Commit
56ae190
·
verified ·
1 Parent(s): 267d9b2

Model card for 263702

Browse files
Files changed (1) hide show
  1. README.md +106 -154
README.md CHANGED
@@ -11,174 +11,126 @@ tags:
11
 
12
  # clockface
13
 
14
- Reads an analog clock and returns the time. Finds the dial first, then reads it.
15
- 11.5 MB of weights.
16
 
17
- > **It does not work on real photographs.** Off by 155.9 minutes on average
18
- > across 1,728 of them. Guessing at random is off by 180. On the synthetic
19
- > renders it trained on it is off by 4.8 minutes. I am publishing it because
20
- > that gap is the only interesting thing here.
21
 
22
- Two weight sets are here. `stage2_best.pt` was trained on renders alone.
23
- `v3mix_stage2.pt` was trained on rougher renders plus 1,528 real photographs,
24
- to see whether real data closed the gap. It did not, and both are published
25
- rather than only the flattering one.
26
 
27
- | weights | version | commit |
28
- |---|---|---|
29
- | `stage2_best.pt` + `stage1_best.pt` | 263601 | `82d3030` |
30
- | `v3mix_stage2.pt` | 263701 | `deed0ca` |
31
 
32
- Both were briefly published under the same version number. The release counter
33
- is derived from git tags, and with no tags it returns 01 forever, so two
34
- different models were stamped 263701. Tags now exist for both, and the stamp
35
- says UNTAGGED when none do rather than emitting a number that looks unique and
36
- is not.
37
 
38
- ## The metric
 
 
 
 
 
39
 
40
- A clock face is a ring of 720 minutes, and error is the shorter way round it. So
41
- 11:58 against 12:02 is 4 minutes, not 716. AM and PM do not exist, because a face
42
- cannot show them: 3:47 and 15:47 are the same reading.
43
 
44
- Random guessing scores 180 minutes. That is the bar.
 
 
45
 
46
- ## What it scores on real photographs
47
-
48
- 1,728 images with third-party labels.
49
-
50
- | | MAE | median | within 5 min | worse than 30 min |
51
- |---|---|---|---|---|
52
- | two stages | 155.9 min | 139.9 | 11.4% | 82.2% |
53
- | whole image, no locator | 151.8 min | 133.7 | 9.2% | 82.5% |
54
- | random guessing | 180 min | | | |
55
-
56
- Dropping the locator makes it slightly better, which tells you how much the
57
- locator is contributing.
58
-
59
- The images are 1,092 film stills from Marclay's *The Clock*, 506 photos from
60
- `moondream/analog-clock-benchmark`, and 131 cropped faces from
61
- `oliverj990/clock-faces-v1-times`. Most of those labels I have not checked. On
62
- the 106 I did check by hand against the image, the error is 148.5 minutes, so bad
63
- labels are not the excuse.
64
-
65
- Split by source, the story is obvious:
66
-
67
- | source | MAE | within 5 min |
68
- |---|---|---|
69
- | cropped faces, dial fills the frame | 87.8 min | 45.0% |
70
- | photographs of rooms and streets | 163.3 min | 9.3% |
71
-
72
- The closer a photo looks to my renders, the better it does. That is a domain gap
73
- and nothing subtler.
74
-
75
- ## What it scores on synthetic data
76
-
77
- 1,200 held-out renders.
78
-
79
- | | MAE | median | within 1 min | within 5 min |
80
- |---|---|---|---|---|
81
- | ground-truth crops | 4.82 min | 0.42 | 79.0% | 92.5% |
82
- | end to end | 5.07 min | 0.43 | 79.0% | 92.2% |
83
-
84
- Do not quote these. The renderer that made the labels also made the pixels, so
85
- the model is being marked by its own teacher. They are here for one reason: 4.8
86
- minutes on renders against 156 on photographs is the whole story.
87
-
88
- ## Does training on real photographs help? No.
89
-
90
- `v3mix_stage2.pt` was trained on a rougher synthetic corpus (sensor noise up to
91
- sigma 42, JPEG down to q25, occlusion, far wider hand geometry) plus 1,528 real
92
- photographs, holding out 200 for evaluation. 106 of those 200 have labels I
93
- checked by hand against the image; the rest are third-party and unverified.
94
-
95
- On the 106 labels I actually verified:
96
-
97
- | weights | trained on | MAE | within 5 min |
98
  |---|---|---|---|
99
- | stage2_best.pt | renders only | 148.5 min | 10.4% |
100
- | v3mix_stage2.pt | rougher renders + 1,528 real photos | 151.9 min | 12.3% |
101
- | random guessing | | 180 min | |
102
-
103
- Adding real photographs moved it from 148.5 to 151.9, which is to say nowhere.
104
 
105
- On the full 200-photo held-out set v3mix scores MAE 120.6 with 24.5% within
106
- five minutes, which looks better and mostly is not: that set contains the
107
- pre-cropped clock faces, which are the easy ones. Its own verified subset scores
108
- 151.9. Composition, not skill.
109
 
110
- Its synthetic score is 35.0 minutes against the older model's 4.8, because the
111
- renders it trained on are much harder. That is the intended direction and it
112
- still did not transfer.
113
 
114
- ## Where it goes wrong
115
 
116
- Of the 1,421 predictions off by more than half an hour, 720 would improve if you
117
- swapped the hour and minute hands. My generator drew hands from a narrow band of
118
- lengths and widths, and the model learned that band instead of learning which
119
- hand is longer. Real clocks do not agree with my band.
 
 
120
 
121
- The confidence signal does not survive the trip to real data. On renders it works
122
- well: keep the most confident 75% and error drops from 4.82 to 1.64 minutes. On
123
- photographs it does nothing, 155.9 to 153.5. The model is confidently wrong,
124
- which is worse than being uncertain.
125
-
126
- The locator finds a dial to within 1.5% of image width on renders. On a photo of
127
- a room where the clock is a smudge on a far wall, it does not.
128
-
129
- ## How it is built
130
-
131
- stage 1 DialLocator 554K params, 2.2 MB image -> centre, radius
132
- stage 2 PretrainedReader 2.3M params, 9.2 MB crop -> two angle distributions
133
-
134
- Stage 2 is MobileNetV3-Small pretrained on ImageNet, with two heads over 180
135
- angular bins each. Targets are von Mises bumps rather than one-hot, so being one
136
- bin out costs less than being ten out, and decoding uses a circular soft-argmax
137
- to get back below bin resolution.
138
-
139
- The same network without ImageNet weights never learns this at all. It sits at
140
- 2·ln(180), a flat distribution, for as long as you care to train it, while
141
- happily memorising 400 images. The pretrained initialisation is the difference
142
- between a model and a plateau.
143
-
144
- It predicts both hands and never the time directly. That costs nothing and buys
145
- two things. The hands disagreeing is a confidence signal. And decoding can work
146
- like a vernier scale, taking precision from the minute hand and only the hour
147
- count from the hour hand, which under noise is about 11 times more accurate than
148
- reading the hour hand alone.
149
-
150
- ## Using it
151
-
152
- import torch
153
- from twostage import PretrainedReader
154
- from model import decode_cls
155
-
156
- reader = PretrainedReader(bins=180, arch="mobilenet_v3_small")
157
- reader.load_state_dict(torch.load("stage2_best.pt")["model"])
158
- hour_logits, minute_logits = reader(crop) # 1x3x256x256, ImageNet normalised
159
- minutes, hour_only, disagreement, sharpness = decode_cls(hour_logits, minute_logits)
160
-
161
- `minutes` is the position on the 720-minute ring. `disagreement` is how far apart
162
- the two hands' readings are, in minutes.
163
-
164
- ## What it is for
165
-
166
- Round analog clock faces. Not digital-analog hybrids, subdials, chronographs or
167
- 24-hour faces.
168
-
169
- Do not use it for anything where a wrong time costs something. It is off by more
170
- than half an hour on 82% of real photographs.
171
-
172
- ## Reproducing
173
-
174
- git clone https://github.com/lyte-codes/clockface
175
- ./eval.py --self-test
176
- python3 predict.py --labels <labels> --root <dir> \
177
- --stage2 stage2_best.pt --stage1 stage1_best.pt --out preds.jsonl
178
- ./eval.py --labels <labels> --pred preds.jsonl
179
-
180
- `--self-test` runs 32 checks on the metric itself, including both directions
181
- across the 12 o'clock seam.
 
 
 
 
 
 
 
182
 
183
  ## Licence
184
 
 
11
 
12
  # clockface
13
 
14
+ Reads an analog clock and returns the time. Finds the dial with a
15
+ COCO-pretrained detector, then reads the crop.
16
 
17
+ > Best measured result on 200 held-out real photographs: **34.0 minutes**
18
+ > mean error, **1.2** median, **49.0%** within one minute.
19
+ > Guessing at random is off by 180 minutes. It is better than it was and it is
20
+ > not a working clock reader.
21
 
22
+ Version 263702, commit `efe5c96`. Reader backbone: `resnet50`.
 
 
 
23
 
24
+ ## Results on real photographs
 
 
 
25
 
26
+ 200 held-out photos, third-party labels, all read through the same pipeline:
27
+ COCO detector finds the dial, the reader reads the crop.
 
 
 
28
 
29
+ | reader | MAE | median | within 1 min | within 5 min | worse than 30 min |
30
+ |---|---|---|---|---|---|
31
+ | `resnet50` | 34.0 min | 1.2 | 49.0% | 73.0% | 22.0% |
32
+ | `resnet18` | 48.5 min | 1.5 | 44.5% | 60.5% | 34.5% |
33
+ | `mobilenet_v3_small` | 55.6 min | 2.5 | 34.5% | 59.0% | 36.5% |
34
+ | random guessing | 180 min | | | | |
35
 
36
+ ## What stage 1 is, and what it is not
 
 
37
 
38
+ Stage 1 is `fasterrcnn_mobilenet_v3_large_fpn` with COCO weights, run on CPU. It
39
+ is not trained here and not fine-tuned. Three ways of producing a crop, same
40
+ reader, same 200 photographs:
41
 
42
+ | crop | MAE | median | within 5 min |
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
43
  |---|---|---|---|
44
+ | whole image, no locator | 120.6 min | 88.3 | 22.5% |
45
+ | a dial locator trained on renders (2.2 MB) | 120.7 min | 67.9 | 25.0% |
46
+ | COCO detector (78 MB) | 73.6 min | 18.4 | 45.5% |
 
 
47
 
48
+ The locator trained on synthetic renders contributes nothing over having no
49
+ locator at all. It was removed. Distilling the detector into it would not have
50
+ helped: a student cannot beat its teacher, and the training would have run back
51
+ through the renderer whose domain gap is the thing being escaped.
52
 
53
+ Detection runs on CPU because sustained torchvision detection on an M2 trips the
54
+ GPU watchdog. On CPU it is 0.135s per image and finds a clock in 93-96% of real
55
+ photographs, having never seen a render.
56
 
57
+ ## How the two hands become one time
58
 
59
+ A clock's hands are geared: at time t the hour hand sits at t/720 of a turn and
60
+ the minute hand at (t mod 60)/60 of a turn, so one number fixes both. The reader
61
+ predicts a distribution for each hand, and the decoder searches the 720 minutes
62
+ the clock could be showing for the one that best explains both. It cannot return
63
+ a pair of angles that describes no time, and a confident minute hand can pull a
64
+ vague hour hand onto the right hour.
65
 
66
+ | decoder | MAE | median | within 5 min | worse than 30 min |
67
+ |---|---|---|---|---|
68
+ | each hand read on its own | 35.3 min | 1.4 | 69.0% | 25.5% |
69
+ | search over times the clock could show | 34.0 min | 1.2 | 73.0% | 22.0% |
70
+
71
+ Same weights, same crops, same photographs; only the decoding differs. The gain
72
+ is in readings that were nearly right already, which is why MAE barely moves --
73
+ MAE here is set by the readings that are wrong by hours, and those are wrong in
74
+ the image, not in the decoding.
75
+
76
+ Scoring the hands in each other's roles, to catch the ones read the wrong way
77
+ round, does not work and is off by default. On these photographs the exchanged
78
+ reading scored a median 0.06 better where the hands really were swapped and 0.11
79
+ worse everywhere else -- the two populations sit on top of each other. A swap is
80
+ not two good angles in the wrong slots; the model is confidently wrong about
81
+ both hands at once and they agree on a wrong but perfectly geared time.
82
+
83
+ ## When to trust a reading
84
+
85
+ The reader ranks its own answers, and the ranking is worth more than the
86
+ headline. Keeping only the half it is surest of:
87
+
88
+ ```
89
+ 200 readings. Keeping only the ones each signal is most sure of:
90
+
91
+ hand agreement
92
+ keep n MAE median <=5min
93
+ 100% 200 34.02 1.25 73.0%
94
+ 75% 150 26.80 1.00 79.3%
95
+ 50% 100 22.64 0.75 86.0%
96
+ 25% 50 30.12 0.75 82.0%
97
+ 10% 20 20.74 0.50 90.0%
98
+ rank correlation with error +0.352 (useful)
99
+
100
+ head sharpness
101
+ keep n MAE median <=5min
102
+ 100% 200 34.02 1.25 73.0%
103
+ 75% 150 16.58 0.75 84.0%
104
+ 50% 100 8.18 0.50 94.0%
105
+ 25% 50 1.91 0.50 98.0%
106
+ 10% 20 4.03 0.50 95.0%
107
+ rank correlation with error -0.506 (useful)
108
+
109
+ search margin
110
+ keep n MAE median <=5min
111
+ 100% 200 34.02 1.25 73.0%
112
+ 75% 150 18.39 0.75 84.0%
113
+ 50% 100 10.73 0.50 92.0%
114
+ 25% 50 6.75 0.50 96.0%
115
+ 10% 20 3.75 0.50 95.0%
116
+ rank correlation with error -0.496 (useful)
117
+
118
+ A caution on reading the low-coverage rows: at 25% of 200 photographs a row
119
+ rests on 50 readings, so small differences between signals there are noise.
120
+ ```
121
+
122
+ The signal this project set out to use for that was the disagreement between the
123
+ hands, which is free to compute. It works on renders and barely ranks anything
124
+ on photographs. The two that do work came out of the decoder instead.
125
+
126
+ ## Honest limits
127
+
128
+ The labels on these 200 photographs are third party. 106 of them were checked by
129
+ hand against the image; the rest were not. The photographs were collected for
130
+ other purposes and skew toward clocks that are already the subject of the frame.
131
+
132
+ There is no first-party test set. A number measured on somebody else's labels is
133
+ worth less than one measured on your own, and that gap is not closed here.
134
 
135
  ## Licence
136