lyte-codes commited on
Commit
02be520
·
verified ·
1 Parent(s): 98f945c

Rewrite the model card in plainer language

Browse files
Files changed (1) hide show
  1. README.md +76 -63
README.md CHANGED
@@ -11,122 +11,135 @@ tags:
11
 
12
  # clockface
13
 
14
- Reads an analog clock face and returns the time. Two stages: find the dial,
15
- then read it. 11.5 MB of weights in total.
16
 
17
- > **This model does not work well on real photographs.** On 1,728 real photos it
18
- > scores **155.9 minutes** mean error, against **180 minutes for guessing at
19
- > random**. It scores 4.8 minutes on the synthetic data it was trained on. It is
20
- > published because the gap between those two numbers is the useful part.
21
 
22
- Version 263701. Trained on synth v2, revision `main` of the dataset above.
23
 
24
  ## The metric
25
 
26
- A clock face is a ring of 720 minutes. Error is the shorter way round it, so it
27
- lies in [0, 360] and 11:58 against 12:02 is 4 minutes, not 716. AM/PM does not
28
- exist: a face cannot show it, so 3:47 and 15:47 are the same reading.
29
 
30
- Guessing uniformly at random scores 180 minutes. That is the number to beat, and
31
- it is the number this model is measured against.
32
 
33
- ## Results
34
 
35
- ### Real photographs (1,728 images, third-party labels)
36
 
37
  | | MAE | median | within 5 min | worse than 30 min |
38
  |---|---|---|---|---|
39
- | two-stage | 155.9 min | 139.9 | 11.4% | 82.2% |
40
  | whole image, no locator | 151.8 min | 133.7 | 9.2% | 82.5% |
41
- | **random guessing** | **180 min** | — | — | — |
42
 
43
- Sources: 1,092 film stills from Marclay's *The Clock*, 506 photographs from
44
- `moondream/analog-clock-benchmark`, 131 cropped faces from
45
- `oliverj990/clock-faces-v1-times`. Their labels are third party and mostly
46
- unverified; restricted to the 106 labels checked by hand against the image, the
47
- result is MAE 148.5, so label noise does not explain it.
48
 
49
- By source, the pattern is a domain gap and not much else:
 
 
 
 
 
 
50
 
51
  | source | MAE | within 5 min |
52
  |---|---|---|
53
- | cropped clock faces (dial fills frame) | 87.8 min | 45.0% |
54
- | photographs of scenes | 163.3 min | 9.3% |
 
 
 
55
 
56
- The closer an image is to the synthetic framing, the better it does.
57
 
58
- ### Synthetic validation (1,200 held-out renders) — training telemetry only
59
 
60
  | | MAE | median | within 1 min | within 5 min |
61
  |---|---|---|---|---|
62
  | ground-truth crops | 4.82 min | 0.42 | 79.0% | 92.5% |
63
  | end to end | 5.07 min | 0.43 | 79.0% | 92.2% |
64
 
65
- **These numbers are not a result.** The model is being scored by the same
66
- renderer that produced its labels. They are here to show the size of the
67
- sim-to-real gap, which is the point: 4.8 minutes on renders, 156 on photographs.
 
 
68
 
69
- ## What goes wrong
 
 
 
70
 
71
- - **Hand confusion.** 720 of the 1,421 gross failures would improve if the hour
72
- and minute hands were swapped. The generator used a narrow band of hand
73
- lengths and widths, and the model learned that convention rather than learning
74
- which hand is longer.
75
- - **The confidence signal does not survive.** The two hands' disagreement is a
76
- good confidence signal on synthetic data (MAE 4.82 falls to 1.64 when keeping
77
- the most confident 75%) and is inert on real photographs (155.9 to 153.5). The
78
- model is confidently wrong, which is worse than being uncertain.
79
- - **Small clocks.** The locator finds a dial to 1.5% of image width on renders
80
- and fails on wide scenes where the clock is a small part of the frame.
81
 
82
- ## Architecture
 
83
 
84
- stage 1 DialLocator 554K params, 2.2 MB image -> dial centre, radius
85
- stage 2 PretrainedReader 2.3M params, 9.2 MB crop -> two angle distributions
86
 
87
- Stage 2 is a MobileNetV3-Small backbone pretrained on ImageNet, with two heads
88
- predicting a distribution over 180 angular bins for each hand. Targets are von
89
- Mises bumps rather than one-hot, and decoding takes a circular soft-argmax.
90
 
91
- An identical from-scratch network never leaves chance on this task: it sits at
92
- 2*ln(180), a uniform distribution, however long it trains, while overfitting 400
93
- images perfectly. ImageNet initialisation is what makes it learn at all.
 
94
 
95
- Both hands are predicted, never the time directly, so their disagreement gives a
96
- confidence signal for free. Decoding is a vernier: the minute hand supplies
97
- precision, the hour hand picks which hour. Measured against reading the hour hand
98
- alone, that is about 11x more accurate under noise.
99
 
100
- ## Usage
 
 
 
 
 
 
101
 
102
  import torch
103
- from twostage import DialLocator, PretrainedReader, crop_dial
104
  from model import decode_cls
105
 
106
  reader = PretrainedReader(bins=180, arch="mobilenet_v3_small")
107
  reader.load_state_dict(torch.load("stage2_best.pt")["model"])
108
- hour_logits, minute_logits = reader(crop) # crop: 1x3x256x256, ImageNet norm
109
  minutes, hour_only, disagreement, sharpness = decode_cls(hour_logits, minute_logits)
110
 
111
- `minutes` is the position on the 720-minute ring; `disagreement` is how far apart
112
  the two hands' readings are, in minutes.
113
 
114
- ## Intended use
115
 
116
- Reading round analog clock faces. Out of scope: digital-analog hybrids, subdials,
117
- chronographs, 24-hour faces.
118
 
119
- **Not suitable for any use where a wrong time carries a cost.** It is wrong by
120
- more than half an hour on 82% of real photographs.
121
 
122
  ## Reproducing
123
 
124
  git clone https://github.com/lyte-codes/clockface
125
- ./eval.py --self-test # 32/32, checks the metric
126
  python3 predict.py --labels <labels> --root <dir> \
127
  --stage2 stage2_best.pt --stage1 stage1_best.pt --out preds.jsonl
128
  ./eval.py --labels <labels> --pred preds.jsonl
129
 
 
 
 
130
  ## Licence
131
 
132
  Apache 2.0. The synthetic dataset is CC BY 4.0.
 
11
 
12
  # clockface
13
 
14
+ Reads an analog clock and returns the time. Finds the dial first, then reads it.
15
+ 11.5 MB of weights.
16
 
17
+ > **It does not work on real photographs.** On 1,728 real photos it is off by
18
+ > 155.9 minutes on average. Guessing at random is off by 180. On the synthetic
19
+ > renders it trained on, it is off by 4.8 minutes. I am publishing it because
20
+ > that gap is the only interesting thing here.
21
 
22
+ Version 263701, trained on synth v2.
23
 
24
  ## The metric
25
 
26
+ A clock face is a ring of 720 minutes, and error is the shorter way round it. So
27
+ 11:58 against 12:02 is 4 minutes, not 716. AM and PM do not exist, because a face
28
+ cannot show them: 3:47 and 15:47 are the same reading.
29
 
30
+ Random guessing scores 180 minutes. That is the bar.
 
31
 
32
+ ## What it scores on real photographs
33
 
34
+ 1,728 images with third-party labels.
35
 
36
  | | MAE | median | within 5 min | worse than 30 min |
37
  |---|---|---|---|---|
38
+ | two stages | 155.9 min | 139.9 | 11.4% | 82.2% |
39
  | whole image, no locator | 151.8 min | 133.7 | 9.2% | 82.5% |
40
+ | random guessing | 180 min | | | |
41
 
42
+ Dropping the locator makes it slightly better, which tells you how much the
43
+ locator is contributing.
 
 
 
44
 
45
+ The images are 1,092 film stills from Marclay's *The Clock*, 506 photos from
46
+ `moondream/analog-clock-benchmark`, and 131 cropped faces from
47
+ `oliverj990/clock-faces-v1-times`. Most of those labels I have not checked. On
48
+ the 106 I did check by hand against the image, the error is 148.5 minutes, so bad
49
+ labels are not the excuse.
50
+
51
+ Split by source, the story is obvious:
52
 
53
  | source | MAE | within 5 min |
54
  |---|---|---|
55
+ | cropped faces, dial fills the frame | 87.8 min | 45.0% |
56
+ | photographs of rooms and streets | 163.3 min | 9.3% |
57
+
58
+ The closer a photo looks to my renders, the better it does. That is a domain gap
59
+ and nothing subtler.
60
 
61
+ ## What it scores on synthetic data
62
 
63
+ 1,200 held-out renders.
64
 
65
  | | MAE | median | within 1 min | within 5 min |
66
  |---|---|---|---|---|
67
  | ground-truth crops | 4.82 min | 0.42 | 79.0% | 92.5% |
68
  | end to end | 5.07 min | 0.43 | 79.0% | 92.2% |
69
 
70
+ Do not quote these. The renderer that made the labels also made the pixels, so
71
+ the model is being marked by its own teacher. They are here for one reason: 4.8
72
+ minutes on renders against 156 on photographs is the whole story.
73
+
74
+ ## Where it goes wrong
75
 
76
+ Of the 1,421 predictions off by more than half an hour, 720 would improve if you
77
+ swapped the hour and minute hands. My generator drew hands from a narrow band of
78
+ lengths and widths, and the model learned that band instead of learning which
79
+ hand is longer. Real clocks do not agree with my band.
80
 
81
+ The confidence signal does not survive the trip to real data. On renders it works
82
+ well: keep the most confident 75% and error drops from 4.82 to 1.64 minutes. On
83
+ photographs it does nothing, 155.9 to 153.5. The model is confidently wrong,
84
+ which is worse than being uncertain.
 
 
 
 
 
 
85
 
86
+ The locator finds a dial to within 1.5% of image width on renders. On a photo of
87
+ a room where the clock is a smudge on a far wall, it does not.
88
 
89
+ ## How it is built
 
90
 
91
+ stage 1 DialLocator 554K params, 2.2 MB image -> centre, radius
92
+ stage 2 PretrainedReader 2.3M params, 9.2 MB crop -> two angle distributions
 
93
 
94
+ Stage 2 is MobileNetV3-Small pretrained on ImageNet, with two heads over 180
95
+ angular bins each. Targets are von Mises bumps rather than one-hot, so being one
96
+ bin out costs less than being ten out, and decoding uses a circular soft-argmax
97
+ to get back below bin resolution.
98
 
99
+ The same network without ImageNet weights never learns this at all. It sits at
100
+ 2·ln(180), a flat distribution, for as long as you care to train it, while
101
+ happily memorising 400 images. The pretrained initialisation is the difference
102
+ between a model and a plateau.
103
 
104
+ It predicts both hands and never the time directly. That costs nothing and buys
105
+ two things. The hands disagreeing is a confidence signal. And decoding can work
106
+ like a vernier scale, taking precision from the minute hand and only the hour
107
+ count from the hour hand, which under noise is about 11 times more accurate than
108
+ reading the hour hand alone.
109
+
110
+ ## Using it
111
 
112
  import torch
113
+ from twostage import PretrainedReader
114
  from model import decode_cls
115
 
116
  reader = PretrainedReader(bins=180, arch="mobilenet_v3_small")
117
  reader.load_state_dict(torch.load("stage2_best.pt")["model"])
118
+ hour_logits, minute_logits = reader(crop) # 1x3x256x256, ImageNet normalised
119
  minutes, hour_only, disagreement, sharpness = decode_cls(hour_logits, minute_logits)
120
 
121
+ `minutes` is the position on the 720-minute ring. `disagreement` is how far apart
122
  the two hands' readings are, in minutes.
123
 
124
+ ## What it is for
125
 
126
+ Round analog clock faces. Not digital-analog hybrids, subdials, chronographs or
127
+ 24-hour faces.
128
 
129
+ Do not use it for anything where a wrong time costs something. It is off by more
130
+ than half an hour on 82% of real photographs.
131
 
132
  ## Reproducing
133
 
134
  git clone https://github.com/lyte-codes/clockface
135
+ ./eval.py --self-test
136
  python3 predict.py --labels <labels> --root <dir> \
137
  --stage2 stage2_best.pt --stage1 stage1_best.pt --out preds.jsonl
138
  ./eval.py --labels <labels> --pred preds.jsonl
139
 
140
+ `--self-test` runs 32 checks on the metric itself, including both directions
141
+ across the 12 o'clock seam.
142
+
143
  ## Licence
144
 
145
  Apache 2.0. The synthetic dataset is CC BY 4.0.