Model card for 263702
Browse files
README.md
CHANGED
|
@@ -11,174 +11,126 @@ tags:
|
|
| 11 |
|
| 12 |
# clockface
|
| 13 |
|
| 14 |
-
Reads an analog clock and returns the time. Finds the dial
|
| 15 |
-
|
| 16 |
|
| 17 |
-
>
|
| 18 |
-
>
|
| 19 |
-
>
|
| 20 |
-
>
|
| 21 |
|
| 22 |
-
|
| 23 |
-
`v3mix_stage2.pt` was trained on rougher renders plus 1,528 real photographs,
|
| 24 |
-
to see whether real data closed the gap. It did not, and both are published
|
| 25 |
-
rather than only the flattering one.
|
| 26 |
|
| 27 |
-
|
| 28 |
-
|---|---|---|
|
| 29 |
-
| `stage2_best.pt` + `stage1_best.pt` | 263601 | `82d3030` |
|
| 30 |
-
| `v3mix_stage2.pt` | 263701 | `deed0ca` |
|
| 31 |
|
| 32 |
-
|
| 33 |
-
|
| 34 |
-
different models were stamped 263701. Tags now exist for both, and the stamp
|
| 35 |
-
says UNTAGGED when none do rather than emitting a number that looks unique and
|
| 36 |
-
is not.
|
| 37 |
|
| 38 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 39 |
|
| 40 |
-
|
| 41 |
-
11:58 against 12:02 is 4 minutes, not 716. AM and PM do not exist, because a face
|
| 42 |
-
cannot show them: 3:47 and 15:47 are the same reading.
|
| 43 |
|
| 44 |
-
|
|
|
|
|
|
|
| 45 |
|
| 46 |
-
|
| 47 |
-
|
| 48 |
-
1,728 images with third-party labels.
|
| 49 |
-
|
| 50 |
-
| | MAE | median | within 5 min | worse than 30 min |
|
| 51 |
-
|---|---|---|---|---|
|
| 52 |
-
| two stages | 155.9 min | 139.9 | 11.4% | 82.2% |
|
| 53 |
-
| whole image, no locator | 151.8 min | 133.7 | 9.2% | 82.5% |
|
| 54 |
-
| random guessing | 180 min | | | |
|
| 55 |
-
|
| 56 |
-
Dropping the locator makes it slightly better, which tells you how much the
|
| 57 |
-
locator is contributing.
|
| 58 |
-
|
| 59 |
-
The images are 1,092 film stills from Marclay's *The Clock*, 506 photos from
|
| 60 |
-
`moondream/analog-clock-benchmark`, and 131 cropped faces from
|
| 61 |
-
`oliverj990/clock-faces-v1-times`. Most of those labels I have not checked. On
|
| 62 |
-
the 106 I did check by hand against the image, the error is 148.5 minutes, so bad
|
| 63 |
-
labels are not the excuse.
|
| 64 |
-
|
| 65 |
-
Split by source, the story is obvious:
|
| 66 |
-
|
| 67 |
-
| source | MAE | within 5 min |
|
| 68 |
-
|---|---|---|
|
| 69 |
-
| cropped faces, dial fills the frame | 87.8 min | 45.0% |
|
| 70 |
-
| photographs of rooms and streets | 163.3 min | 9.3% |
|
| 71 |
-
|
| 72 |
-
The closer a photo looks to my renders, the better it does. That is a domain gap
|
| 73 |
-
and nothing subtler.
|
| 74 |
-
|
| 75 |
-
## What it scores on synthetic data
|
| 76 |
-
|
| 77 |
-
1,200 held-out renders.
|
| 78 |
-
|
| 79 |
-
| | MAE | median | within 1 min | within 5 min |
|
| 80 |
-
|---|---|---|---|---|
|
| 81 |
-
| ground-truth crops | 4.82 min | 0.42 | 79.0% | 92.5% |
|
| 82 |
-
| end to end | 5.07 min | 0.43 | 79.0% | 92.2% |
|
| 83 |
-
|
| 84 |
-
Do not quote these. The renderer that made the labels also made the pixels, so
|
| 85 |
-
the model is being marked by its own teacher. They are here for one reason: 4.8
|
| 86 |
-
minutes on renders against 156 on photographs is the whole story.
|
| 87 |
-
|
| 88 |
-
## Does training on real photographs help? No.
|
| 89 |
-
|
| 90 |
-
`v3mix_stage2.pt` was trained on a rougher synthetic corpus (sensor noise up to
|
| 91 |
-
sigma 42, JPEG down to q25, occlusion, far wider hand geometry) plus 1,528 real
|
| 92 |
-
photographs, holding out 200 for evaluation. 106 of those 200 have labels I
|
| 93 |
-
checked by hand against the image; the rest are third-party and unverified.
|
| 94 |
-
|
| 95 |
-
On the 106 labels I actually verified:
|
| 96 |
-
|
| 97 |
-
| weights | trained on | MAE | within 5 min |
|
| 98 |
|---|---|---|---|
|
| 99 |
-
|
|
| 100 |
-
|
|
| 101 |
-
|
|
| 102 |
-
|
| 103 |
-
Adding real photographs moved it from 148.5 to 151.9, which is to say nowhere.
|
| 104 |
|
| 105 |
-
|
| 106 |
-
|
| 107 |
-
|
| 108 |
-
|
| 109 |
|
| 110 |
-
|
| 111 |
-
|
| 112 |
-
|
| 113 |
|
| 114 |
-
##
|
| 115 |
|
| 116 |
-
|
| 117 |
-
|
| 118 |
-
|
| 119 |
-
|
|
|
|
|
|
|
| 120 |
|
| 121 |
-
|
| 122 |
-
|
| 123 |
-
|
| 124 |
-
|
| 125 |
-
|
| 126 |
-
|
| 127 |
-
|
| 128 |
-
|
| 129 |
-
|
| 130 |
-
|
| 131 |
-
|
| 132 |
-
|
| 133 |
-
|
| 134 |
-
|
| 135 |
-
|
| 136 |
-
|
| 137 |
-
|
| 138 |
-
|
| 139 |
-
|
| 140 |
-
|
| 141 |
-
|
| 142 |
-
|
| 143 |
-
|
| 144 |
-
|
| 145 |
-
|
| 146 |
-
|
| 147 |
-
|
| 148 |
-
|
| 149 |
-
|
| 150 |
-
|
| 151 |
-
|
| 152 |
-
|
| 153 |
-
|
| 154 |
-
|
| 155 |
-
|
| 156 |
-
|
| 157 |
-
|
| 158 |
-
|
| 159 |
-
|
| 160 |
-
|
| 161 |
-
|
| 162 |
-
|
| 163 |
-
|
| 164 |
-
|
| 165 |
-
|
| 166 |
-
|
| 167 |
-
|
| 168 |
-
|
| 169 |
-
|
| 170 |
-
|
| 171 |
-
|
| 172 |
-
|
| 173 |
-
|
| 174 |
-
|
| 175 |
-
|
| 176 |
-
|
| 177 |
-
|
| 178 |
-
|
| 179 |
-
|
| 180 |
-
|
| 181 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 182 |
|
| 183 |
## Licence
|
| 184 |
|
|
|
|
| 11 |
|
| 12 |
# clockface
|
| 13 |
|
| 14 |
+
Reads an analog clock and returns the time. Finds the dial with a
|
| 15 |
+
COCO-pretrained detector, then reads the crop.
|
| 16 |
|
| 17 |
+
> Best measured result on 200 held-out real photographs: **34.0 minutes**
|
| 18 |
+
> mean error, **1.2** median, **49.0%** within one minute.
|
| 19 |
+
> Guessing at random is off by 180 minutes. It is better than it was and it is
|
| 20 |
+
> not a working clock reader.
|
| 21 |
|
| 22 |
+
Version 263702, commit `efe5c96`. Reader backbone: `resnet50`.
|
|
|
|
|
|
|
|
|
|
| 23 |
|
| 24 |
+
## Results on real photographs
|
|
|
|
|
|
|
|
|
|
| 25 |
|
| 26 |
+
200 held-out photos, third-party labels, all read through the same pipeline:
|
| 27 |
+
COCO detector finds the dial, the reader reads the crop.
|
|
|
|
|
|
|
|
|
|
| 28 |
|
| 29 |
+
| reader | MAE | median | within 1 min | within 5 min | worse than 30 min |
|
| 30 |
+
|---|---|---|---|---|---|
|
| 31 |
+
| `resnet50` | 34.0 min | 1.2 | 49.0% | 73.0% | 22.0% |
|
| 32 |
+
| `resnet18` | 48.5 min | 1.5 | 44.5% | 60.5% | 34.5% |
|
| 33 |
+
| `mobilenet_v3_small` | 55.6 min | 2.5 | 34.5% | 59.0% | 36.5% |
|
| 34 |
+
| random guessing | 180 min | | | | |
|
| 35 |
|
| 36 |
+
## What stage 1 is, and what it is not
|
|
|
|
|
|
|
| 37 |
|
| 38 |
+
Stage 1 is `fasterrcnn_mobilenet_v3_large_fpn` with COCO weights, run on CPU. It
|
| 39 |
+
is not trained here and not fine-tuned. Three ways of producing a crop, same
|
| 40 |
+
reader, same 200 photographs:
|
| 41 |
|
| 42 |
+
| crop | MAE | median | within 5 min |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 43 |
|---|---|---|---|
|
| 44 |
+
| whole image, no locator | 120.6 min | 88.3 | 22.5% |
|
| 45 |
+
| a dial locator trained on renders (2.2 MB) | 120.7 min | 67.9 | 25.0% |
|
| 46 |
+
| COCO detector (78 MB) | 73.6 min | 18.4 | 45.5% |
|
|
|
|
|
|
|
| 47 |
|
| 48 |
+
The locator trained on synthetic renders contributes nothing over having no
|
| 49 |
+
locator at all. It was removed. Distilling the detector into it would not have
|
| 50 |
+
helped: a student cannot beat its teacher, and the training would have run back
|
| 51 |
+
through the renderer whose domain gap is the thing being escaped.
|
| 52 |
|
| 53 |
+
Detection runs on CPU because sustained torchvision detection on an M2 trips the
|
| 54 |
+
GPU watchdog. On CPU it is 0.135s per image and finds a clock in 93-96% of real
|
| 55 |
+
photographs, having never seen a render.
|
| 56 |
|
| 57 |
+
## How the two hands become one time
|
| 58 |
|
| 59 |
+
A clock's hands are geared: at time t the hour hand sits at t/720 of a turn and
|
| 60 |
+
the minute hand at (t mod 60)/60 of a turn, so one number fixes both. The reader
|
| 61 |
+
predicts a distribution for each hand, and the decoder searches the 720 minutes
|
| 62 |
+
the clock could be showing for the one that best explains both. It cannot return
|
| 63 |
+
a pair of angles that describes no time, and a confident minute hand can pull a
|
| 64 |
+
vague hour hand onto the right hour.
|
| 65 |
|
| 66 |
+
| decoder | MAE | median | within 5 min | worse than 30 min |
|
| 67 |
+
|---|---|---|---|---|
|
| 68 |
+
| each hand read on its own | 35.3 min | 1.4 | 69.0% | 25.5% |
|
| 69 |
+
| search over times the clock could show | 34.0 min | 1.2 | 73.0% | 22.0% |
|
| 70 |
+
|
| 71 |
+
Same weights, same crops, same photographs; only the decoding differs. The gain
|
| 72 |
+
is in readings that were nearly right already, which is why MAE barely moves --
|
| 73 |
+
MAE here is set by the readings that are wrong by hours, and those are wrong in
|
| 74 |
+
the image, not in the decoding.
|
| 75 |
+
|
| 76 |
+
Scoring the hands in each other's roles, to catch the ones read the wrong way
|
| 77 |
+
round, does not work and is off by default. On these photographs the exchanged
|
| 78 |
+
reading scored a median 0.06 better where the hands really were swapped and 0.11
|
| 79 |
+
worse everywhere else -- the two populations sit on top of each other. A swap is
|
| 80 |
+
not two good angles in the wrong slots; the model is confidently wrong about
|
| 81 |
+
both hands at once and they agree on a wrong but perfectly geared time.
|
| 82 |
+
|
| 83 |
+
## When to trust a reading
|
| 84 |
+
|
| 85 |
+
The reader ranks its own answers, and the ranking is worth more than the
|
| 86 |
+
headline. Keeping only the half it is surest of:
|
| 87 |
+
|
| 88 |
+
```
|
| 89 |
+
200 readings. Keeping only the ones each signal is most sure of:
|
| 90 |
+
|
| 91 |
+
hand agreement
|
| 92 |
+
keep n MAE median <=5min
|
| 93 |
+
100% 200 34.02 1.25 73.0%
|
| 94 |
+
75% 150 26.80 1.00 79.3%
|
| 95 |
+
50% 100 22.64 0.75 86.0%
|
| 96 |
+
25% 50 30.12 0.75 82.0%
|
| 97 |
+
10% 20 20.74 0.50 90.0%
|
| 98 |
+
rank correlation with error +0.352 (useful)
|
| 99 |
+
|
| 100 |
+
head sharpness
|
| 101 |
+
keep n MAE median <=5min
|
| 102 |
+
100% 200 34.02 1.25 73.0%
|
| 103 |
+
75% 150 16.58 0.75 84.0%
|
| 104 |
+
50% 100 8.18 0.50 94.0%
|
| 105 |
+
25% 50 1.91 0.50 98.0%
|
| 106 |
+
10% 20 4.03 0.50 95.0%
|
| 107 |
+
rank correlation with error -0.506 (useful)
|
| 108 |
+
|
| 109 |
+
search margin
|
| 110 |
+
keep n MAE median <=5min
|
| 111 |
+
100% 200 34.02 1.25 73.0%
|
| 112 |
+
75% 150 18.39 0.75 84.0%
|
| 113 |
+
50% 100 10.73 0.50 92.0%
|
| 114 |
+
25% 50 6.75 0.50 96.0%
|
| 115 |
+
10% 20 3.75 0.50 95.0%
|
| 116 |
+
rank correlation with error -0.496 (useful)
|
| 117 |
+
|
| 118 |
+
A caution on reading the low-coverage rows: at 25% of 200 photographs a row
|
| 119 |
+
rests on 50 readings, so small differences between signals there are noise.
|
| 120 |
+
```
|
| 121 |
+
|
| 122 |
+
The signal this project set out to use for that was the disagreement between the
|
| 123 |
+
hands, which is free to compute. It works on renders and barely ranks anything
|
| 124 |
+
on photographs. The two that do work came out of the decoder instead.
|
| 125 |
+
|
| 126 |
+
## Honest limits
|
| 127 |
+
|
| 128 |
+
The labels on these 200 photographs are third party. 106 of them were checked by
|
| 129 |
+
hand against the image; the rest were not. The photographs were collected for
|
| 130 |
+
other purposes and skew toward clocks that are already the subject of the frame.
|
| 131 |
+
|
| 132 |
+
There is no first-party test set. A number measured on somebody else's labels is
|
| 133 |
+
worth less than one measured on your own, and that gap is not closed here.
|
| 134 |
|
| 135 |
## Licence
|
| 136 |
|