Rewrite the model card in plainer language
Browse files
README.md
CHANGED
|
@@ -11,122 +11,135 @@ tags:
|
|
| 11 |
|
| 12 |
# clockface
|
| 13 |
|
| 14 |
-
Reads an analog clock
|
| 15 |
-
|
| 16 |
|
| 17 |
-
> **
|
| 18 |
-
>
|
| 19 |
-
>
|
| 20 |
-
>
|
| 21 |
|
| 22 |
-
Version 263701
|
| 23 |
|
| 24 |
## The metric
|
| 25 |
|
| 26 |
-
A clock face is a ring of 720 minutes
|
| 27 |
-
|
| 28 |
-
|
| 29 |
|
| 30 |
-
|
| 31 |
-
it is the number this model is measured against.
|
| 32 |
|
| 33 |
-
##
|
| 34 |
|
| 35 |
-
|
| 36 |
|
| 37 |
| | MAE | median | within 5 min | worse than 30 min |
|
| 38 |
|---|---|---|---|---|
|
| 39 |
-
| two
|
| 40 |
| whole image, no locator | 151.8 min | 133.7 | 9.2% | 82.5% |
|
| 41 |
-
|
|
| 42 |
|
| 43 |
-
|
| 44 |
-
|
| 45 |
-
`oliverj990/clock-faces-v1-times`. Their labels are third party and mostly
|
| 46 |
-
unverified; restricted to the 106 labels checked by hand against the image, the
|
| 47 |
-
result is MAE 148.5, so label noise does not explain it.
|
| 48 |
|
| 49 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 50 |
|
| 51 |
| source | MAE | within 5 min |
|
| 52 |
|---|---|---|
|
| 53 |
-
| cropped
|
| 54 |
-
| photographs of
|
|
|
|
|
|
|
|
|
|
| 55 |
|
| 56 |
-
|
| 57 |
|
| 58 |
-
|
| 59 |
|
| 60 |
| | MAE | median | within 1 min | within 5 min |
|
| 61 |
|---|---|---|---|---|
|
| 62 |
| ground-truth crops | 4.82 min | 0.42 | 79.0% | 92.5% |
|
| 63 |
| end to end | 5.07 min | 0.43 | 79.0% | 92.2% |
|
| 64 |
|
| 65 |
-
|
| 66 |
-
|
| 67 |
-
|
|
|
|
|
|
|
| 68 |
|
| 69 |
-
|
|
|
|
|
|
|
|
|
|
| 70 |
|
| 71 |
-
|
| 72 |
-
|
| 73 |
-
|
| 74 |
-
|
| 75 |
-
- **The confidence signal does not survive.** The two hands' disagreement is a
|
| 76 |
-
good confidence signal on synthetic data (MAE 4.82 falls to 1.64 when keeping
|
| 77 |
-
the most confident 75%) and is inert on real photographs (155.9 to 153.5). The
|
| 78 |
-
model is confidently wrong, which is worse than being uncertain.
|
| 79 |
-
- **Small clocks.** The locator finds a dial to 1.5% of image width on renders
|
| 80 |
-
and fails on wide scenes where the clock is a small part of the frame.
|
| 81 |
|
| 82 |
-
|
|
|
|
| 83 |
|
| 84 |
-
|
| 85 |
-
stage 2 PretrainedReader 2.3M params, 9.2 MB crop -> two angle distributions
|
| 86 |
|
| 87 |
-
|
| 88 |
-
|
| 89 |
-
Mises bumps rather than one-hot, and decoding takes a circular soft-argmax.
|
| 90 |
|
| 91 |
-
|
| 92 |
-
|
| 93 |
-
|
|
|
|
| 94 |
|
| 95 |
-
|
| 96 |
-
|
| 97 |
-
|
| 98 |
-
|
| 99 |
|
| 100 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 101 |
|
| 102 |
import torch
|
| 103 |
-
from twostage import
|
| 104 |
from model import decode_cls
|
| 105 |
|
| 106 |
reader = PretrainedReader(bins=180, arch="mobilenet_v3_small")
|
| 107 |
reader.load_state_dict(torch.load("stage2_best.pt")["model"])
|
| 108 |
-
hour_logits, minute_logits = reader(crop)
|
| 109 |
minutes, hour_only, disagreement, sharpness = decode_cls(hour_logits, minute_logits)
|
| 110 |
|
| 111 |
-
`minutes` is the position on the 720-minute ring
|
| 112 |
the two hands' readings are, in minutes.
|
| 113 |
|
| 114 |
-
##
|
| 115 |
|
| 116 |
-
|
| 117 |
-
|
| 118 |
|
| 119 |
-
|
| 120 |
-
|
| 121 |
|
| 122 |
## Reproducing
|
| 123 |
|
| 124 |
git clone https://github.com/lyte-codes/clockface
|
| 125 |
-
./eval.py --self-test
|
| 126 |
python3 predict.py --labels <labels> --root <dir> \
|
| 127 |
--stage2 stage2_best.pt --stage1 stage1_best.pt --out preds.jsonl
|
| 128 |
./eval.py --labels <labels> --pred preds.jsonl
|
| 129 |
|
|
|
|
|
|
|
|
|
|
| 130 |
## Licence
|
| 131 |
|
| 132 |
Apache 2.0. The synthetic dataset is CC BY 4.0.
|
|
|
|
| 11 |
|
| 12 |
# clockface
|
| 13 |
|
| 14 |
+
Reads an analog clock and returns the time. Finds the dial first, then reads it.
|
| 15 |
+
11.5 MB of weights.
|
| 16 |
|
| 17 |
+
> **It does not work on real photographs.** On 1,728 real photos it is off by
|
| 18 |
+
> 155.9 minutes on average. Guessing at random is off by 180. On the synthetic
|
| 19 |
+
> renders it trained on, it is off by 4.8 minutes. I am publishing it because
|
| 20 |
+
> that gap is the only interesting thing here.
|
| 21 |
|
| 22 |
+
Version 263701, trained on synth v2.
|
| 23 |
|
| 24 |
## The metric
|
| 25 |
|
| 26 |
+
A clock face is a ring of 720 minutes, and error is the shorter way round it. So
|
| 27 |
+
11:58 against 12:02 is 4 minutes, not 716. AM and PM do not exist, because a face
|
| 28 |
+
cannot show them: 3:47 and 15:47 are the same reading.
|
| 29 |
|
| 30 |
+
Random guessing scores 180 minutes. That is the bar.
|
|
|
|
| 31 |
|
| 32 |
+
## What it scores on real photographs
|
| 33 |
|
| 34 |
+
1,728 images with third-party labels.
|
| 35 |
|
| 36 |
| | MAE | median | within 5 min | worse than 30 min |
|
| 37 |
|---|---|---|---|---|
|
| 38 |
+
| two stages | 155.9 min | 139.9 | 11.4% | 82.2% |
|
| 39 |
| whole image, no locator | 151.8 min | 133.7 | 9.2% | 82.5% |
|
| 40 |
+
| random guessing | 180 min | | | |
|
| 41 |
|
| 42 |
+
Dropping the locator makes it slightly better, which tells you how much the
|
| 43 |
+
locator is contributing.
|
|
|
|
|
|
|
|
|
|
| 44 |
|
| 45 |
+
The images are 1,092 film stills from Marclay's *The Clock*, 506 photos from
|
| 46 |
+
`moondream/analog-clock-benchmark`, and 131 cropped faces from
|
| 47 |
+
`oliverj990/clock-faces-v1-times`. Most of those labels I have not checked. On
|
| 48 |
+
the 106 I did check by hand against the image, the error is 148.5 minutes, so bad
|
| 49 |
+
labels are not the excuse.
|
| 50 |
+
|
| 51 |
+
Split by source, the story is obvious:
|
| 52 |
|
| 53 |
| source | MAE | within 5 min |
|
| 54 |
|---|---|---|
|
| 55 |
+
| cropped faces, dial fills the frame | 87.8 min | 45.0% |
|
| 56 |
+
| photographs of rooms and streets | 163.3 min | 9.3% |
|
| 57 |
+
|
| 58 |
+
The closer a photo looks to my renders, the better it does. That is a domain gap
|
| 59 |
+
and nothing subtler.
|
| 60 |
|
| 61 |
+
## What it scores on synthetic data
|
| 62 |
|
| 63 |
+
1,200 held-out renders.
|
| 64 |
|
| 65 |
| | MAE | median | within 1 min | within 5 min |
|
| 66 |
|---|---|---|---|---|
|
| 67 |
| ground-truth crops | 4.82 min | 0.42 | 79.0% | 92.5% |
|
| 68 |
| end to end | 5.07 min | 0.43 | 79.0% | 92.2% |
|
| 69 |
|
| 70 |
+
Do not quote these. The renderer that made the labels also made the pixels, so
|
| 71 |
+
the model is being marked by its own teacher. They are here for one reason: 4.8
|
| 72 |
+
minutes on renders against 156 on photographs is the whole story.
|
| 73 |
+
|
| 74 |
+
## Where it goes wrong
|
| 75 |
|
| 76 |
+
Of the 1,421 predictions off by more than half an hour, 720 would improve if you
|
| 77 |
+
swapped the hour and minute hands. My generator drew hands from a narrow band of
|
| 78 |
+
lengths and widths, and the model learned that band instead of learning which
|
| 79 |
+
hand is longer. Real clocks do not agree with my band.
|
| 80 |
|
| 81 |
+
The confidence signal does not survive the trip to real data. On renders it works
|
| 82 |
+
well: keep the most confident 75% and error drops from 4.82 to 1.64 minutes. On
|
| 83 |
+
photographs it does nothing, 155.9 to 153.5. The model is confidently wrong,
|
| 84 |
+
which is worse than being uncertain.
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 85 |
|
| 86 |
+
The locator finds a dial to within 1.5% of image width on renders. On a photo of
|
| 87 |
+
a room where the clock is a smudge on a far wall, it does not.
|
| 88 |
|
| 89 |
+
## How it is built
|
|
|
|
| 90 |
|
| 91 |
+
stage 1 DialLocator 554K params, 2.2 MB image -> centre, radius
|
| 92 |
+
stage 2 PretrainedReader 2.3M params, 9.2 MB crop -> two angle distributions
|
|
|
|
| 93 |
|
| 94 |
+
Stage 2 is MobileNetV3-Small pretrained on ImageNet, with two heads over 180
|
| 95 |
+
angular bins each. Targets are von Mises bumps rather than one-hot, so being one
|
| 96 |
+
bin out costs less than being ten out, and decoding uses a circular soft-argmax
|
| 97 |
+
to get back below bin resolution.
|
| 98 |
|
| 99 |
+
The same network without ImageNet weights never learns this at all. It sits at
|
| 100 |
+
2·ln(180), a flat distribution, for as long as you care to train it, while
|
| 101 |
+
happily memorising 400 images. The pretrained initialisation is the difference
|
| 102 |
+
between a model and a plateau.
|
| 103 |
|
| 104 |
+
It predicts both hands and never the time directly. That costs nothing and buys
|
| 105 |
+
two things. The hands disagreeing is a confidence signal. And decoding can work
|
| 106 |
+
like a vernier scale, taking precision from the minute hand and only the hour
|
| 107 |
+
count from the hour hand, which under noise is about 11 times more accurate than
|
| 108 |
+
reading the hour hand alone.
|
| 109 |
+
|
| 110 |
+
## Using it
|
| 111 |
|
| 112 |
import torch
|
| 113 |
+
from twostage import PretrainedReader
|
| 114 |
from model import decode_cls
|
| 115 |
|
| 116 |
reader = PretrainedReader(bins=180, arch="mobilenet_v3_small")
|
| 117 |
reader.load_state_dict(torch.load("stage2_best.pt")["model"])
|
| 118 |
+
hour_logits, minute_logits = reader(crop) # 1x3x256x256, ImageNet normalised
|
| 119 |
minutes, hour_only, disagreement, sharpness = decode_cls(hour_logits, minute_logits)
|
| 120 |
|
| 121 |
+
`minutes` is the position on the 720-minute ring. `disagreement` is how far apart
|
| 122 |
the two hands' readings are, in minutes.
|
| 123 |
|
| 124 |
+
## What it is for
|
| 125 |
|
| 126 |
+
Round analog clock faces. Not digital-analog hybrids, subdials, chronographs or
|
| 127 |
+
24-hour faces.
|
| 128 |
|
| 129 |
+
Do not use it for anything where a wrong time costs something. It is off by more
|
| 130 |
+
than half an hour on 82% of real photographs.
|
| 131 |
|
| 132 |
## Reproducing
|
| 133 |
|
| 134 |
git clone https://github.com/lyte-codes/clockface
|
| 135 |
+
./eval.py --self-test
|
| 136 |
python3 predict.py --labels <labels> --root <dir> \
|
| 137 |
--stage2 stage2_best.pt --stage1 stage1_best.pt --out preds.jsonl
|
| 138 |
./eval.py --labels <labels> --pred preds.jsonl
|
| 139 |
|
| 140 |
+
`--self-test` runs 32 checks on the metric itself, including both directions
|
| 141 |
+
across the 12 o'clock seam.
|
| 142 |
+
|
| 143 |
## Licence
|
| 144 |
|
| 145 |
Apache 2.0. The synthetic dataset is CC BY 4.0.
|