lyte-codes commited on
Commit
f76be3b
·
verified ·
1 Parent(s): c522b5c

Add v3mix weights and the negative result that came with them

Browse files
Files changed (1) hide show
  1. README.md +35 -4
README.md CHANGED
@@ -14,12 +14,17 @@ tags:
14
  Reads an analog clock and returns the time. Finds the dial first, then reads it.
15
  11.5 MB of weights.
16
 
17
- > **It does not work on real photographs.** On 1,728 real photos it is off by
18
- > 155.9 minutes on average. Guessing at random is off by 180. On the synthetic
19
- > renders it trained on, it is off by 4.8 minutes. I am publishing it because
20
  > that gap is the only interesting thing here.
21
 
22
- Version 263701, trained on synth v2.
 
 
 
 
 
23
 
24
  ## The metric
25
 
@@ -71,6 +76,32 @@ Do not quote these. The renderer that made the labels also made the pixels, so
71
  the model is being marked by its own teacher. They are here for one reason: 4.8
72
  minutes on renders against 156 on photographs is the whole story.
73
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
74
  ## Where it goes wrong
75
 
76
  Of the 1,421 predictions off by more than half an hour, 720 would improve if you
 
14
  Reads an analog clock and returns the time. Finds the dial first, then reads it.
15
  11.5 MB of weights.
16
 
17
+ > **It does not work on real photographs.** Off by 155.9 minutes on average
18
+ > across 1,728 of them. Guessing at random is off by 180. On the synthetic
19
+ > renders it trained on it is off by 4.8 minutes. I am publishing it because
20
  > that gap is the only interesting thing here.
21
 
22
+ Two weight sets are here. `stage2_best.pt` was trained on renders alone.
23
+ `v3mix_stage2.pt` was trained on rougher renders plus 1,528 real photographs,
24
+ to see whether real data closed the gap. It did not, and both are published
25
+ rather than only the flattering one.
26
+
27
+ Version 263701.
28
 
29
  ## The metric
30
 
 
76
  the model is being marked by its own teacher. They are here for one reason: 4.8
77
  minutes on renders against 156 on photographs is the whole story.
78
 
79
+ ## Does training on real photographs help? No.
80
+
81
+ `v3mix_stage2.pt` was trained on a rougher synthetic corpus (sensor noise up to
82
+ sigma 42, JPEG down to q25, occlusion, far wider hand geometry) plus 1,528 real
83
+ photographs, holding out 200 for evaluation. 106 of those 200 have labels I
84
+ checked by hand against the image; the rest are third-party and unverified.
85
+
86
+ On the 106 labels I actually verified:
87
+
88
+ | weights | trained on | MAE | within 5 min |
89
+ |---|---|---|---|
90
+ | stage2_best.pt | renders only | 148.5 min | 10.4% |
91
+ | v3mix_stage2.pt | rougher renders + 1,528 real photos | 151.9 min | 12.3% |
92
+ | random guessing | | 180 min | |
93
+
94
+ Adding real photographs moved it from 148.5 to 151.9, which is to say nowhere.
95
+
96
+ On the full 200-photo held-out set v3mix scores MAE 120.6 with 24.5% within
97
+ five minutes, which looks better and mostly is not: that set contains the
98
+ pre-cropped clock faces, which are the easy ones. Its own verified subset scores
99
+ 151.9. Composition, not skill.
100
+
101
+ Its synthetic score is 35.0 minutes against the older model's 4.8, because the
102
+ renders it trained on are much harder. That is the intended direction and it
103
+ still did not transfer.
104
+
105
  ## Where it goes wrong
106
 
107
  Of the 1,421 predictions off by more than half an hour, 720 would improve if you