File size: 29,139 Bytes
1995ff1
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
305
306
307
308
309
310
311
312
313
314
315
316
317
318
319
320
321
322
323
324
325
326
327
328
329
330
331
332
333
334
335
336
337
338
339
340
341
342
343
344
345
346
347
348
349
350
351
352
353
354
355
356
357
358
359
360
361
362
363
364
365
366
367
368
369
370
371
372
373
374
375
376
377
378
379
380
381
382
383
384
385
386
387
388
389
390
391
392
393
394
395
396
397
398
399
400
401
402
403
404
405
406
407
408
409
410
411
412
413
414
415
416
417
418
419
420
421
422
423
424
425
426
427
428
429
430
431
432
433
434
435
436
437
438
439
440
441
442
443
444
445
446
447
448
449
450
451
452
453
454
455
456
457
458
459
460
461
462
463
464
465
466
467
468
469
470
471
472
473
474
475
476
477
478
479
480
481
482
483
484
485
486
487
488
489
490
491
492
493
494
495
496
497
498
499
500
501
502
503
504
505
506
507
508
509
510
511
512
513
514
515
516
517
518
519
520
521
522
523
524
525
526
527
528
529
530
531
532
533
534
535
536
537
538
539
540
541
542
543
544
545
546
547
548
549
550
551
552
553
554
555
556
557
558
559
560
561
562
563
564
565
566
567
568
569
570
571
572
573
574
575
576
577
578
579
580
581
582
583
584
585
586
587
588
589
590
591
592
593
594
595
596
597
598
599
600
601
602
603
604
605
606
607
608
609
610
611
612
613
614
615
616
617
618
619
620
621
622
623
624
625
626
627
628
629
630
631
632
633
634
635
636
637
638
639
640
641
642
643
644
645
646
647
648
649
650
651
652
653
654
655
656
657
658
659
660
661
662
663
664
665
666
667
668
669
670
671
672
\documentclass[aspectratio=169,10pt]{beamer}
\usetheme{default}
\usecolortheme{default}
\setbeamertemplate{navigation symbols}{}
\setbeamertemplate{footline}[frame number]
\setbeamerfont{frametitle}{size=\large,series=\bfseries}
\setbeamerfont{framesubtitle}{size=\small,series=\mdseries,shape=\itshape}

\usepackage{booktabs}
\usepackage{array}
\usepackage{colortbl}
\usepackage{xcolor}
\usepackage[utf8]{inputenc}
\usepackage{textcomp}
\usepackage{amsmath}
\usepackage{amssymb}

\definecolor{hl}{RGB}{0,90,158}
\definecolor{bad}{RGB}{178,34,34}
\definecolor{good}{RGB}{0,125,75}
\definecolor{mute}{RGB}{110,110,110}

\newcommand{\hl}[1]{\textcolor{hl}{\textbf{#1}}}
\newcommand{\bad}[1]{\textcolor{bad}{\textbf{#1}}}
\newcommand{\good}[1]{\textcolor{good}{\textbf{#1}}}

\title{The In-Submap Accuracy Problem}
\subtitle{Evidence, Baselines, and the Frame-BA Pivot}
\author{Qi Zhang}
\date{2026-04-21}

\begin{document}

\begin{frame}
\titlepage
\end{frame}

% ==================================================================
%  PART 1 -- EVIDENCE (RAW TABLES)
% ==================================================================

% ------------------------------------------------------------------
\begin{frame}{Problem in one sentence}
\vspace{0.5em}
V4 SLAM-Former retrieves loops correctly but the submaps themselves
carry \bad{$\sim$15\% ATE of traveled distance}. Any pose-graph built
on top of this cannot beat the same ceiling.
\vspace{1.0em}

\textbf{What you'll see in this deck:}
\begin{enumerate}
  \item Per-submap Sim(3)-aligned ATE --- 10 real scenes (KITTI 00 + 9 TUM fr1).
  \item Retrieval overlap --- voxel-based Q$\leftrightarrow$H and H$\leftrightarrow$H on predicted pointclouds.
  \item $\Delta t$ retrieval --- KITTI 00 full distribution.
  \item M3-SLAM benchmark and per-seq oracle split.
  \item Top-K and bn\_every sweeps (full tables).
  \item Indoor and outdoor cross-dataset scoreboard.
  \item Matcher eval (TUM fr1/xyz, ScanNet++ iPhone) and Sim(3) frame-BA per scene.
  \item Methods: descriptor extraction, matching, Sim(3) optimization.
\end{enumerate}
\end{frame}

% ------------------------------------------------------------------
\begin{frame}{V4 SLAM-Former pipeline (context)}
\small
\begin{center}
\begin{tabular}{p{2.2cm}|p{4.5cm}|p{5.6cm}}
\toprule
\textbf{Stage} & \textbf{Unit} & \textbf{What it outputs} \\
\midrule
Backbone    & per-frame tokens   & depth, pose, feature maps (V4-B'' ckpt-1) \\
Frontend    & bn\_every=6 frames & streaming \emph{submap} (local 3D + poses) \\
Scorer      & importance\_topk   & top-K historical submaps via Q$\cdot$V attention \\
Backend     & retrieved submaps  & refined submap poses with cross-submap context \\
Sim(3) stitch & submap overlap   & chained world trajectory (\texttt{enable\_sim3\_stitch}) \\
\bottomrule
\end{tabular}
\end{center}
\vspace{0.6em}
\textbf{Unit of measurement:} one submap = 6 frames of local pose $+$ pointcloud.
In-submap ATE = Umeyama-Sim(3)-aligned trajectory error inside a single submap,
before any stitching. Measured by \texttt{scripts/insubmap\_accuracy.py}.
\end{frame}

% ------------------------------------------------------------------
\begin{frame}{Evidence A --- per-submap ATE (full table)}
\framesubtitle{V4-B'' ckpt-1, Sim(3) Umeyama, measured 2026-04-21}
\scriptsize
\begin{center}
\begin{tabular}{l r r r r r r}
\toprule
\textbf{scene} & \textbf{\#submaps} & \textbf{frames/sm} & \textbf{median ATE (m)} & \textbf{p95 (m)} & \textbf{traveled/sm (m)} & \textbf{ATE/dist} \\
\midrule
KITTI 00 & 389 & 8 & \bad{0.644} & 1.82  & 5.12 & \bad{14.2\%} \\
\midrule
TUM xyz   & 8  & 4 & 0.085 & 0.154 & 0.36 & \bad{31.0\%} \\
TUM desk  & 11 & 4 & 0.051 & 0.112 & 0.36 & 12.4\% \\
TUM desk2 & 26 & 4 & 0.063 & 0.114 & 0.51 & 13.0\% \\
TUM room  & 52 & 4 & 0.039 & 0.125 & 0.32 & 13.6\% \\
TUM 360   & 39 & 5 & 0.044 & 0.081 & 0.26 & 17.0\% \\
TUM rpy   & 36 & 4 & 0.017 & 0.032 & 0.08 & 20.7\% \\
TUM teddy & 23 & 4 & 0.049 & 0.130 & 0.34 & 10.9\% \\
TUM floor & 30 & 4 & 0.040 & 0.110 & 0.26 & 15.7\% \\
TUM plant & 32 & 4 & 0.055 & 0.134 & 0.42 & 15.3\% \\
\midrule
\textbf{TUM mean} & & & \textbf{0.050} & \textbf{0.110} & \textbf{0.32} & \bad{\textbf{16.6\%}} \\
\bottomrule
\end{tabular}
\end{center}
\vspace{0.3em}
{\footnotesize full\_token and top5\_dual\_queue modes produce bitwise-identical dumps under
\texttt{--submap\_fetch\_source frontend} (md5 verified). Numbers hold for both inference modes.}
\end{frame}

% ------------------------------------------------------------------
\begin{frame}{Evidence B --- retrieval overlap (probe2, full sweep)}
\framesubtitle{Voxel-based pairwise overlap on V4-B'' predicted pointclouds}
\scriptsize
\begin{center}
\begin{tabular}{l l r r r r l}
\toprule
\textbf{scene} & \textbf{voxel} & \textbf{Q$\leftrightarrow$H} & \textbf{H$\leftrightarrow$H} & \textbf{frac H$\leftrightarrow$H$>$5\%} & \textbf{\#submap pairs} & \textbf{verdict} \\
\midrule
7Sc office/seq-02 & 0.05 m & 0.339 & \good{0.370} & 84.1\% & --- & healthy \\
TUM fr1/desk      & 0.05 m & 0.458 & \good{0.393} & 76.7\% & --- & healthy \\
KITTI 00          & 0.50 m & 0.361 & \good{0.499} & 92.5\% & 389 & healthy \\
KITTI 00          & 1.00 m & 0.444 & \good{0.674} & 98.9\% & 389 & healthy \\
KITTI 00          & 2.00 m & 0.629 & \good{0.829} & 98.9\% & 389 & healthy \\
\midrule
\multicolumn{7}{l}{\textbf{For reference --- hook-translation artifact (probe1, pre-fix):}} \\
7Sc office/seq-02 & 0.05 m & 0.056 & \bad{0.008} & 4.4\%  & --- & ARTIFACT \\
7Sc office/seq-02 & 0.40 m & 0.117 & \bad{0.026} & 4.4\%  & --- & ARTIFACT \\
\bottomrule
\end{tabular}
\end{center}
\vspace{0.5em}
{\footnotesize
Retrieved sids were row-by-row \emph{identical} between probe1 and probe2; only the
pointcloud-loader got fixed (float \texttt{frame\_id} parsing + timestamp-stem filename lookup).
Early ``7Sc rejected / TTA 14$\times$ lift'' reading was driven entirely by the load miss.
}
\end{frame}

% ------------------------------------------------------------------
\begin{frame}{Evidence C --- $\Delta t$ reach of the scorer (full distribution)}
\framesubtitle{KITTI 00, V4-B'' ckpt-1, importance\_topk=20: 215 submaps, 4070 fetches}
\scriptsize
\textbf{$\Delta t$ distribution over all 4070 fetches:}
\begin{center}
\begin{tabular}{l r r r r r r r}
\toprule
\textbf{$\Delta t$ bin} & 1--3 & 4--9 & 10--19 & 20--49 & 50--99 & 100+ & \textbf{total $\ge 50$} \\
\midrule
share of fetches & 2.8\% & 7.9\% & 9.9\% & 22.9\% & \good{27.4\%} & \good{29.0\%} & \good{56.4\%} \\
\bottomrule
\end{tabular}
\end{center}
\vspace{0.6em}
\textbf{Per-query loop closure rate:}
\begin{center}
\begin{tabular}{l r r r r}
\toprule
\textbf{stat} & median $\Delta t$ & mean $\Delta t$ & max $\Delta t$ & queries with $\exists\Delta t\ge 50$ \\
\midrule
KITTI 00 (k=20) & 60 & 70.3 & 214 & \good{76.7\%} \\
KITTI 02 (k=10) & 48 & --- & --- & 72.7\% \\
TUM fr1/teddy (topall, bn=4) & 11 & --- & 34 & 0\% (seq too short) \\
\bottomrule
\end{tabular}
\end{center}
\vspace{0.6em}
\textbf{Anchor-dominance sanity check (KITTI 00):} fetched sid=0 in only 5.2\% of the 4070 fetches.
Top-fetched $\Delta t$ values: 2, 3, 5--11 ($\sim$50 fetches each) --- recent history is fetched
but the long tail is thick. Training-time finetune only supervises $\Delta t\le 3$; inference runs out to 214.
\end{frame}

% ==================================================================
%  PART 2 -- BASELINE SCOREBOARD
% ==================================================================

% ------------------------------------------------------------------
\begin{frame}{M3-SLAM benchmark --- full paper table}
\framesubtitle{ATE RMSE (m) -- lower is better. Source: m3-slam/sec/experiments.tex.}
\small
\begin{center}
\begin{tabular}{l r r r r r r}
\toprule
\textbf{Dataset} & MonoGS & DROID-SLAM & LEAP-VO & VGGT-SLAM & Pi3X-slam & \textbf{M3-SLAM} \\
\midrule
ScanNet++ (20 scenes)  & 0.623 & 0.346  & 1.081  & 0.182 & 0.137 & \good{0.065} \\
ScanNetV2 (10 scenes)  & 0.217 & 0.106  & 0.852  & 0.073 & 0.141 & \good{0.051} \\
Waymo (9 seqs)         & 7.535 & 16.61  & 13.17  & 1.295 & 2.410 & \good{0.773} \\
KITTI (8 seqs)         & 8.212 & 18.07  & 10.99  & 2.521 & 2.938 & \good{0.890} \\
\bottomrule
\end{tabular}
\end{center}
\vspace{0.8em}

\textbf{V4-B'' ckpt-1 on the same axis:}
\begin{center}
\begin{tabular}{l r r r}
\toprule
\textbf{Dataset} & V4-B'' ckpt-1 & M3-SLAM & gap \\
\midrule
TUM fr1 mean (9 scenes)   & 0.224 & \textit{n/a}  & same order as M3 indoor \\
KITTI 00 (frame-BA floor) & \bad{172.6} & 0.890 & \bad{194$\times$} \\
\bottomrule
\end{tabular}
\end{center}

\textit{Same backbone family (Pi3X-class), same monocular streaming regime. Architectural
differences only: matching head, pose-guided local search, Sim(3) global BA.}
\end{frame}

% ------------------------------------------------------------------
\begin{frame}{Oracle chain-TRS --- full per-backbone split}
\framesubtitle{V4-oracle chain-TRS(6,4): oracle substitutes GT poses for 6-frame submap overlap alignment}
\scriptsize
\textbf{KITTI (FINAL, avg over 11 sequences, lower is better):}
\begin{center}
\begin{tabular}{l r r r}
\toprule
\textbf{Backbone} & top-5 ATE (m) & \#seqs won vs others & per-seq winners \\
\midrule
V4-B'  (paper + scannetpp + mvs\_synth + vK) & 72.5 & 0/11 & --- \\
V4-B'' (V3 ckpt-2 + same data mix)           & \good{41.3} & 6/11 & 02, 03, 04, 07, 08, 10 \\
no-vK  (paper + scannetpp + mvs\_synth)      & \good{40.7} & 5/11 & 00, 01, 05, 06, 09 \\
\bottomrule
\end{tabular}
\end{center}
\vspace{0.4em}
\textbf{TUM fr1 (avg over 8 sequences):}
\begin{center}
\begin{tabular}{l r r r}
\toprule
\textbf{Backbone} & top-5 & top-all & gap to paper target (0.045) \\
\midrule
V4-B'  & 0.183 & 0.165 & 3.7$\times$ \\
V4-B'' & 0.160 & \good{0.102} & 2.3$\times$ \\
no-vK  & 0.145 & 0.136 & 3.0$\times$ \\
\bottomrule
\end{tabular}
\end{center}
\vspace{0.4em}
With \emph{perfect} overlap alignment we still get 40.7 m KITTI / 0.102 m TUM.
Stitch is not the bottleneck; the 15\% in-submap ATE is.
\end{frame}

% ------------------------------------------------------------------
\begin{frame}{Ablation --- top-K retrieval, full sweep}
\framesubtitle{\S15.13, V4-oracle chain-TRS, bn\_every=6 fixed, same ckpts as scoreboard}
\small
\textbf{TUM fr1 (V4-B'' ckpt-1, all 8/8 seqs complete):}
\begin{center}
\begin{tabular}{l r r r r r}
\toprule
top-K & 5 & 10 & \textbf{20} & 50 & all \\
\midrule
mean ATE (m) & 0.160 & 0.106 & \good{0.100} & 0.102 & 0.102 \\
\bottomrule
\end{tabular}
\end{center}
\textit{Monotone down k=5$\to$10$\to$20, then plateau. TUM sweet spot = k=20.}

\vspace{0.8em}
\textbf{KITTI (no-vK ckpt-1, 11/11 at k=5/10/20, 6/11 at k=50):}
\begin{center}
\begin{tabular}{l r r r r}
\toprule
top-K & \textbf{5} & 10 & 20 & 50 \\
\midrule
mean ATE (m) & \good{40.7} & 48.5 & 56.2 & 69.4 \\
\bottomrule
\end{tabular}
\end{center}
\textit{Monotone \emph{up} --- adding history \emph{hurts}.  KITTI sweet spot = k=5.}

\vspace{0.6em}
\textbf{Domain-specific top-K is required.} KITTI forward motion has no persistent
loop back until kilometre scale, so large top-K just injects irrelevant tokens.
TUM sees the same scene from many angles; more retrieval = more coverage.
\end{frame}

% ------------------------------------------------------------------
\begin{frame}{Ablation --- frames per submap (\texttt{bn\_every})}
\framesubtitle{\S15.14, V4-oracle chain-TRS, top-K=5 fixed}
\small
\begin{center}
\begin{tabular}{l r r r l}
\toprule
\textbf{dataset} & bn=4 & \textbf{bn=6} & bn=10 & coverage \\
\midrule
TUM fr1 (V4-B'')       & 0.174 & \good{0.160} & 0.170 & 8/8 at every bn \\
KITTI (no-vK)          & 75.7 & \good{40.7} & 77.5 & bn4=8/11, bn6=11/11, bn10=9/11 \\
\bottomrule
\end{tabular}
\end{center}
\vspace{0.8em}

\textbf{Both domains prefer the default bn=6.}
bn=4 fragments context before the scorer can use it; bn=10 accumulates more depth bias
inside the window. Same pattern as top-K: neither knob closes the gap to M3.

\vspace{0.6em}
\textbf{Canonical config} (carried through all subsequent experiments): V4-B'' backbone
+ bn\_every=6 + top-K=20 on TUM / top-K=5 on KITTI.
\end{frame}

% ------------------------------------------------------------------
\begin{frame}{Cross-dataset scoreboard --- indoor (V4-B'' ckpt-1)}
\framesubtitle{Evaluated 2026-04-20}
\small
\begin{center}
\begin{tabular}{l r r r r r}
\toprule
\textbf{benchmark} & \textbf{\#seqs} & \textbf{mean ATE (m)} & \textbf{AUC@3$^\circ$} & \textbf{AUC@30$^\circ$} & \textbf{status} \\
\midrule
7-Scenes            & 17 & 0.193 & $\sim$0.05 & 0.642 & ATE at paper parity, rotation weak \\
NRGBD               & 8  & 0.252 & $\sim$0.01 & 0.780 & ATE at paper parity (OpenGL fix required) \\
TUM fr1             & 9  & 0.224 & \bad{0.000} & 0.368 & AUC@3$^\circ$=0 systemic \\
\bottomrule
\end{tabular}
\end{center}
\vspace{0.8em}

\textbf{TUM fr1 per-scene ATE (V4-B'' ckpt-1):}
\begin{center}
\begin{tabular}{l r r r r r r r r r r}
\toprule
scene & 360 & desk & desk2 & floor & plant & room & rpy & teddy & xyz & \textbf{mean} \\
\midrule
ATE (m) & 0.17 & 0.11 & \good{0.017} & 0.13 & --- & \bad{0.73} & 0.24 & 0.23 & 0.10 & 0.224 \\
\bottomrule
\end{tabular}
\end{center}
\vspace{0.3em}
{\footnotesize Rotation AUC@3$^\circ$=0 even on scenes where translation is cm-scale $\Rightarrow$
indoor rotation is under-fit, not a geometric failure.}
\end{frame}

% ------------------------------------------------------------------
\begin{frame}{Cross-dataset scoreboard --- outdoor (V4-B'' ckpt-1)}
\framesubtitle{Evaluated 2026-04-20}
\small
\begin{center}
\begin{tabular}{l r r r r}
\toprule
\textbf{benchmark} & \textbf{mean ATE (m)} & \textbf{paper target (m)} & \textbf{M3 (m)} & \textbf{gap vs paper / M3} \\
\midrule
ETH3D (median over 13 scenes) & 0.63  & --- & --- & AUC@3$^\circ$=0.01 vs paper 0.67 ($30\times$) \\
Oxford Spires / keble  & 25.25 & 6--12 & --- & 2--4$\times$ \\
Oxford Spires / christ & 36.15 & 6--12 & --- & 3--6$\times$ \\
Oxford Spires mean     & 30.7  & 6--12 & --- & 3--5$\times$ \\
KITTI (mean)           & \bad{148}   & 14.9  & 0.89 & \bad{10$\times$ paper, 166$\times$ M3} \\
\bottomrule
\end{tabular}
\end{center}
\vspace{0.8em}

\textbf{Oxford Spires alignment coverage:} n\_matched $\approx$ 100 / 3840 patches
$\Rightarrow$ \bad{2.6\% alignment coverage}. Most frames never align under the current
matcher. The 25--36 m errors are driven by unaligned frames, not by local geometry
precision.
\vspace{0.3em}

{\footnotesize ETH3D F1$>$0 on 4/13 scenes only. Bad F1 is a coverage issue, not a depth one.}
\end{frame}

% ==================================================================
%  PART 3 -- METHODS (MATCHING + DESCRIPTORS + SIM(3) OPTIMIZATION)
% ==================================================================

% ------------------------------------------------------------------
\begin{frame}{Matching system pipeline}
\vspace{-0.3em}
\begin{center}
\resizebox{\linewidth}{!}{\includegraphics{matching_pipeline_figure.pdf}}
\end{center}
\vspace{-0.7em}
{\footnotesize
\textbf{Four stages, all inference-time, no training:}
(1) SLAM-Former backbone returns per-patch tokens $F\in\mathbb{R}^{P\times d}$ and local 3D $X\in\mathbb{R}^{P\times 3}$;
(2) rank patches by confidence $c_p$, keep top-$K$, L2-normalize the token as the descriptor;
(3) mutual-NN cosine match + Lowe ratio + $3\!\times\!3$ quadratic sub-pixel refinement;
(4) Sim(3) frame BA (scipy LSMR, 7 DoF per frame, velocity-smoothness scale prior $\lambda_{\rm scale}=0.1$).
No learned matching head; reuses backbone tokens as features.}
\end{frame}

% ------------------------------------------------------------------
\begin{frame}{Matching --- how descriptors are extracted}
\small
\textbf{Ours (\texttt{frame\_matcher.py:select\_frame\_keypoints}):}
\begin{itemize}\small
  \item Backbone returns per-patch tokens $\mathbf{t}_{\rm patch}\in\mathbb{R}^{H_p\times W_p\times D}$
        ($H_p=W_p=37$ at target\_size=518, patch=14$\times$14 px, $D$=token dim).
  \item Per-patch confidence: mean of pixel-wise conf over the 14$\times$14 cell.
  \item Patch-center 3D point $\mathbf{X}_i^i\in\mathbb{R}^3$ from the predicted local pointcloud;
        patch-center normalized camera UV $(u,v) = (x/z, y/z)$ of that 3D point.
  \item Top-$K$ patches by confidence, with $z\in(z_{\min},z_{\max})$ finite.
  \item Descriptor = patch token, L2-normalized, stored as fp16.
\end{itemize}
\vspace{0.6em}

\textbf{M3-SLAM (from \texttt{refer/m3-slam/sec/method.tex}):}
\begin{itemize}\small
  \item Separate \textbf{matching head} on top of Pi3X: $\mathrm{DPT}_{\rm desc}+$2-layer MLP with GELU.
  \item Outputs dense descriptors $\mathbf{D}\in\mathbb{R}^{H\times W\times d}$
        ($d=24$ by default, \emph{per pixel}, not per patch) and matching confidence
        $\mathbf{Q}\in\mathbb{R}^{H\times W\times 1}$.
  \item Trained with InfoNCE symmetric loss against GT pixel correspondences;
        encoder / decoder / point head frozen.
\end{itemize}
\vspace{0.4em}
{\footnotesize \textbf{The structural gap:} our ``descriptor'' is a token over a
14$\times$14 patch, giving $37^2 \approx 1369$ candidates per frame; M3's is
\emph{per pixel}, giving $518^2 \approx 268{,}000$ candidates per frame. A single
token integrates over $\sim$200 pixels of depth variation.}
\end{frame}

% ------------------------------------------------------------------
\begin{frame}{Matching --- how pairs are formed}
\small
\textbf{Ours (\texttt{frame\_matcher.py:match\_frames}):}
\begin{enumerate}\small
  \item Mutual-nearest-neighbour on cosine similarity between L2-normalized token descriptors.
  \item Lowe's ratio test on cosine: keep pair if $\cos_{\rm top2}\le r\cdot\cos_{\rm top1}$,
        typical $r\in[0.8,0.9]$.
  \item Sub-pixel refinement via quadratic fit on the 3$\times$3 cos-sim neighborhood of
        the peak, offset clipped to $[-0.5,+0.5]$ patch units.
  \item Output: \texttt{MatchEdge} with indices $i\leftrightarrow j$, cos-sim, and refined $(u_j,v_j)$.
\end{enumerate}
\vspace{0.6em}

\textbf{M3-SLAM (Sec. 3.1.3 --- ``Dense matching for SLAM''):}
\begin{enumerate}\small
  \item Pose-guided initialization: project $\mathbf{X}_i$ into frame $j$ using current $(T_i, T_j)$
        $$\mathbf{X}_i^j = T_j^{-1}\,T_i\,\mathbf{X}_i.$$
  \item For each pixel $\mathbf{p}_i$, search a local window of radius $r$ around the projected location
        in frame $j$, selecting $\mathbf{p}^*_j = \arg\max_{\mathbf{p}\in\mathcal{N}_r} \big\langle \mathbf{D}_i(\mathbf{p}_i), \mathbf{D}_j(\mathbf{p})\big\rangle$
        (cosine similarity).
  \item Dynamic masking: a motion map $\mathbf{M}_i\in[0,1]^{H\times W}$ down-weights pixels with
        low warped-vs-observed descriptor consistency.
\end{enumerate}
\vspace{0.4em}
{\footnotesize \textbf{Key difference:} M3's matching is $O(N\cdot r^2)$ local geometric search;
ours is $O(K^2)$ global descriptor matching with no pose prior. M3's search radius $r$ can be
as small as a few pixels \emph{because the pose initialization is good}.}
\end{frame}

% ------------------------------------------------------------------
\begin{frame}{Matcher precision --- how it actually performs}
\framesubtitle{GT-supervised M3-style eval (\texttt{eval\_matcher\_gt.py}), $\tau=0.6$, err threshold 0.10 m}
\small
\textbf{TUM fr1/xyz (48 frames, patch-level):}
\begin{center}
\begin{tabular}{l r r r r}
\toprule
err threshold & 5 cm & 10 cm & 20 cm & 50 cm \\
\midrule
precision & \bad{$\sim$8\%} & --- & 55\% & 88\% \\
\bottomrule
\end{tabular}
\end{center}
median patch-pair 3D error $\approx$ \textbf{19 cm} (patches = 14$\times$14 px $\approx$ 17 original pixels).

\vspace{0.8em}
\textbf{ScanNet++ matcher eval (\#22058916, STRIDE=5 iPhone frames, 3 scenes):}
\begin{center}
\begin{tabular}{l r r r r r}
\toprule
\textbf{scene} & \textbf{\#pairs} & \textbf{\#matches} & \textbf{\#correct} & \textbf{precision} & \textbf{median err (m)} \\
\midrule
036bce3393  & 51 & 1292 &  75 & \bad{5.8\%}  & 1.99 \\
079a326597  & 41 & 1319 & 172 & \bad{13.0\%} & 0.68 \\
07ff1c45bb  & 53 & 1286 &  59 & \bad{4.6\%}  & 1.31 \\
\midrule
mean        & 48 & 1299 & 102 & \bad{7.8\%}  & 1.33 \\
\bottomrule
\end{tabular}
\end{center}

\vspace{0.3em}
{\footnotesize iPhone STRIDE=5 keyframe gaps of 0.3--0.8 s produce wide view changes.
Precision 5--13\% vs TUM 30\% baseline is the natural consequence of patch-level
descriptors under large disparity.}
\end{frame}

% ------------------------------------------------------------------
\begin{frame}{Sim(3) optimization --- parameterization}
\small
\textbf{Pose as a similarity transform} $T\in \mathbf{Sim}(3)$:
$$
T = \begin{bmatrix} s\mathbf{R} & \mathbf{t} \\ \mathbf{0}^\top & 1 \end{bmatrix},
\qquad \mathbf{R}\in\mathbf{SO}(3),\ \mathbf{t}\in\mathbb{R}^3,\ s\in\mathbb{R}_{>0}
$$
\vspace{0.4em}

\textbf{Lie algebra representation} $\boldsymbol{\tau}\in\mathfrak{sim}(3)\cong\mathbb{R}^7$:
\begin{itemize}
  \item 3 for axis-angle rotation $\boldsymbol{\omega}\in\mathbb{R}^3$, 3 for translation
        $\boldsymbol{\rho}\in\mathbb{R}^3$, 1 for log-scale $\lambda\in\mathbb{R}$.
  \item Update via left-multiplication on the manifold:
        $T \leftarrow \exp(\boldsymbol{\tau})\circ T$.
\end{itemize}
\vspace{0.8em}

\textbf{Ours (\texttt{frame\_ba.py:FrameBA}, \texttt{use\_sim3=True}):}
\begin{itemize}
  \item State vector $\boldsymbol{\theta}\in\mathbb{R}^{7(N-1)}$: 6 DoF SE(3) + 1 DoF log-scale per frame,
        frame 0 pinned at identity (gauge).
  \item Local state reconstruction:
        $T_i = \big(\mathbf{R}_i \mid \mathbf{t}_i\big),\ \log s_i = \lambda_i$
        $\Rightarrow$ world point
        $\mathbf{X}_i^w = e^{\lambda_i}\mathbf{R}_i\,\mathbf{X}_i^i + \mathbf{t}_i.$
\end{itemize}
\vspace{0.5em}
{\footnotesize M3 also uses $\mathfrak{sim}(3)$ on the Lie algebra with the left-plus update operator.
Difference from ours: M3 has $\mathbf{Sim}(3)$ \emph{tracking} against a single reference keyframe plus global BA;
ours is a single monolithic BA over all matched frames.}
\end{frame}

% ------------------------------------------------------------------
\begin{frame}{Sim(3) optimization --- residuals and solver}
\small
\textbf{Ours --- per match} $(m \text{ in } i) \leftrightarrow (n \text{ in } j)$:
\begin{enumerate}
  \item Unproject the i-side patch to world: $\mathbf{X}^w_m = e^{\lambda_i}\mathbf{R}_i\mathbf{X}_i^i(m) + \mathbf{t}_i$.
  \item Reproject into j-camera: $\mathbf{X}_j^{\rm cam}(m) = \mathbf{R}_j^\top(\mathbf{X}^w_m - \mathbf{t}_j)\cdot e^{-\lambda_j}$.
  \item Project to normalized pixel: $\hat{\mathbf{u}}_j(m) = \mathbf{X}_j^{\rm cam}(m)_{[:2]}/\mathbf{X}_j^{\rm cam}(m)_{[2]}$.
  \item Residual: $r_m = \cos(\theta_{mn})\cdot(\mathbf{u}_j^{\rm obs}(n) - \hat{\mathbf{u}}_j(m))\in\mathbb{R}^2$.
\end{enumerate}
Plus a \textbf{velocity smoothness / scale prior}:
$$r_{\rm scale}(i) = w_{\rm prior}\cdot(\lambda_{i+1}-\lambda_i),\qquad w_{\rm prior}=0.1.$$

\vspace{0.4em}
\textbf{Solver:} \texttt{scipy.optimize.least\_squares} with LSMR sparse Jacobian, Huber loss
(scale=$\sigma_{\rm reproj}$), max iterations 500, tolerance $10^{-8}$.

\vspace{0.6em}
\textbf{M3 tracking residual (\texttt{method.tex:215-228}):}
$$
E_{\rm track} = \sum_{(m,n)\in\mathcal{M}(I_f,I_k)}\mathbf{M}_k\,
\rho\!\left(\left\|\frac{\mathbf{p}_{k,n} - \phi(T_{kf}\,\mathbf{X}^f_{f,m})}{w(q_{m,n})}\right\|^2\right)
$$
\textbf{M3 global BA:}
$$
E_g = \sum_{(i,j)\in\mathcal{E}}\sum_{(m,n)\in\mathcal{M}(I_j,I_i)}\mathbf{M}_i\,
\rho\!\left(\left\|\frac{\mathbf{p}_{i,n} - \phi(T_{ij}\,\tilde{\mathbf{X}}^j_{j,m})}{w(q_{m,n})}\right\|^2\right)
$$
Huber robust loss $\rho$, per-match confidence weight $w(q_{m,n})\!=\!\sqrt{\mathbf{Q}^f_f(m)\mathbf{Q}^f_k(n)}$.
\end{frame}

% ==================================================================
%  PART 4 -- RESULTS OF THE PIVOT + PATH FORWARD
% ==================================================================

% ------------------------------------------------------------------
\begin{frame}{Sim(3) frame-BA --- TUM fr1 per-scene}
\framesubtitle{Job 22058142, 9 scenes, SE(3) baseline vs Sim(3) with \texttt{scale\_prior\_weight=0.1}}
\scriptsize
\begin{center}
\begin{tabular}{l r r r l}
\toprule
\textbf{scene} & \textbf{SE(3) ATE} & \textbf{Sim(3) ATE} & \textbf{$\Delta$ (Sim-SE)} & \textbf{winner} \\
\midrule
360    & 0.164 & 0.159 & $-0.005$ & Sim(3) \\
desk   & 0.699 & 0.615 & $-0.085$ & Sim(3) \\
desk2  & \good{0.011} & 0.012 & $+0.001$ & tie \\
floor  & 0.183 & 0.178 & $-0.005$ & Sim(3) \\
plant  & --- & --- & --- & {\color{mute}Umeyama degenerate (TUM data issue)} \\
room   & 0.283 & 0.298 & $+0.016$ & SE(3) \\
rpy    & 0.044 & 0.047 & $+0.003$ & tie \\
teddy  & 0.576 & 0.685 & $+0.109$ & SE(3) \\
xyz    & 0.097 & 0.105 & $+0.008$ & tie \\
\midrule
\textbf{mean} & \textbf{0.257} & \textbf{0.262} & \bad{\textbf{+0.005}} & SE(3) by 0.5 cm \\
\bottomrule
\end{tabular}
\end{center}
\vspace{0.5em}
Sim(3) wins on 360 / desk / floor (scenes with larger rigid drift), loses on teddy (108$\times$108 hand-scale),
and is a wash elsewhere. TUM indoor has almost no scale drift to absorb; the extra DoF over-flexes.
\texttt{scale\_prior\_weight=0.1} is too permissive for TUM (per-frame scale wanders despite the prior).
\end{frame}

% ------------------------------------------------------------------
\begin{frame}{Sim(3) frame-BA --- KITTI 00 full variant table}
\framesubtitle{Job 22058145, seq 00 only, 500 iterations, 109 min wall time}
\small
\textbf{Sim(3)-aligned ATE (m) for all output variants of the same dump:}
\begin{center}
\begin{tabular}{l r l}
\toprule
\textbf{variant} & \textbf{ATE (m)} & \textbf{description} \\
\midrule
stitch     & 173.05  & online Sim(3) stitch, per-submap (default pipeline) \\
raw        & 192.94  & frontend poses, no stitch, no BA \\
offline    & 173.18  & final-pass re-stitch after run ends \\
posegraph  & \good{172.62} & SubmapPoseGraph (gtsam Similarity3 factor graph) \\
frame\_ba  & 173.19  & per-frame Sim(3) BA over the matcher graph \\
\midrule
SE(3) frame\_ba (comparison) & 172.6 & Same runs w/o the 7th DoF \\
\bottomrule
\end{tabular}
\end{center}
\vspace{0.6em}

\textbf{Per-frame scale DoF stats (Sim(3) run):}
\begin{center}
\begin{tabular}{l r r r r}
\toprule
statistic & median & min & max & stdev \\
\midrule
$e^\lambda$ & \hl{0.9999} & 0.83 & 1.02 & 0.013 \\
\bottomrule
\end{tabular}
\end{center}
Cost function dropped \textbf{20$\times$} (8.0e5 $\to$ 4.0e4); ATE moved by 0.5 m. The
extra per-frame scale DoF is barely used because backbone depth is scale-consistent
but \emph{absolute-bias} (Umeyama $\Rightarrow$ 22\% mean depth error over 784-m Karlsruhe loop).
\end{frame}

% ------------------------------------------------------------------
\begin{frame}{Root-cause summary table --- what each hypothesis predicts vs what we see}
\small
\begin{center}
\begin{tabular}{p{4.0cm} l l l}
\toprule
\textbf{hypothesis} & \textbf{prediction} & \textbf{observation} & \textbf{verdict} \\
\midrule
Per-frame scale drift & Sim(3) $\gg$ SE(3) on KITTI & Sim(3) = SE(3) $\pm$ 0.6 m & \bad{FALSIFIED} \\
Retrieval collapse    & small $\Delta t$, small H$\leftrightarrow$H & $\Delta t$ med=60, H$\leftrightarrow$H=0.50--0.83 & \bad{FALSIFIED} \\
Bad overlap stitch    & oracle $\gg$ raw-stitch & oracle 40.7 vs raw 192.9 m (KITTI) & (stitch contributes) \\
Depth absolute bias   & Umeyama scale $\ne 1$, ATE persists & Umeyama 22\% depth error, ATE 172 m & \good{SUPPORTED} \\
Patch-level matching cap & matcher precision capped & TUM@5cm=8\%, ScnNet++=7.8\% & \good{SUPPORTED} \\
\bottomrule
\end{tabular}
\end{center}
\vspace{0.8em}

Two root causes supported:
(1) absolute-scale depth bias of the monocular backbone on outdoor imagery;
(2) 14$\times$14 patch resolution of our matcher cannot resolve 5--20 cm correspondence
for BA to act on. Submap-level stitch knobs do not move the ATE regardless of configuration.
\end{frame}

% ------------------------------------------------------------------
\begin{frame}{Path forward --- prioritized}
\small
\begin{center}
\begin{tabular}{r p{7.3cm} p{3.3cm}}
\toprule
\# & \textbf{action} & \textbf{expected lever} \\
\midrule
1 & Train a pixel-level matching head (DPT+MLP) on V4 backbone, InfoNCE loss vs vKITTI2/ScanNet++ GT correspondences. & Closes TUM@5cm gap (8$\to$$>$50\%). \\
2 & Pose-guided local search at inference: project $\mathbf{X}_i$ into $I_j$, refine inside radius $r$. & Lifts matcher precision under wide disparity. \\
3 & Frame-level SE(3) graph BA with new matcher as 2D-3D residuals. & Sim(3) only buys 0.6 m; lever is matching. \\
4 & Rotation-underfit fix on indoor: loss rebalancing on pose head. & AUC@3$^\circ$ from 0 to paper range. \\
5 & Add real KITTI imagery to V4 training mix. & Closes 10$\times$ outdoor gap (vKITTI2-only). \\
6 & De-prioritize submap-level PGO variants. & They sit on the 15\% substrate regardless. \\
\bottomrule
\end{tabular}
\end{center}
\vspace{0.8em}

\textbf{What we keep as-is:} V4-B'' ckpt-1 backbone, bn\_every=6, domain-specific top-K
(20 TUM / 5 KITTI), Sim(3) stitch + freeze\_history\_writeback, scorer + overlap retrieval,
submap abstraction itself.
\end{frame}

\end{document}