trained_model / reports /insubmap_problem_data_deck.tex
qizhangslam's picture
Upload folder using huggingface_hub
1995ff1 verified
Raw History Blame Contribute Delete
29.1 kB
\documentclass[aspectratio=169,10pt]{beamer}
\usetheme{default}
\usecolortheme{default}
\setbeamertemplate{navigation symbols}{}
\setbeamertemplate{footline}[frame number]
\setbeamerfont{frametitle}{size=\large,series=\bfseries}
\setbeamerfont{framesubtitle}{size=\small,series=\mdseries,shape=\itshape}
\usepackage{booktabs}
\usepackage{array}
\usepackage{colortbl}
\usepackage{xcolor}
\usepackage[utf8]{inputenc}
\usepackage{textcomp}
\usepackage{amsmath}
\usepackage{amssymb}
\definecolor{hl}{RGB}{0,90,158}
\definecolor{bad}{RGB}{178,34,34}
\definecolor{good}{RGB}{0,125,75}
\definecolor{mute}{RGB}{110,110,110}
\newcommand{\hl}[1]{\textcolor{hl}{\textbf{#1}}}
\newcommand{\bad}[1]{\textcolor{bad}{\textbf{#1}}}
\newcommand{\good}[1]{\textcolor{good}{\textbf{#1}}}
\title{The In-Submap Accuracy Problem}
\subtitle{Evidence, Baselines, and the Frame-BA Pivot}
\author{Qi Zhang}
\date{2026-04-21}
\begin{document}
\begin{frame}
\titlepage
\end{frame}
% ==================================================================
% PART 1 -- EVIDENCE (RAW TABLES)
% ==================================================================
% ------------------------------------------------------------------
\begin{frame}{Problem in one sentence}
\vspace{0.5em}
V4 SLAM-Former retrieves loops correctly but the submaps themselves
carry \bad{$\sim$15\% ATE of traveled distance}. Any pose-graph built
on top of this cannot beat the same ceiling.
\vspace{1.0em}
\textbf{What you'll see in this deck:}
\begin{enumerate}
\item Per-submap Sim(3)-aligned ATE --- 10 real scenes (KITTI 00 + 9 TUM fr1).
\item Retrieval overlap --- voxel-based Q$\leftrightarrow$H and H$\leftrightarrow$H on predicted pointclouds.
\item $\Delta t$ retrieval --- KITTI 00 full distribution.
\item M3-SLAM benchmark and per-seq oracle split.
\item Top-K and bn\_every sweeps (full tables).
\item Indoor and outdoor cross-dataset scoreboard.
\item Matcher eval (TUM fr1/xyz, ScanNet++ iPhone) and Sim(3) frame-BA per scene.
\item Methods: descriptor extraction, matching, Sim(3) optimization.
\end{enumerate}
\end{frame}
% ------------------------------------------------------------------
\begin{frame}{V4 SLAM-Former pipeline (context)}
\small
\begin{center}
\begin{tabular}{p{2.2cm}|p{4.5cm}|p{5.6cm}}
\toprule
\textbf{Stage} & \textbf{Unit} & \textbf{What it outputs} \\
\midrule
Backbone & per-frame tokens & depth, pose, feature maps (V4-B'' ckpt-1) \\
Frontend & bn\_every=6 frames & streaming \emph{submap} (local 3D + poses) \\
Scorer & importance\_topk & top-K historical submaps via Q$\cdot$V attention \\
Backend & retrieved submaps & refined submap poses with cross-submap context \\
Sim(3) stitch & submap overlap & chained world trajectory (\texttt{enable\_sim3\_stitch}) \\
\bottomrule
\end{tabular}
\end{center}
\vspace{0.6em}
\textbf{Unit of measurement:} one submap = 6 frames of local pose $+$ pointcloud.
In-submap ATE = Umeyama-Sim(3)-aligned trajectory error inside a single submap,
before any stitching. Measured by \texttt{scripts/insubmap\_accuracy.py}.
\end{frame}
% ------------------------------------------------------------------
\begin{frame}{Evidence A --- per-submap ATE (full table)}
\framesubtitle{V4-B'' ckpt-1, Sim(3) Umeyama, measured 2026-04-21}
\scriptsize
\begin{center}
\begin{tabular}{l r r r r r r}
\toprule
\textbf{scene} & \textbf{\#submaps} & \textbf{frames/sm} & \textbf{median ATE (m)} & \textbf{p95 (m)} & \textbf{traveled/sm (m)} & \textbf{ATE/dist} \\
\midrule
KITTI 00 & 389 & 8 & \bad{0.644} & 1.82 & 5.12 & \bad{14.2\%} \\
\midrule
TUM xyz & 8 & 4 & 0.085 & 0.154 & 0.36 & \bad{31.0\%} \\
TUM desk & 11 & 4 & 0.051 & 0.112 & 0.36 & 12.4\% \\
TUM desk2 & 26 & 4 & 0.063 & 0.114 & 0.51 & 13.0\% \\
TUM room & 52 & 4 & 0.039 & 0.125 & 0.32 & 13.6\% \\
TUM 360 & 39 & 5 & 0.044 & 0.081 & 0.26 & 17.0\% \\
TUM rpy & 36 & 4 & 0.017 & 0.032 & 0.08 & 20.7\% \\
TUM teddy & 23 & 4 & 0.049 & 0.130 & 0.34 & 10.9\% \\
TUM floor & 30 & 4 & 0.040 & 0.110 & 0.26 & 15.7\% \\
TUM plant & 32 & 4 & 0.055 & 0.134 & 0.42 & 15.3\% \\
\midrule
\textbf{TUM mean} & & & \textbf{0.050} & \textbf{0.110} & \textbf{0.32} & \bad{\textbf{16.6\%}} \\
\bottomrule
\end{tabular}
\end{center}
\vspace{0.3em}
{\footnotesize full\_token and top5\_dual\_queue modes produce bitwise-identical dumps under
\texttt{--submap\_fetch\_source frontend} (md5 verified). Numbers hold for both inference modes.}
\end{frame}
% ------------------------------------------------------------------
\begin{frame}{Evidence B --- retrieval overlap (probe2, full sweep)}
\framesubtitle{Voxel-based pairwise overlap on V4-B'' predicted pointclouds}
\scriptsize
\begin{center}
\begin{tabular}{l l r r r r l}
\toprule
\textbf{scene} & \textbf{voxel} & \textbf{Q$\leftrightarrow$H} & \textbf{H$\leftrightarrow$H} & \textbf{frac H$\leftrightarrow$H$>$5\%} & \textbf{\#submap pairs} & \textbf{verdict} \\
\midrule
7Sc office/seq-02 & 0.05 m & 0.339 & \good{0.370} & 84.1\% & --- & healthy \\
TUM fr1/desk & 0.05 m & 0.458 & \good{0.393} & 76.7\% & --- & healthy \\
KITTI 00 & 0.50 m & 0.361 & \good{0.499} & 92.5\% & 389 & healthy \\
KITTI 00 & 1.00 m & 0.444 & \good{0.674} & 98.9\% & 389 & healthy \\
KITTI 00 & 2.00 m & 0.629 & \good{0.829} & 98.9\% & 389 & healthy \\
\midrule
\multicolumn{7}{l}{\textbf{For reference --- hook-translation artifact (probe1, pre-fix):}} \\
7Sc office/seq-02 & 0.05 m & 0.056 & \bad{0.008} & 4.4\% & --- & ARTIFACT \\
7Sc office/seq-02 & 0.40 m & 0.117 & \bad{0.026} & 4.4\% & --- & ARTIFACT \\
\bottomrule
\end{tabular}
\end{center}
\vspace{0.5em}
{\footnotesize
Retrieved sids were row-by-row \emph{identical} between probe1 and probe2; only the
pointcloud-loader got fixed (float \texttt{frame\_id} parsing + timestamp-stem filename lookup).
Early ``7Sc rejected / TTA 14$\times$ lift'' reading was driven entirely by the load miss.
}
\end{frame}
% ------------------------------------------------------------------
\begin{frame}{Evidence C --- $\Delta t$ reach of the scorer (full distribution)}
\framesubtitle{KITTI 00, V4-B'' ckpt-1, importance\_topk=20: 215 submaps, 4070 fetches}
\scriptsize
\textbf{$\Delta t$ distribution over all 4070 fetches:}
\begin{center}
\begin{tabular}{l r r r r r r r}
\toprule
\textbf{$\Delta t$ bin} & 1--3 & 4--9 & 10--19 & 20--49 & 50--99 & 100+ & \textbf{total $\ge 50$} \\
\midrule
share of fetches & 2.8\% & 7.9\% & 9.9\% & 22.9\% & \good{27.4\%} & \good{29.0\%} & \good{56.4\%} \\
\bottomrule
\end{tabular}
\end{center}
\vspace{0.6em}
\textbf{Per-query loop closure rate:}
\begin{center}
\begin{tabular}{l r r r r}
\toprule
\textbf{stat} & median $\Delta t$ & mean $\Delta t$ & max $\Delta t$ & queries with $\exists\Delta t\ge 50$ \\
\midrule
KITTI 00 (k=20) & 60 & 70.3 & 214 & \good{76.7\%} \\
KITTI 02 (k=10) & 48 & --- & --- & 72.7\% \\
TUM fr1/teddy (topall, bn=4) & 11 & --- & 34 & 0\% (seq too short) \\
\bottomrule
\end{tabular}
\end{center}
\vspace{0.6em}
\textbf{Anchor-dominance sanity check (KITTI 00):} fetched sid=0 in only 5.2\% of the 4070 fetches.
Top-fetched $\Delta t$ values: 2, 3, 5--11 ($\sim$50 fetches each) --- recent history is fetched
but the long tail is thick. Training-time finetune only supervises $\Delta t\le 3$; inference runs out to 214.
\end{frame}
% ==================================================================
% PART 2 -- BASELINE SCOREBOARD
% ==================================================================
% ------------------------------------------------------------------
\begin{frame}{M3-SLAM benchmark --- full paper table}
\framesubtitle{ATE RMSE (m) -- lower is better. Source: m3-slam/sec/experiments.tex.}
\small
\begin{center}
\begin{tabular}{l r r r r r r}
\toprule
\textbf{Dataset} & MonoGS & DROID-SLAM & LEAP-VO & VGGT-SLAM & Pi3X-slam & \textbf{M3-SLAM} \\
\midrule
ScanNet++ (20 scenes) & 0.623 & 0.346 & 1.081 & 0.182 & 0.137 & \good{0.065} \\
ScanNetV2 (10 scenes) & 0.217 & 0.106 & 0.852 & 0.073 & 0.141 & \good{0.051} \\
Waymo (9 seqs) & 7.535 & 16.61 & 13.17 & 1.295 & 2.410 & \good{0.773} \\
KITTI (8 seqs) & 8.212 & 18.07 & 10.99 & 2.521 & 2.938 & \good{0.890} \\
\bottomrule
\end{tabular}
\end{center}
\vspace{0.8em}
\textbf{V4-B'' ckpt-1 on the same axis:}
\begin{center}
\begin{tabular}{l r r r}
\toprule
\textbf{Dataset} & V4-B'' ckpt-1 & M3-SLAM & gap \\
\midrule
TUM fr1 mean (9 scenes) & 0.224 & \textit{n/a} & same order as M3 indoor \\
KITTI 00 (frame-BA floor) & \bad{172.6} & 0.890 & \bad{194$\times$} \\
\bottomrule
\end{tabular}
\end{center}
\textit{Same backbone family (Pi3X-class), same monocular streaming regime. Architectural
differences only: matching head, pose-guided local search, Sim(3) global BA.}
\end{frame}
% ------------------------------------------------------------------
\begin{frame}{Oracle chain-TRS --- full per-backbone split}
\framesubtitle{V4-oracle chain-TRS(6,4): oracle substitutes GT poses for 6-frame submap overlap alignment}
\scriptsize
\textbf{KITTI (FINAL, avg over 11 sequences, lower is better):}
\begin{center}
\begin{tabular}{l r r r}
\toprule
\textbf{Backbone} & top-5 ATE (m) & \#seqs won vs others & per-seq winners \\
\midrule
V4-B' (paper + scannetpp + mvs\_synth + vK) & 72.5 & 0/11 & --- \\
V4-B'' (V3 ckpt-2 + same data mix) & \good{41.3} & 6/11 & 02, 03, 04, 07, 08, 10 \\
no-vK (paper + scannetpp + mvs\_synth) & \good{40.7} & 5/11 & 00, 01, 05, 06, 09 \\
\bottomrule
\end{tabular}
\end{center}
\vspace{0.4em}
\textbf{TUM fr1 (avg over 8 sequences):}
\begin{center}
\begin{tabular}{l r r r}
\toprule
\textbf{Backbone} & top-5 & top-all & gap to paper target (0.045) \\
\midrule
V4-B' & 0.183 & 0.165 & 3.7$\times$ \\
V4-B'' & 0.160 & \good{0.102} & 2.3$\times$ \\
no-vK & 0.145 & 0.136 & 3.0$\times$ \\
\bottomrule
\end{tabular}
\end{center}
\vspace{0.4em}
With \emph{perfect} overlap alignment we still get 40.7 m KITTI / 0.102 m TUM.
Stitch is not the bottleneck; the 15\% in-submap ATE is.
\end{frame}
% ------------------------------------------------------------------
\begin{frame}{Ablation --- top-K retrieval, full sweep}
\framesubtitle{\S15.13, V4-oracle chain-TRS, bn\_every=6 fixed, same ckpts as scoreboard}
\small
\textbf{TUM fr1 (V4-B'' ckpt-1, all 8/8 seqs complete):}
\begin{center}
\begin{tabular}{l r r r r r}
\toprule
top-K & 5 & 10 & \textbf{20} & 50 & all \\
\midrule
mean ATE (m) & 0.160 & 0.106 & \good{0.100} & 0.102 & 0.102 \\
\bottomrule
\end{tabular}
\end{center}
\textit{Monotone down k=5$\to$10$\to$20, then plateau. TUM sweet spot = k=20.}
\vspace{0.8em}
\textbf{KITTI (no-vK ckpt-1, 11/11 at k=5/10/20, 6/11 at k=50):}
\begin{center}
\begin{tabular}{l r r r r}
\toprule
top-K & \textbf{5} & 10 & 20 & 50 \\
\midrule
mean ATE (m) & \good{40.7} & 48.5 & 56.2 & 69.4 \\
\bottomrule
\end{tabular}
\end{center}
\textit{Monotone \emph{up} --- adding history \emph{hurts}. KITTI sweet spot = k=5.}
\vspace{0.6em}
\textbf{Domain-specific top-K is required.} KITTI forward motion has no persistent
loop back until kilometre scale, so large top-K just injects irrelevant tokens.
TUM sees the same scene from many angles; more retrieval = more coverage.
\end{frame}
% ------------------------------------------------------------------
\begin{frame}{Ablation --- frames per submap (\texttt{bn\_every})}
\framesubtitle{\S15.14, V4-oracle chain-TRS, top-K=5 fixed}
\small
\begin{center}
\begin{tabular}{l r r r l}
\toprule
\textbf{dataset} & bn=4 & \textbf{bn=6} & bn=10 & coverage \\
\midrule
TUM fr1 (V4-B'') & 0.174 & \good{0.160} & 0.170 & 8/8 at every bn \\
KITTI (no-vK) & 75.7 & \good{40.7} & 77.5 & bn4=8/11, bn6=11/11, bn10=9/11 \\
\bottomrule
\end{tabular}
\end{center}
\vspace{0.8em}
\textbf{Both domains prefer the default bn=6.}
bn=4 fragments context before the scorer can use it; bn=10 accumulates more depth bias
inside the window. Same pattern as top-K: neither knob closes the gap to M3.
\vspace{0.6em}
\textbf{Canonical config} (carried through all subsequent experiments): V4-B'' backbone
+ bn\_every=6 + top-K=20 on TUM / top-K=5 on KITTI.
\end{frame}
% ------------------------------------------------------------------
\begin{frame}{Cross-dataset scoreboard --- indoor (V4-B'' ckpt-1)}
\framesubtitle{Evaluated 2026-04-20}
\small
\begin{center}
\begin{tabular}{l r r r r r}
\toprule
\textbf{benchmark} & \textbf{\#seqs} & \textbf{mean ATE (m)} & \textbf{AUC@3$^\circ$} & \textbf{AUC@30$^\circ$} & \textbf{status} \\
\midrule
7-Scenes & 17 & 0.193 & $\sim$0.05 & 0.642 & ATE at paper parity, rotation weak \\
NRGBD & 8 & 0.252 & $\sim$0.01 & 0.780 & ATE at paper parity (OpenGL fix required) \\
TUM fr1 & 9 & 0.224 & \bad{0.000} & 0.368 & AUC@3$^\circ$=0 systemic \\
\bottomrule
\end{tabular}
\end{center}
\vspace{0.8em}
\textbf{TUM fr1 per-scene ATE (V4-B'' ckpt-1):}
\begin{center}
\begin{tabular}{l r r r r r r r r r r}
\toprule
scene & 360 & desk & desk2 & floor & plant & room & rpy & teddy & xyz & \textbf{mean} \\
\midrule
ATE (m) & 0.17 & 0.11 & \good{0.017} & 0.13 & --- & \bad{0.73} & 0.24 & 0.23 & 0.10 & 0.224 \\
\bottomrule
\end{tabular}
\end{center}
\vspace{0.3em}
{\footnotesize Rotation AUC@3$^\circ$=0 even on scenes where translation is cm-scale $\Rightarrow$
indoor rotation is under-fit, not a geometric failure.}
\end{frame}
% ------------------------------------------------------------------
\begin{frame}{Cross-dataset scoreboard --- outdoor (V4-B'' ckpt-1)}
\framesubtitle{Evaluated 2026-04-20}
\small
\begin{center}
\begin{tabular}{l r r r r}
\toprule
\textbf{benchmark} & \textbf{mean ATE (m)} & \textbf{paper target (m)} & \textbf{M3 (m)} & \textbf{gap vs paper / M3} \\
\midrule
ETH3D (median over 13 scenes) & 0.63 & --- & --- & AUC@3$^\circ$=0.01 vs paper 0.67 ($30\times$) \\
Oxford Spires / keble & 25.25 & 6--12 & --- & 2--4$\times$ \\
Oxford Spires / christ & 36.15 & 6--12 & --- & 3--6$\times$ \\
Oxford Spires mean & 30.7 & 6--12 & --- & 3--5$\times$ \\
KITTI (mean) & \bad{148} & 14.9 & 0.89 & \bad{10$\times$ paper, 166$\times$ M3} \\
\bottomrule
\end{tabular}
\end{center}
\vspace{0.8em}
\textbf{Oxford Spires alignment coverage:} n\_matched $\approx$ 100 / 3840 patches
$\Rightarrow$ \bad{2.6\% alignment coverage}. Most frames never align under the current
matcher. The 25--36 m errors are driven by unaligned frames, not by local geometry
precision.
\vspace{0.3em}
{\footnotesize ETH3D F1$>$0 on 4/13 scenes only. Bad F1 is a coverage issue, not a depth one.}
\end{frame}
% ==================================================================
% PART 3 -- METHODS (MATCHING + DESCRIPTORS + SIM(3) OPTIMIZATION)
% ==================================================================
% ------------------------------------------------------------------
\begin{frame}{Matching system pipeline}
\vspace{-0.3em}
\begin{center}
\resizebox{\linewidth}{!}{\includegraphics{matching_pipeline_figure.pdf}}
\end{center}
\vspace{-0.7em}
{\footnotesize
\textbf{Four stages, all inference-time, no training:}
(1) SLAM-Former backbone returns per-patch tokens $F\in\mathbb{R}^{P\times d}$ and local 3D $X\in\mathbb{R}^{P\times 3}$;
(2) rank patches by confidence $c_p$, keep top-$K$, L2-normalize the token as the descriptor;
(3) mutual-NN cosine match + Lowe ratio + $3\!\times\!3$ quadratic sub-pixel refinement;
(4) Sim(3) frame BA (scipy LSMR, 7 DoF per frame, velocity-smoothness scale prior $\lambda_{\rm scale}=0.1$).
No learned matching head; reuses backbone tokens as features.}
\end{frame}
% ------------------------------------------------------------------
\begin{frame}{Matching --- how descriptors are extracted}
\small
\textbf{Ours (\texttt{frame\_matcher.py:select\_frame\_keypoints}):}
\begin{itemize}\small
\item Backbone returns per-patch tokens $\mathbf{t}_{\rm patch}\in\mathbb{R}^{H_p\times W_p\times D}$
($H_p=W_p=37$ at target\_size=518, patch=14$\times$14 px, $D$=token dim).
\item Per-patch confidence: mean of pixel-wise conf over the 14$\times$14 cell.
\item Patch-center 3D point $\mathbf{X}_i^i\in\mathbb{R}^3$ from the predicted local pointcloud;
patch-center normalized camera UV $(u,v) = (x/z, y/z)$ of that 3D point.
\item Top-$K$ patches by confidence, with $z\in(z_{\min},z_{\max})$ finite.
\item Descriptor = patch token, L2-normalized, stored as fp16.
\end{itemize}
\vspace{0.6em}
\textbf{M3-SLAM (from \texttt{refer/m3-slam/sec/method.tex}):}
\begin{itemize}\small
\item Separate \textbf{matching head} on top of Pi3X: $\mathrm{DPT}_{\rm desc}+$2-layer MLP with GELU.
\item Outputs dense descriptors $\mathbf{D}\in\mathbb{R}^{H\times W\times d}$
($d=24$ by default, \emph{per pixel}, not per patch) and matching confidence
$\mathbf{Q}\in\mathbb{R}^{H\times W\times 1}$.
\item Trained with InfoNCE symmetric loss against GT pixel correspondences;
encoder / decoder / point head frozen.
\end{itemize}
\vspace{0.4em}
{\footnotesize \textbf{The structural gap:} our ``descriptor'' is a token over a
14$\times$14 patch, giving $37^2 \approx 1369$ candidates per frame; M3's is
\emph{per pixel}, giving $518^2 \approx 268{,}000$ candidates per frame. A single
token integrates over $\sim$200 pixels of depth variation.}
\end{frame}
% ------------------------------------------------------------------
\begin{frame}{Matching --- how pairs are formed}
\small
\textbf{Ours (\texttt{frame\_matcher.py:match\_frames}):}
\begin{enumerate}\small
\item Mutual-nearest-neighbour on cosine similarity between L2-normalized token descriptors.
\item Lowe's ratio test on cosine: keep pair if $\cos_{\rm top2}\le r\cdot\cos_{\rm top1}$,
typical $r\in[0.8,0.9]$.
\item Sub-pixel refinement via quadratic fit on the 3$\times$3 cos-sim neighborhood of
the peak, offset clipped to $[-0.5,+0.5]$ patch units.
\item Output: \texttt{MatchEdge} with indices $i\leftrightarrow j$, cos-sim, and refined $(u_j,v_j)$.
\end{enumerate}
\vspace{0.6em}
\textbf{M3-SLAM (Sec. 3.1.3 --- ``Dense matching for SLAM''):}
\begin{enumerate}\small
\item Pose-guided initialization: project $\mathbf{X}_i$ into frame $j$ using current $(T_i, T_j)$
$$\mathbf{X}_i^j = T_j^{-1}\,T_i\,\mathbf{X}_i.$$
\item For each pixel $\mathbf{p}_i$, search a local window of radius $r$ around the projected location
in frame $j$, selecting $\mathbf{p}^*_j = \arg\max_{\mathbf{p}\in\mathcal{N}_r} \big\langle \mathbf{D}_i(\mathbf{p}_i), \mathbf{D}_j(\mathbf{p})\big\rangle$
(cosine similarity).
\item Dynamic masking: a motion map $\mathbf{M}_i\in[0,1]^{H\times W}$ down-weights pixels with
low warped-vs-observed descriptor consistency.
\end{enumerate}
\vspace{0.4em}
{\footnotesize \textbf{Key difference:} M3's matching is $O(N\cdot r^2)$ local geometric search;
ours is $O(K^2)$ global descriptor matching with no pose prior. M3's search radius $r$ can be
as small as a few pixels \emph{because the pose initialization is good}.}
\end{frame}
% ------------------------------------------------------------------
\begin{frame}{Matcher precision --- how it actually performs}
\framesubtitle{GT-supervised M3-style eval (\texttt{eval\_matcher\_gt.py}), $\tau=0.6$, err threshold 0.10 m}
\small
\textbf{TUM fr1/xyz (48 frames, patch-level):}
\begin{center}
\begin{tabular}{l r r r r}
\toprule
err threshold & 5 cm & 10 cm & 20 cm & 50 cm \\
\midrule
precision & \bad{$\sim$8\%} & --- & 55\% & 88\% \\
\bottomrule
\end{tabular}
\end{center}
median patch-pair 3D error $\approx$ \textbf{19 cm} (patches = 14$\times$14 px $\approx$ 17 original pixels).
\vspace{0.8em}
\textbf{ScanNet++ matcher eval (\#22058916, STRIDE=5 iPhone frames, 3 scenes):}
\begin{center}
\begin{tabular}{l r r r r r}
\toprule
\textbf{scene} & \textbf{\#pairs} & \textbf{\#matches} & \textbf{\#correct} & \textbf{precision} & \textbf{median err (m)} \\
\midrule
036bce3393 & 51 & 1292 & 75 & \bad{5.8\%} & 1.99 \\
079a326597 & 41 & 1319 & 172 & \bad{13.0\%} & 0.68 \\
07ff1c45bb & 53 & 1286 & 59 & \bad{4.6\%} & 1.31 \\
\midrule
mean & 48 & 1299 & 102 & \bad{7.8\%} & 1.33 \\
\bottomrule
\end{tabular}
\end{center}
\vspace{0.3em}
{\footnotesize iPhone STRIDE=5 keyframe gaps of 0.3--0.8 s produce wide view changes.
Precision 5--13\% vs TUM 30\% baseline is the natural consequence of patch-level
descriptors under large disparity.}
\end{frame}
% ------------------------------------------------------------------
\begin{frame}{Sim(3) optimization --- parameterization}
\small
\textbf{Pose as a similarity transform} $T\in \mathbf{Sim}(3)$:
$$
T = \begin{bmatrix} s\mathbf{R} & \mathbf{t} \\ \mathbf{0}^\top & 1 \end{bmatrix},
\qquad \mathbf{R}\in\mathbf{SO}(3),\ \mathbf{t}\in\mathbb{R}^3,\ s\in\mathbb{R}_{>0}
$$
\vspace{0.4em}
\textbf{Lie algebra representation} $\boldsymbol{\tau}\in\mathfrak{sim}(3)\cong\mathbb{R}^7$:
\begin{itemize}
\item 3 for axis-angle rotation $\boldsymbol{\omega}\in\mathbb{R}^3$, 3 for translation
$\boldsymbol{\rho}\in\mathbb{R}^3$, 1 for log-scale $\lambda\in\mathbb{R}$.
\item Update via left-multiplication on the manifold:
$T \leftarrow \exp(\boldsymbol{\tau})\circ T$.
\end{itemize}
\vspace{0.8em}
\textbf{Ours (\texttt{frame\_ba.py:FrameBA}, \texttt{use\_sim3=True}):}
\begin{itemize}
\item State vector $\boldsymbol{\theta}\in\mathbb{R}^{7(N-1)}$: 6 DoF SE(3) + 1 DoF log-scale per frame,
frame 0 pinned at identity (gauge).
\item Local state reconstruction:
$T_i = \big(\mathbf{R}_i \mid \mathbf{t}_i\big),\ \log s_i = \lambda_i$
$\Rightarrow$ world point
$\mathbf{X}_i^w = e^{\lambda_i}\mathbf{R}_i\,\mathbf{X}_i^i + \mathbf{t}_i.$
\end{itemize}
\vspace{0.5em}
{\footnotesize M3 also uses $\mathfrak{sim}(3)$ on the Lie algebra with the left-plus update operator.
Difference from ours: M3 has $\mathbf{Sim}(3)$ \emph{tracking} against a single reference keyframe plus global BA;
ours is a single monolithic BA over all matched frames.}
\end{frame}
% ------------------------------------------------------------------
\begin{frame}{Sim(3) optimization --- residuals and solver}
\small
\textbf{Ours --- per match} $(m \text{ in } i) \leftrightarrow (n \text{ in } j)$:
\begin{enumerate}
\item Unproject the i-side patch to world: $\mathbf{X}^w_m = e^{\lambda_i}\mathbf{R}_i\mathbf{X}_i^i(m) + \mathbf{t}_i$.
\item Reproject into j-camera: $\mathbf{X}_j^{\rm cam}(m) = \mathbf{R}_j^\top(\mathbf{X}^w_m - \mathbf{t}_j)\cdot e^{-\lambda_j}$.
\item Project to normalized pixel: $\hat{\mathbf{u}}_j(m) = \mathbf{X}_j^{\rm cam}(m)_{[:2]}/\mathbf{X}_j^{\rm cam}(m)_{[2]}$.
\item Residual: $r_m = \cos(\theta_{mn})\cdot(\mathbf{u}_j^{\rm obs}(n) - \hat{\mathbf{u}}_j(m))\in\mathbb{R}^2$.
\end{enumerate}
Plus a \textbf{velocity smoothness / scale prior}:
$$r_{\rm scale}(i) = w_{\rm prior}\cdot(\lambda_{i+1}-\lambda_i),\qquad w_{\rm prior}=0.1.$$
\vspace{0.4em}
\textbf{Solver:} \texttt{scipy.optimize.least\_squares} with LSMR sparse Jacobian, Huber loss
(scale=$\sigma_{\rm reproj}$), max iterations 500, tolerance $10^{-8}$.
\vspace{0.6em}
\textbf{M3 tracking residual (\texttt{method.tex:215-228}):}
$$
E_{\rm track} = \sum_{(m,n)\in\mathcal{M}(I_f,I_k)}\mathbf{M}_k\,
\rho\!\left(\left\|\frac{\mathbf{p}_{k,n} - \phi(T_{kf}\,\mathbf{X}^f_{f,m})}{w(q_{m,n})}\right\|^2\right)
$$
\textbf{M3 global BA:}
$$
E_g = \sum_{(i,j)\in\mathcal{E}}\sum_{(m,n)\in\mathcal{M}(I_j,I_i)}\mathbf{M}_i\,
\rho\!\left(\left\|\frac{\mathbf{p}_{i,n} - \phi(T_{ij}\,\tilde{\mathbf{X}}^j_{j,m})}{w(q_{m,n})}\right\|^2\right)
$$
Huber robust loss $\rho$, per-match confidence weight $w(q_{m,n})\!=\!\sqrt{\mathbf{Q}^f_f(m)\mathbf{Q}^f_k(n)}$.
\end{frame}
% ==================================================================
% PART 4 -- RESULTS OF THE PIVOT + PATH FORWARD
% ==================================================================
% ------------------------------------------------------------------
\begin{frame}{Sim(3) frame-BA --- TUM fr1 per-scene}
\framesubtitle{Job 22058142, 9 scenes, SE(3) baseline vs Sim(3) with \texttt{scale\_prior\_weight=0.1}}
\scriptsize
\begin{center}
\begin{tabular}{l r r r l}
\toprule
\textbf{scene} & \textbf{SE(3) ATE} & \textbf{Sim(3) ATE} & \textbf{$\Delta$ (Sim-SE)} & \textbf{winner} \\
\midrule
360 & 0.164 & 0.159 & $-0.005$ & Sim(3) \\
desk & 0.699 & 0.615 & $-0.085$ & Sim(3) \\
desk2 & \good{0.011} & 0.012 & $+0.001$ & tie \\
floor & 0.183 & 0.178 & $-0.005$ & Sim(3) \\
plant & --- & --- & --- & {\color{mute}Umeyama degenerate (TUM data issue)} \\
room & 0.283 & 0.298 & $+0.016$ & SE(3) \\
rpy & 0.044 & 0.047 & $+0.003$ & tie \\
teddy & 0.576 & 0.685 & $+0.109$ & SE(3) \\
xyz & 0.097 & 0.105 & $+0.008$ & tie \\
\midrule
\textbf{mean} & \textbf{0.257} & \textbf{0.262} & \bad{\textbf{+0.005}} & SE(3) by 0.5 cm \\
\bottomrule
\end{tabular}
\end{center}
\vspace{0.5em}
Sim(3) wins on 360 / desk / floor (scenes with larger rigid drift), loses on teddy (108$\times$108 hand-scale),
and is a wash elsewhere. TUM indoor has almost no scale drift to absorb; the extra DoF over-flexes.
\texttt{scale\_prior\_weight=0.1} is too permissive for TUM (per-frame scale wanders despite the prior).
\end{frame}
% ------------------------------------------------------------------
\begin{frame}{Sim(3) frame-BA --- KITTI 00 full variant table}
\framesubtitle{Job 22058145, seq 00 only, 500 iterations, 109 min wall time}
\small
\textbf{Sim(3)-aligned ATE (m) for all output variants of the same dump:}
\begin{center}
\begin{tabular}{l r l}
\toprule
\textbf{variant} & \textbf{ATE (m)} & \textbf{description} \\
\midrule
stitch & 173.05 & online Sim(3) stitch, per-submap (default pipeline) \\
raw & 192.94 & frontend poses, no stitch, no BA \\
offline & 173.18 & final-pass re-stitch after run ends \\
posegraph & \good{172.62} & SubmapPoseGraph (gtsam Similarity3 factor graph) \\
frame\_ba & 173.19 & per-frame Sim(3) BA over the matcher graph \\
\midrule
SE(3) frame\_ba (comparison) & 172.6 & Same runs w/o the 7th DoF \\
\bottomrule
\end{tabular}
\end{center}
\vspace{0.6em}
\textbf{Per-frame scale DoF stats (Sim(3) run):}
\begin{center}
\begin{tabular}{l r r r r}
\toprule
statistic & median & min & max & stdev \\
\midrule
$e^\lambda$ & \hl{0.9999} & 0.83 & 1.02 & 0.013 \\
\bottomrule
\end{tabular}
\end{center}
Cost function dropped \textbf{20$\times$} (8.0e5 $\to$ 4.0e4); ATE moved by 0.5 m. The
extra per-frame scale DoF is barely used because backbone depth is scale-consistent
but \emph{absolute-bias} (Umeyama $\Rightarrow$ 22\% mean depth error over 784-m Karlsruhe loop).
\end{frame}
% ------------------------------------------------------------------
\begin{frame}{Root-cause summary table --- what each hypothesis predicts vs what we see}
\small
\begin{center}
\begin{tabular}{p{4.0cm} l l l}
\toprule
\textbf{hypothesis} & \textbf{prediction} & \textbf{observation} & \textbf{verdict} \\
\midrule
Per-frame scale drift & Sim(3) $\gg$ SE(3) on KITTI & Sim(3) = SE(3) $\pm$ 0.6 m & \bad{FALSIFIED} \\
Retrieval collapse & small $\Delta t$, small H$\leftrightarrow$H & $\Delta t$ med=60, H$\leftrightarrow$H=0.50--0.83 & \bad{FALSIFIED} \\
Bad overlap stitch & oracle $\gg$ raw-stitch & oracle 40.7 vs raw 192.9 m (KITTI) & (stitch contributes) \\
Depth absolute bias & Umeyama scale $\ne 1$, ATE persists & Umeyama 22\% depth error, ATE 172 m & \good{SUPPORTED} \\
Patch-level matching cap & matcher precision capped & TUM@5cm=8\%, ScnNet++=7.8\% & \good{SUPPORTED} \\
\bottomrule
\end{tabular}
\end{center}
\vspace{0.8em}
Two root causes supported:
(1) absolute-scale depth bias of the monocular backbone on outdoor imagery;
(2) 14$\times$14 patch resolution of our matcher cannot resolve 5--20 cm correspondence
for BA to act on. Submap-level stitch knobs do not move the ATE regardless of configuration.
\end{frame}
% ------------------------------------------------------------------
\begin{frame}{Path forward --- prioritized}
\small
\begin{center}
\begin{tabular}{r p{7.3cm} p{3.3cm}}
\toprule
\# & \textbf{action} & \textbf{expected lever} \\
\midrule
1 & Train a pixel-level matching head (DPT+MLP) on V4 backbone, InfoNCE loss vs vKITTI2/ScanNet++ GT correspondences. & Closes TUM@5cm gap (8$\to$$>$50\%). \\
2 & Pose-guided local search at inference: project $\mathbf{X}_i$ into $I_j$, refine inside radius $r$. & Lifts matcher precision under wide disparity. \\
3 & Frame-level SE(3) graph BA with new matcher as 2D-3D residuals. & Sim(3) only buys 0.6 m; lever is matching. \\
4 & Rotation-underfit fix on indoor: loss rebalancing on pose head. & AUC@3$^\circ$ from 0 to paper range. \\
5 & Add real KITTI imagery to V4 training mix. & Closes 10$\times$ outdoor gap (vKITTI2-only). \\
6 & De-prioritize submap-level PGO variants. & They sit on the 15\% substrate regardless. \\
\bottomrule
\end{tabular}
\end{center}
\vspace{0.8em}
\textbf{What we keep as-is:} V4-B'' ckpt-1 backbone, bn\_every=6, domain-specific top-K
(20 TUM / 5 KITTI), Sim(3) stitch + freeze\_history\_writeback, scorer + overlap retrieval,
submap abstraction itself.
\end{frame}
\end{document}