Download reports/insubmap_problem_data_deck.tex from qizhangslam/trained_model: direct link, hf CLI and curl.
- Browser
- Download file 29.1 kB
-
https://huggingface.co/qizhangslam/trained_model/resolve/main/reports/insubmap_problem_data_deck.tex
- Command line
-
hf download hf://qizhangslam/trained_model/reports/insubmap_problem_data_deck.tex
-
curl -L -o insubmap_problem_data_deck.tex https://huggingface.co/qizhangslam/trained_model/resolve/main/reports/insubmap_problem_data_deck.tex
29.1 kB
| \documentclass[aspectratio=169,10pt]{beamer} | |
| \usetheme{default} | |
| \usecolortheme{default} | |
| \setbeamertemplate{navigation symbols}{} | |
| \setbeamertemplate{footline}[frame number] | |
| \setbeamerfont{frametitle}{size=\large,series=\bfseries} | |
| \setbeamerfont{framesubtitle}{size=\small,series=\mdseries,shape=\itshape} | |
| \usepackage{booktabs} | |
| \usepackage{array} | |
| \usepackage{colortbl} | |
| \usepackage{xcolor} | |
| \usepackage[utf8]{inputenc} | |
| \usepackage{textcomp} | |
| \usepackage{amsmath} | |
| \usepackage{amssymb} | |
| \definecolor{hl}{RGB}{0,90,158} | |
| \definecolor{bad}{RGB}{178,34,34} | |
| \definecolor{good}{RGB}{0,125,75} | |
| \definecolor{mute}{RGB}{110,110,110} | |
| \newcommand{\hl}[1]{\textcolor{hl}{\textbf{#1}}} | |
| \newcommand{\bad}[1]{\textcolor{bad}{\textbf{#1}}} | |
| \newcommand{\good}[1]{\textcolor{good}{\textbf{#1}}} | |
| \title{The In-Submap Accuracy Problem} | |
| \subtitle{Evidence, Baselines, and the Frame-BA Pivot} | |
| \author{Qi Zhang} | |
| \date{2026-04-21} | |
| \begin{document} | |
| \begin{frame} | |
| \titlepage | |
| \end{frame} | |
| % ================================================================== | |
| % PART 1 -- EVIDENCE (RAW TABLES) | |
| % ================================================================== | |
| % ------------------------------------------------------------------ | |
| \begin{frame}{Problem in one sentence} | |
| \vspace{0.5em} | |
| V4 SLAM-Former retrieves loops correctly but the submaps themselves | |
| carry \bad{$\sim$15\% ATE of traveled distance}. Any pose-graph built | |
| on top of this cannot beat the same ceiling. | |
| \vspace{1.0em} | |
| \textbf{What you'll see in this deck:} | |
| \begin{enumerate} | |
| \item Per-submap Sim(3)-aligned ATE --- 10 real scenes (KITTI 00 + 9 TUM fr1). | |
| \item Retrieval overlap --- voxel-based Q$\leftrightarrow$H and H$\leftrightarrow$H on predicted pointclouds. | |
| \item $\Delta t$ retrieval --- KITTI 00 full distribution. | |
| \item M3-SLAM benchmark and per-seq oracle split. | |
| \item Top-K and bn\_every sweeps (full tables). | |
| \item Indoor and outdoor cross-dataset scoreboard. | |
| \item Matcher eval (TUM fr1/xyz, ScanNet++ iPhone) and Sim(3) frame-BA per scene. | |
| \item Methods: descriptor extraction, matching, Sim(3) optimization. | |
| \end{enumerate} | |
| \end{frame} | |
| % ------------------------------------------------------------------ | |
| \begin{frame}{V4 SLAM-Former pipeline (context)} | |
| \small | |
| \begin{center} | |
| \begin{tabular}{p{2.2cm}|p{4.5cm}|p{5.6cm}} | |
| \toprule | |
| \textbf{Stage} & \textbf{Unit} & \textbf{What it outputs} \\ | |
| \midrule | |
| Backbone & per-frame tokens & depth, pose, feature maps (V4-B'' ckpt-1) \\ | |
| Frontend & bn\_every=6 frames & streaming \emph{submap} (local 3D + poses) \\ | |
| Scorer & importance\_topk & top-K historical submaps via Q$\cdot$V attention \\ | |
| Backend & retrieved submaps & refined submap poses with cross-submap context \\ | |
| Sim(3) stitch & submap overlap & chained world trajectory (\texttt{enable\_sim3\_stitch}) \\ | |
| \bottomrule | |
| \end{tabular} | |
| \end{center} | |
| \vspace{0.6em} | |
| \textbf{Unit of measurement:} one submap = 6 frames of local pose $+$ pointcloud. | |
| In-submap ATE = Umeyama-Sim(3)-aligned trajectory error inside a single submap, | |
| before any stitching. Measured by \texttt{scripts/insubmap\_accuracy.py}. | |
| \end{frame} | |
| % ------------------------------------------------------------------ | |
| \begin{frame}{Evidence A --- per-submap ATE (full table)} | |
| \framesubtitle{V4-B'' ckpt-1, Sim(3) Umeyama, measured 2026-04-21} | |
| \scriptsize | |
| \begin{center} | |
| \begin{tabular}{l r r r r r r} | |
| \toprule | |
| \textbf{scene} & \textbf{\#submaps} & \textbf{frames/sm} & \textbf{median ATE (m)} & \textbf{p95 (m)} & \textbf{traveled/sm (m)} & \textbf{ATE/dist} \\ | |
| \midrule | |
| KITTI 00 & 389 & 8 & \bad{0.644} & 1.82 & 5.12 & \bad{14.2\%} \\ | |
| \midrule | |
| TUM xyz & 8 & 4 & 0.085 & 0.154 & 0.36 & \bad{31.0\%} \\ | |
| TUM desk & 11 & 4 & 0.051 & 0.112 & 0.36 & 12.4\% \\ | |
| TUM desk2 & 26 & 4 & 0.063 & 0.114 & 0.51 & 13.0\% \\ | |
| TUM room & 52 & 4 & 0.039 & 0.125 & 0.32 & 13.6\% \\ | |
| TUM 360 & 39 & 5 & 0.044 & 0.081 & 0.26 & 17.0\% \\ | |
| TUM rpy & 36 & 4 & 0.017 & 0.032 & 0.08 & 20.7\% \\ | |
| TUM teddy & 23 & 4 & 0.049 & 0.130 & 0.34 & 10.9\% \\ | |
| TUM floor & 30 & 4 & 0.040 & 0.110 & 0.26 & 15.7\% \\ | |
| TUM plant & 32 & 4 & 0.055 & 0.134 & 0.42 & 15.3\% \\ | |
| \midrule | |
| \textbf{TUM mean} & & & \textbf{0.050} & \textbf{0.110} & \textbf{0.32} & \bad{\textbf{16.6\%}} \\ | |
| \bottomrule | |
| \end{tabular} | |
| \end{center} | |
| \vspace{0.3em} | |
| {\footnotesize full\_token and top5\_dual\_queue modes produce bitwise-identical dumps under | |
| \texttt{--submap\_fetch\_source frontend} (md5 verified). Numbers hold for both inference modes.} | |
| \end{frame} | |
| % ------------------------------------------------------------------ | |
| \begin{frame}{Evidence B --- retrieval overlap (probe2, full sweep)} | |
| \framesubtitle{Voxel-based pairwise overlap on V4-B'' predicted pointclouds} | |
| \scriptsize | |
| \begin{center} | |
| \begin{tabular}{l l r r r r l} | |
| \toprule | |
| \textbf{scene} & \textbf{voxel} & \textbf{Q$\leftrightarrow$H} & \textbf{H$\leftrightarrow$H} & \textbf{frac H$\leftrightarrow$H$>$5\%} & \textbf{\#submap pairs} & \textbf{verdict} \\ | |
| \midrule | |
| 7Sc office/seq-02 & 0.05 m & 0.339 & \good{0.370} & 84.1\% & --- & healthy \\ | |
| TUM fr1/desk & 0.05 m & 0.458 & \good{0.393} & 76.7\% & --- & healthy \\ | |
| KITTI 00 & 0.50 m & 0.361 & \good{0.499} & 92.5\% & 389 & healthy \\ | |
| KITTI 00 & 1.00 m & 0.444 & \good{0.674} & 98.9\% & 389 & healthy \\ | |
| KITTI 00 & 2.00 m & 0.629 & \good{0.829} & 98.9\% & 389 & healthy \\ | |
| \midrule | |
| \multicolumn{7}{l}{\textbf{For reference --- hook-translation artifact (probe1, pre-fix):}} \\ | |
| 7Sc office/seq-02 & 0.05 m & 0.056 & \bad{0.008} & 4.4\% & --- & ARTIFACT \\ | |
| 7Sc office/seq-02 & 0.40 m & 0.117 & \bad{0.026} & 4.4\% & --- & ARTIFACT \\ | |
| \bottomrule | |
| \end{tabular} | |
| \end{center} | |
| \vspace{0.5em} | |
| {\footnotesize | |
| Retrieved sids were row-by-row \emph{identical} between probe1 and probe2; only the | |
| pointcloud-loader got fixed (float \texttt{frame\_id} parsing + timestamp-stem filename lookup). | |
| Early ``7Sc rejected / TTA 14$\times$ lift'' reading was driven entirely by the load miss. | |
| } | |
| \end{frame} | |
| % ------------------------------------------------------------------ | |
| \begin{frame}{Evidence C --- $\Delta t$ reach of the scorer (full distribution)} | |
| \framesubtitle{KITTI 00, V4-B'' ckpt-1, importance\_topk=20: 215 submaps, 4070 fetches} | |
| \scriptsize | |
| \textbf{$\Delta t$ distribution over all 4070 fetches:} | |
| \begin{center} | |
| \begin{tabular}{l r r r r r r r} | |
| \toprule | |
| \textbf{$\Delta t$ bin} & 1--3 & 4--9 & 10--19 & 20--49 & 50--99 & 100+ & \textbf{total $\ge 50$} \\ | |
| \midrule | |
| share of fetches & 2.8\% & 7.9\% & 9.9\% & 22.9\% & \good{27.4\%} & \good{29.0\%} & \good{56.4\%} \\ | |
| \bottomrule | |
| \end{tabular} | |
| \end{center} | |
| \vspace{0.6em} | |
| \textbf{Per-query loop closure rate:} | |
| \begin{center} | |
| \begin{tabular}{l r r r r} | |
| \toprule | |
| \textbf{stat} & median $\Delta t$ & mean $\Delta t$ & max $\Delta t$ & queries with $\exists\Delta t\ge 50$ \\ | |
| \midrule | |
| KITTI 00 (k=20) & 60 & 70.3 & 214 & \good{76.7\%} \\ | |
| KITTI 02 (k=10) & 48 & --- & --- & 72.7\% \\ | |
| TUM fr1/teddy (topall, bn=4) & 11 & --- & 34 & 0\% (seq too short) \\ | |
| \bottomrule | |
| \end{tabular} | |
| \end{center} | |
| \vspace{0.6em} | |
| \textbf{Anchor-dominance sanity check (KITTI 00):} fetched sid=0 in only 5.2\% of the 4070 fetches. | |
| Top-fetched $\Delta t$ values: 2, 3, 5--11 ($\sim$50 fetches each) --- recent history is fetched | |
| but the long tail is thick. Training-time finetune only supervises $\Delta t\le 3$; inference runs out to 214. | |
| \end{frame} | |
| % ================================================================== | |
| % PART 2 -- BASELINE SCOREBOARD | |
| % ================================================================== | |
| % ------------------------------------------------------------------ | |
| \begin{frame}{M3-SLAM benchmark --- full paper table} | |
| \framesubtitle{ATE RMSE (m) -- lower is better. Source: m3-slam/sec/experiments.tex.} | |
| \small | |
| \begin{center} | |
| \begin{tabular}{l r r r r r r} | |
| \toprule | |
| \textbf{Dataset} & MonoGS & DROID-SLAM & LEAP-VO & VGGT-SLAM & Pi3X-slam & \textbf{M3-SLAM} \\ | |
| \midrule | |
| ScanNet++ (20 scenes) & 0.623 & 0.346 & 1.081 & 0.182 & 0.137 & \good{0.065} \\ | |
| ScanNetV2 (10 scenes) & 0.217 & 0.106 & 0.852 & 0.073 & 0.141 & \good{0.051} \\ | |
| Waymo (9 seqs) & 7.535 & 16.61 & 13.17 & 1.295 & 2.410 & \good{0.773} \\ | |
| KITTI (8 seqs) & 8.212 & 18.07 & 10.99 & 2.521 & 2.938 & \good{0.890} \\ | |
| \bottomrule | |
| \end{tabular} | |
| \end{center} | |
| \vspace{0.8em} | |
| \textbf{V4-B'' ckpt-1 on the same axis:} | |
| \begin{center} | |
| \begin{tabular}{l r r r} | |
| \toprule | |
| \textbf{Dataset} & V4-B'' ckpt-1 & M3-SLAM & gap \\ | |
| \midrule | |
| TUM fr1 mean (9 scenes) & 0.224 & \textit{n/a} & same order as M3 indoor \\ | |
| KITTI 00 (frame-BA floor) & \bad{172.6} & 0.890 & \bad{194$\times$} \\ | |
| \bottomrule | |
| \end{tabular} | |
| \end{center} | |
| \textit{Same backbone family (Pi3X-class), same monocular streaming regime. Architectural | |
| differences only: matching head, pose-guided local search, Sim(3) global BA.} | |
| \end{frame} | |
| % ------------------------------------------------------------------ | |
| \begin{frame}{Oracle chain-TRS --- full per-backbone split} | |
| \framesubtitle{V4-oracle chain-TRS(6,4): oracle substitutes GT poses for 6-frame submap overlap alignment} | |
| \scriptsize | |
| \textbf{KITTI (FINAL, avg over 11 sequences, lower is better):} | |
| \begin{center} | |
| \begin{tabular}{l r r r} | |
| \toprule | |
| \textbf{Backbone} & top-5 ATE (m) & \#seqs won vs others & per-seq winners \\ | |
| \midrule | |
| V4-B' (paper + scannetpp + mvs\_synth + vK) & 72.5 & 0/11 & --- \\ | |
| V4-B'' (V3 ckpt-2 + same data mix) & \good{41.3} & 6/11 & 02, 03, 04, 07, 08, 10 \\ | |
| no-vK (paper + scannetpp + mvs\_synth) & \good{40.7} & 5/11 & 00, 01, 05, 06, 09 \\ | |
| \bottomrule | |
| \end{tabular} | |
| \end{center} | |
| \vspace{0.4em} | |
| \textbf{TUM fr1 (avg over 8 sequences):} | |
| \begin{center} | |
| \begin{tabular}{l r r r} | |
| \toprule | |
| \textbf{Backbone} & top-5 & top-all & gap to paper target (0.045) \\ | |
| \midrule | |
| V4-B' & 0.183 & 0.165 & 3.7$\times$ \\ | |
| V4-B'' & 0.160 & \good{0.102} & 2.3$\times$ \\ | |
| no-vK & 0.145 & 0.136 & 3.0$\times$ \\ | |
| \bottomrule | |
| \end{tabular} | |
| \end{center} | |
| \vspace{0.4em} | |
| With \emph{perfect} overlap alignment we still get 40.7 m KITTI / 0.102 m TUM. | |
| Stitch is not the bottleneck; the 15\% in-submap ATE is. | |
| \end{frame} | |
| % ------------------------------------------------------------------ | |
| \begin{frame}{Ablation --- top-K retrieval, full sweep} | |
| \framesubtitle{\S15.13, V4-oracle chain-TRS, bn\_every=6 fixed, same ckpts as scoreboard} | |
| \small | |
| \textbf{TUM fr1 (V4-B'' ckpt-1, all 8/8 seqs complete):} | |
| \begin{center} | |
| \begin{tabular}{l r r r r r} | |
| \toprule | |
| top-K & 5 & 10 & \textbf{20} & 50 & all \\ | |
| \midrule | |
| mean ATE (m) & 0.160 & 0.106 & \good{0.100} & 0.102 & 0.102 \\ | |
| \bottomrule | |
| \end{tabular} | |
| \end{center} | |
| \textit{Monotone down k=5$\to$10$\to$20, then plateau. TUM sweet spot = k=20.} | |
| \vspace{0.8em} | |
| \textbf{KITTI (no-vK ckpt-1, 11/11 at k=5/10/20, 6/11 at k=50):} | |
| \begin{center} | |
| \begin{tabular}{l r r r r} | |
| \toprule | |
| top-K & \textbf{5} & 10 & 20 & 50 \\ | |
| \midrule | |
| mean ATE (m) & \good{40.7} & 48.5 & 56.2 & 69.4 \\ | |
| \bottomrule | |
| \end{tabular} | |
| \end{center} | |
| \textit{Monotone \emph{up} --- adding history \emph{hurts}. KITTI sweet spot = k=5.} | |
| \vspace{0.6em} | |
| \textbf{Domain-specific top-K is required.} KITTI forward motion has no persistent | |
| loop back until kilometre scale, so large top-K just injects irrelevant tokens. | |
| TUM sees the same scene from many angles; more retrieval = more coverage. | |
| \end{frame} | |
| % ------------------------------------------------------------------ | |
| \begin{frame}{Ablation --- frames per submap (\texttt{bn\_every})} | |
| \framesubtitle{\S15.14, V4-oracle chain-TRS, top-K=5 fixed} | |
| \small | |
| \begin{center} | |
| \begin{tabular}{l r r r l} | |
| \toprule | |
| \textbf{dataset} & bn=4 & \textbf{bn=6} & bn=10 & coverage \\ | |
| \midrule | |
| TUM fr1 (V4-B'') & 0.174 & \good{0.160} & 0.170 & 8/8 at every bn \\ | |
| KITTI (no-vK) & 75.7 & \good{40.7} & 77.5 & bn4=8/11, bn6=11/11, bn10=9/11 \\ | |
| \bottomrule | |
| \end{tabular} | |
| \end{center} | |
| \vspace{0.8em} | |
| \textbf{Both domains prefer the default bn=6.} | |
| bn=4 fragments context before the scorer can use it; bn=10 accumulates more depth bias | |
| inside the window. Same pattern as top-K: neither knob closes the gap to M3. | |
| \vspace{0.6em} | |
| \textbf{Canonical config} (carried through all subsequent experiments): V4-B'' backbone | |
| + bn\_every=6 + top-K=20 on TUM / top-K=5 on KITTI. | |
| \end{frame} | |
| % ------------------------------------------------------------------ | |
| \begin{frame}{Cross-dataset scoreboard --- indoor (V4-B'' ckpt-1)} | |
| \framesubtitle{Evaluated 2026-04-20} | |
| \small | |
| \begin{center} | |
| \begin{tabular}{l r r r r r} | |
| \toprule | |
| \textbf{benchmark} & \textbf{\#seqs} & \textbf{mean ATE (m)} & \textbf{AUC@3$^\circ$} & \textbf{AUC@30$^\circ$} & \textbf{status} \\ | |
| \midrule | |
| 7-Scenes & 17 & 0.193 & $\sim$0.05 & 0.642 & ATE at paper parity, rotation weak \\ | |
| NRGBD & 8 & 0.252 & $\sim$0.01 & 0.780 & ATE at paper parity (OpenGL fix required) \\ | |
| TUM fr1 & 9 & 0.224 & \bad{0.000} & 0.368 & AUC@3$^\circ$=0 systemic \\ | |
| \bottomrule | |
| \end{tabular} | |
| \end{center} | |
| \vspace{0.8em} | |
| \textbf{TUM fr1 per-scene ATE (V4-B'' ckpt-1):} | |
| \begin{center} | |
| \begin{tabular}{l r r r r r r r r r r} | |
| \toprule | |
| scene & 360 & desk & desk2 & floor & plant & room & rpy & teddy & xyz & \textbf{mean} \\ | |
| \midrule | |
| ATE (m) & 0.17 & 0.11 & \good{0.017} & 0.13 & --- & \bad{0.73} & 0.24 & 0.23 & 0.10 & 0.224 \\ | |
| \bottomrule | |
| \end{tabular} | |
| \end{center} | |
| \vspace{0.3em} | |
| {\footnotesize Rotation AUC@3$^\circ$=0 even on scenes where translation is cm-scale $\Rightarrow$ | |
| indoor rotation is under-fit, not a geometric failure.} | |
| \end{frame} | |
| % ------------------------------------------------------------------ | |
| \begin{frame}{Cross-dataset scoreboard --- outdoor (V4-B'' ckpt-1)} | |
| \framesubtitle{Evaluated 2026-04-20} | |
| \small | |
| \begin{center} | |
| \begin{tabular}{l r r r r} | |
| \toprule | |
| \textbf{benchmark} & \textbf{mean ATE (m)} & \textbf{paper target (m)} & \textbf{M3 (m)} & \textbf{gap vs paper / M3} \\ | |
| \midrule | |
| ETH3D (median over 13 scenes) & 0.63 & --- & --- & AUC@3$^\circ$=0.01 vs paper 0.67 ($30\times$) \\ | |
| Oxford Spires / keble & 25.25 & 6--12 & --- & 2--4$\times$ \\ | |
| Oxford Spires / christ & 36.15 & 6--12 & --- & 3--6$\times$ \\ | |
| Oxford Spires mean & 30.7 & 6--12 & --- & 3--5$\times$ \\ | |
| KITTI (mean) & \bad{148} & 14.9 & 0.89 & \bad{10$\times$ paper, 166$\times$ M3} \\ | |
| \bottomrule | |
| \end{tabular} | |
| \end{center} | |
| \vspace{0.8em} | |
| \textbf{Oxford Spires alignment coverage:} n\_matched $\approx$ 100 / 3840 patches | |
| $\Rightarrow$ \bad{2.6\% alignment coverage}. Most frames never align under the current | |
| matcher. The 25--36 m errors are driven by unaligned frames, not by local geometry | |
| precision. | |
| \vspace{0.3em} | |
| {\footnotesize ETH3D F1$>$0 on 4/13 scenes only. Bad F1 is a coverage issue, not a depth one.} | |
| \end{frame} | |
| % ================================================================== | |
| % PART 3 -- METHODS (MATCHING + DESCRIPTORS + SIM(3) OPTIMIZATION) | |
| % ================================================================== | |
| % ------------------------------------------------------------------ | |
| \begin{frame}{Matching system pipeline} | |
| \vspace{-0.3em} | |
| \begin{center} | |
| \resizebox{\linewidth}{!}{\includegraphics{matching_pipeline_figure.pdf}} | |
| \end{center} | |
| \vspace{-0.7em} | |
| {\footnotesize | |
| \textbf{Four stages, all inference-time, no training:} | |
| (1) SLAM-Former backbone returns per-patch tokens $F\in\mathbb{R}^{P\times d}$ and local 3D $X\in\mathbb{R}^{P\times 3}$; | |
| (2) rank patches by confidence $c_p$, keep top-$K$, L2-normalize the token as the descriptor; | |
| (3) mutual-NN cosine match + Lowe ratio + $3\!\times\!3$ quadratic sub-pixel refinement; | |
| (4) Sim(3) frame BA (scipy LSMR, 7 DoF per frame, velocity-smoothness scale prior $\lambda_{\rm scale}=0.1$). | |
| No learned matching head; reuses backbone tokens as features.} | |
| \end{frame} | |
| % ------------------------------------------------------------------ | |
| \begin{frame}{Matching --- how descriptors are extracted} | |
| \small | |
| \textbf{Ours (\texttt{frame\_matcher.py:select\_frame\_keypoints}):} | |
| \begin{itemize}\small | |
| \item Backbone returns per-patch tokens $\mathbf{t}_{\rm patch}\in\mathbb{R}^{H_p\times W_p\times D}$ | |
| ($H_p=W_p=37$ at target\_size=518, patch=14$\times$14 px, $D$=token dim). | |
| \item Per-patch confidence: mean of pixel-wise conf over the 14$\times$14 cell. | |
| \item Patch-center 3D point $\mathbf{X}_i^i\in\mathbb{R}^3$ from the predicted local pointcloud; | |
| patch-center normalized camera UV $(u,v) = (x/z, y/z)$ of that 3D point. | |
| \item Top-$K$ patches by confidence, with $z\in(z_{\min},z_{\max})$ finite. | |
| \item Descriptor = patch token, L2-normalized, stored as fp16. | |
| \end{itemize} | |
| \vspace{0.6em} | |
| \textbf{M3-SLAM (from \texttt{refer/m3-slam/sec/method.tex}):} | |
| \begin{itemize}\small | |
| \item Separate \textbf{matching head} on top of Pi3X: $\mathrm{DPT}_{\rm desc}+$2-layer MLP with GELU. | |
| \item Outputs dense descriptors $\mathbf{D}\in\mathbb{R}^{H\times W\times d}$ | |
| ($d=24$ by default, \emph{per pixel}, not per patch) and matching confidence | |
| $\mathbf{Q}\in\mathbb{R}^{H\times W\times 1}$. | |
| \item Trained with InfoNCE symmetric loss against GT pixel correspondences; | |
| encoder / decoder / point head frozen. | |
| \end{itemize} | |
| \vspace{0.4em} | |
| {\footnotesize \textbf{The structural gap:} our ``descriptor'' is a token over a | |
| 14$\times$14 patch, giving $37^2 \approx 1369$ candidates per frame; M3's is | |
| \emph{per pixel}, giving $518^2 \approx 268{,}000$ candidates per frame. A single | |
| token integrates over $\sim$200 pixels of depth variation.} | |
| \end{frame} | |
| % ------------------------------------------------------------------ | |
| \begin{frame}{Matching --- how pairs are formed} | |
| \small | |
| \textbf{Ours (\texttt{frame\_matcher.py:match\_frames}):} | |
| \begin{enumerate}\small | |
| \item Mutual-nearest-neighbour on cosine similarity between L2-normalized token descriptors. | |
| \item Lowe's ratio test on cosine: keep pair if $\cos_{\rm top2}\le r\cdot\cos_{\rm top1}$, | |
| typical $r\in[0.8,0.9]$. | |
| \item Sub-pixel refinement via quadratic fit on the 3$\times$3 cos-sim neighborhood of | |
| the peak, offset clipped to $[-0.5,+0.5]$ patch units. | |
| \item Output: \texttt{MatchEdge} with indices $i\leftrightarrow j$, cos-sim, and refined $(u_j,v_j)$. | |
| \end{enumerate} | |
| \vspace{0.6em} | |
| \textbf{M3-SLAM (Sec. 3.1.3 --- ``Dense matching for SLAM''):} | |
| \begin{enumerate}\small | |
| \item Pose-guided initialization: project $\mathbf{X}_i$ into frame $j$ using current $(T_i, T_j)$ | |
| $$\mathbf{X}_i^j = T_j^{-1}\,T_i\,\mathbf{X}_i.$$ | |
| \item For each pixel $\mathbf{p}_i$, search a local window of radius $r$ around the projected location | |
| in frame $j$, selecting $\mathbf{p}^*_j = \arg\max_{\mathbf{p}\in\mathcal{N}_r} \big\langle \mathbf{D}_i(\mathbf{p}_i), \mathbf{D}_j(\mathbf{p})\big\rangle$ | |
| (cosine similarity). | |
| \item Dynamic masking: a motion map $\mathbf{M}_i\in[0,1]^{H\times W}$ down-weights pixels with | |
| low warped-vs-observed descriptor consistency. | |
| \end{enumerate} | |
| \vspace{0.4em} | |
| {\footnotesize \textbf{Key difference:} M3's matching is $O(N\cdot r^2)$ local geometric search; | |
| ours is $O(K^2)$ global descriptor matching with no pose prior. M3's search radius $r$ can be | |
| as small as a few pixels \emph{because the pose initialization is good}.} | |
| \end{frame} | |
| % ------------------------------------------------------------------ | |
| \begin{frame}{Matcher precision --- how it actually performs} | |
| \framesubtitle{GT-supervised M3-style eval (\texttt{eval\_matcher\_gt.py}), $\tau=0.6$, err threshold 0.10 m} | |
| \small | |
| \textbf{TUM fr1/xyz (48 frames, patch-level):} | |
| \begin{center} | |
| \begin{tabular}{l r r r r} | |
| \toprule | |
| err threshold & 5 cm & 10 cm & 20 cm & 50 cm \\ | |
| \midrule | |
| precision & \bad{$\sim$8\%} & --- & 55\% & 88\% \\ | |
| \bottomrule | |
| \end{tabular} | |
| \end{center} | |
| median patch-pair 3D error $\approx$ \textbf{19 cm} (patches = 14$\times$14 px $\approx$ 17 original pixels). | |
| \vspace{0.8em} | |
| \textbf{ScanNet++ matcher eval (\#22058916, STRIDE=5 iPhone frames, 3 scenes):} | |
| \begin{center} | |
| \begin{tabular}{l r r r r r} | |
| \toprule | |
| \textbf{scene} & \textbf{\#pairs} & \textbf{\#matches} & \textbf{\#correct} & \textbf{precision} & \textbf{median err (m)} \\ | |
| \midrule | |
| 036bce3393 & 51 & 1292 & 75 & \bad{5.8\%} & 1.99 \\ | |
| 079a326597 & 41 & 1319 & 172 & \bad{13.0\%} & 0.68 \\ | |
| 07ff1c45bb & 53 & 1286 & 59 & \bad{4.6\%} & 1.31 \\ | |
| \midrule | |
| mean & 48 & 1299 & 102 & \bad{7.8\%} & 1.33 \\ | |
| \bottomrule | |
| \end{tabular} | |
| \end{center} | |
| \vspace{0.3em} | |
| {\footnotesize iPhone STRIDE=5 keyframe gaps of 0.3--0.8 s produce wide view changes. | |
| Precision 5--13\% vs TUM 30\% baseline is the natural consequence of patch-level | |
| descriptors under large disparity.} | |
| \end{frame} | |
| % ------------------------------------------------------------------ | |
| \begin{frame}{Sim(3) optimization --- parameterization} | |
| \small | |
| \textbf{Pose as a similarity transform} $T\in \mathbf{Sim}(3)$: | |
| $$ | |
| T = \begin{bmatrix} s\mathbf{R} & \mathbf{t} \\ \mathbf{0}^\top & 1 \end{bmatrix}, | |
| \qquad \mathbf{R}\in\mathbf{SO}(3),\ \mathbf{t}\in\mathbb{R}^3,\ s\in\mathbb{R}_{>0} | |
| $$ | |
| \vspace{0.4em} | |
| \textbf{Lie algebra representation} $\boldsymbol{\tau}\in\mathfrak{sim}(3)\cong\mathbb{R}^7$: | |
| \begin{itemize} | |
| \item 3 for axis-angle rotation $\boldsymbol{\omega}\in\mathbb{R}^3$, 3 for translation | |
| $\boldsymbol{\rho}\in\mathbb{R}^3$, 1 for log-scale $\lambda\in\mathbb{R}$. | |
| \item Update via left-multiplication on the manifold: | |
| $T \leftarrow \exp(\boldsymbol{\tau})\circ T$. | |
| \end{itemize} | |
| \vspace{0.8em} | |
| \textbf{Ours (\texttt{frame\_ba.py:FrameBA}, \texttt{use\_sim3=True}):} | |
| \begin{itemize} | |
| \item State vector $\boldsymbol{\theta}\in\mathbb{R}^{7(N-1)}$: 6 DoF SE(3) + 1 DoF log-scale per frame, | |
| frame 0 pinned at identity (gauge). | |
| \item Local state reconstruction: | |
| $T_i = \big(\mathbf{R}_i \mid \mathbf{t}_i\big),\ \log s_i = \lambda_i$ | |
| $\Rightarrow$ world point | |
| $\mathbf{X}_i^w = e^{\lambda_i}\mathbf{R}_i\,\mathbf{X}_i^i + \mathbf{t}_i.$ | |
| \end{itemize} | |
| \vspace{0.5em} | |
| {\footnotesize M3 also uses $\mathfrak{sim}(3)$ on the Lie algebra with the left-plus update operator. | |
| Difference from ours: M3 has $\mathbf{Sim}(3)$ \emph{tracking} against a single reference keyframe plus global BA; | |
| ours is a single monolithic BA over all matched frames.} | |
| \end{frame} | |
| % ------------------------------------------------------------------ | |
| \begin{frame}{Sim(3) optimization --- residuals and solver} | |
| \small | |
| \textbf{Ours --- per match} $(m \text{ in } i) \leftrightarrow (n \text{ in } j)$: | |
| \begin{enumerate} | |
| \item Unproject the i-side patch to world: $\mathbf{X}^w_m = e^{\lambda_i}\mathbf{R}_i\mathbf{X}_i^i(m) + \mathbf{t}_i$. | |
| \item Reproject into j-camera: $\mathbf{X}_j^{\rm cam}(m) = \mathbf{R}_j^\top(\mathbf{X}^w_m - \mathbf{t}_j)\cdot e^{-\lambda_j}$. | |
| \item Project to normalized pixel: $\hat{\mathbf{u}}_j(m) = \mathbf{X}_j^{\rm cam}(m)_{[:2]}/\mathbf{X}_j^{\rm cam}(m)_{[2]}$. | |
| \item Residual: $r_m = \cos(\theta_{mn})\cdot(\mathbf{u}_j^{\rm obs}(n) - \hat{\mathbf{u}}_j(m))\in\mathbb{R}^2$. | |
| \end{enumerate} | |
| Plus a \textbf{velocity smoothness / scale prior}: | |
| $$r_{\rm scale}(i) = w_{\rm prior}\cdot(\lambda_{i+1}-\lambda_i),\qquad w_{\rm prior}=0.1.$$ | |
| \vspace{0.4em} | |
| \textbf{Solver:} \texttt{scipy.optimize.least\_squares} with LSMR sparse Jacobian, Huber loss | |
| (scale=$\sigma_{\rm reproj}$), max iterations 500, tolerance $10^{-8}$. | |
| \vspace{0.6em} | |
| \textbf{M3 tracking residual (\texttt{method.tex:215-228}):} | |
| $$ | |
| E_{\rm track} = \sum_{(m,n)\in\mathcal{M}(I_f,I_k)}\mathbf{M}_k\, | |
| \rho\!\left(\left\|\frac{\mathbf{p}_{k,n} - \phi(T_{kf}\,\mathbf{X}^f_{f,m})}{w(q_{m,n})}\right\|^2\right) | |
| $$ | |
| \textbf{M3 global BA:} | |
| $$ | |
| E_g = \sum_{(i,j)\in\mathcal{E}}\sum_{(m,n)\in\mathcal{M}(I_j,I_i)}\mathbf{M}_i\, | |
| \rho\!\left(\left\|\frac{\mathbf{p}_{i,n} - \phi(T_{ij}\,\tilde{\mathbf{X}}^j_{j,m})}{w(q_{m,n})}\right\|^2\right) | |
| $$ | |
| Huber robust loss $\rho$, per-match confidence weight $w(q_{m,n})\!=\!\sqrt{\mathbf{Q}^f_f(m)\mathbf{Q}^f_k(n)}$. | |
| \end{frame} | |
| % ================================================================== | |
| % PART 4 -- RESULTS OF THE PIVOT + PATH FORWARD | |
| % ================================================================== | |
| % ------------------------------------------------------------------ | |
| \begin{frame}{Sim(3) frame-BA --- TUM fr1 per-scene} | |
| \framesubtitle{Job 22058142, 9 scenes, SE(3) baseline vs Sim(3) with \texttt{scale\_prior\_weight=0.1}} | |
| \scriptsize | |
| \begin{center} | |
| \begin{tabular}{l r r r l} | |
| \toprule | |
| \textbf{scene} & \textbf{SE(3) ATE} & \textbf{Sim(3) ATE} & \textbf{$\Delta$ (Sim-SE)} & \textbf{winner} \\ | |
| \midrule | |
| 360 & 0.164 & 0.159 & $-0.005$ & Sim(3) \\ | |
| desk & 0.699 & 0.615 & $-0.085$ & Sim(3) \\ | |
| desk2 & \good{0.011} & 0.012 & $+0.001$ & tie \\ | |
| floor & 0.183 & 0.178 & $-0.005$ & Sim(3) \\ | |
| plant & --- & --- & --- & {\color{mute}Umeyama degenerate (TUM data issue)} \\ | |
| room & 0.283 & 0.298 & $+0.016$ & SE(3) \\ | |
| rpy & 0.044 & 0.047 & $+0.003$ & tie \\ | |
| teddy & 0.576 & 0.685 & $+0.109$ & SE(3) \\ | |
| xyz & 0.097 & 0.105 & $+0.008$ & tie \\ | |
| \midrule | |
| \textbf{mean} & \textbf{0.257} & \textbf{0.262} & \bad{\textbf{+0.005}} & SE(3) by 0.5 cm \\ | |
| \bottomrule | |
| \end{tabular} | |
| \end{center} | |
| \vspace{0.5em} | |
| Sim(3) wins on 360 / desk / floor (scenes with larger rigid drift), loses on teddy (108$\times$108 hand-scale), | |
| and is a wash elsewhere. TUM indoor has almost no scale drift to absorb; the extra DoF over-flexes. | |
| \texttt{scale\_prior\_weight=0.1} is too permissive for TUM (per-frame scale wanders despite the prior). | |
| \end{frame} | |
| % ------------------------------------------------------------------ | |
| \begin{frame}{Sim(3) frame-BA --- KITTI 00 full variant table} | |
| \framesubtitle{Job 22058145, seq 00 only, 500 iterations, 109 min wall time} | |
| \small | |
| \textbf{Sim(3)-aligned ATE (m) for all output variants of the same dump:} | |
| \begin{center} | |
| \begin{tabular}{l r l} | |
| \toprule | |
| \textbf{variant} & \textbf{ATE (m)} & \textbf{description} \\ | |
| \midrule | |
| stitch & 173.05 & online Sim(3) stitch, per-submap (default pipeline) \\ | |
| raw & 192.94 & frontend poses, no stitch, no BA \\ | |
| offline & 173.18 & final-pass re-stitch after run ends \\ | |
| posegraph & \good{172.62} & SubmapPoseGraph (gtsam Similarity3 factor graph) \\ | |
| frame\_ba & 173.19 & per-frame Sim(3) BA over the matcher graph \\ | |
| \midrule | |
| SE(3) frame\_ba (comparison) & 172.6 & Same runs w/o the 7th DoF \\ | |
| \bottomrule | |
| \end{tabular} | |
| \end{center} | |
| \vspace{0.6em} | |
| \textbf{Per-frame scale DoF stats (Sim(3) run):} | |
| \begin{center} | |
| \begin{tabular}{l r r r r} | |
| \toprule | |
| statistic & median & min & max & stdev \\ | |
| \midrule | |
| $e^\lambda$ & \hl{0.9999} & 0.83 & 1.02 & 0.013 \\ | |
| \bottomrule | |
| \end{tabular} | |
| \end{center} | |
| Cost function dropped \textbf{20$\times$} (8.0e5 $\to$ 4.0e4); ATE moved by 0.5 m. The | |
| extra per-frame scale DoF is barely used because backbone depth is scale-consistent | |
| but \emph{absolute-bias} (Umeyama $\Rightarrow$ 22\% mean depth error over 784-m Karlsruhe loop). | |
| \end{frame} | |
| % ------------------------------------------------------------------ | |
| \begin{frame}{Root-cause summary table --- what each hypothesis predicts vs what we see} | |
| \small | |
| \begin{center} | |
| \begin{tabular}{p{4.0cm} l l l} | |
| \toprule | |
| \textbf{hypothesis} & \textbf{prediction} & \textbf{observation} & \textbf{verdict} \\ | |
| \midrule | |
| Per-frame scale drift & Sim(3) $\gg$ SE(3) on KITTI & Sim(3) = SE(3) $\pm$ 0.6 m & \bad{FALSIFIED} \\ | |
| Retrieval collapse & small $\Delta t$, small H$\leftrightarrow$H & $\Delta t$ med=60, H$\leftrightarrow$H=0.50--0.83 & \bad{FALSIFIED} \\ | |
| Bad overlap stitch & oracle $\gg$ raw-stitch & oracle 40.7 vs raw 192.9 m (KITTI) & (stitch contributes) \\ | |
| Depth absolute bias & Umeyama scale $\ne 1$, ATE persists & Umeyama 22\% depth error, ATE 172 m & \good{SUPPORTED} \\ | |
| Patch-level matching cap & matcher precision capped & TUM@5cm=8\%, ScnNet++=7.8\% & \good{SUPPORTED} \\ | |
| \bottomrule | |
| \end{tabular} | |
| \end{center} | |
| \vspace{0.8em} | |
| Two root causes supported: | |
| (1) absolute-scale depth bias of the monocular backbone on outdoor imagery; | |
| (2) 14$\times$14 patch resolution of our matcher cannot resolve 5--20 cm correspondence | |
| for BA to act on. Submap-level stitch knobs do not move the ATE regardless of configuration. | |
| \end{frame} | |
| % ------------------------------------------------------------------ | |
| \begin{frame}{Path forward --- prioritized} | |
| \small | |
| \begin{center} | |
| \begin{tabular}{r p{7.3cm} p{3.3cm}} | |
| \toprule | |
| \# & \textbf{action} & \textbf{expected lever} \\ | |
| \midrule | |
| 1 & Train a pixel-level matching head (DPT+MLP) on V4 backbone, InfoNCE loss vs vKITTI2/ScanNet++ GT correspondences. & Closes TUM@5cm gap (8$\to$$>$50\%). \\ | |
| 2 & Pose-guided local search at inference: project $\mathbf{X}_i$ into $I_j$, refine inside radius $r$. & Lifts matcher precision under wide disparity. \\ | |
| 3 & Frame-level SE(3) graph BA with new matcher as 2D-3D residuals. & Sim(3) only buys 0.6 m; lever is matching. \\ | |
| 4 & Rotation-underfit fix on indoor: loss rebalancing on pose head. & AUC@3$^\circ$ from 0 to paper range. \\ | |
| 5 & Add real KITTI imagery to V4 training mix. & Closes 10$\times$ outdoor gap (vKITTI2-only). \\ | |
| 6 & De-prioritize submap-level PGO variants. & They sit on the 15\% substrate regardless. \\ | |
| \bottomrule | |
| \end{tabular} | |
| \end{center} | |
| \vspace{0.8em} | |
| \textbf{What we keep as-is:} V4-B'' ckpt-1 backbone, bn\_every=6, domain-specific top-K | |
| (20 TUM / 5 KITTI), Sim(3) stitch + freeze\_history\_writeback, scorer + overlap retrieval, | |
| submap abstraction itself. | |
| \end{frame} | |
| \end{document} | |