You need to agree to share your contact information to access this model
This repository is publicly accessible, but you have to accept the conditions to access its files and content.
These weights are trained on NVIDIA PhysicalAI-AV data (TanitAD research program). Access is granted per request for research/evaluation use only; you agree not to redistribute.
Log in or Sign Up to review the conditions and access this model content.
TanitAD β REF-C-base / medium (anchored diffusion, step 29,999 / 30,000 FINAL)
Reference arm C of the TanitAD program: a DiffusionDrive-style anchored-diffusion direct planner β the budget-matched non-world-model control for the hierarchical 4-brain flagship. This is the base (medium) rung, 104.2 M params, of the three-size REF-C ladder (small 54.7 M Β· base 104.2 M Β· XL 251.9 M).
This is the rung the program actually recommends. It ties the 2.42Γ-larger XL on every shipping metric while running 1.32β2.02Γ faster.
Registry key: refc-diffusion-base-v21-30k Β· TanitEval arm key: refc-base-30k
Source of truth for every number below: Project Steering/MODEL_REGISTRY.md Β§4.3 and the raw eval
JSON taniteval/results/refc-base-30k.json. Evidence class: MEASURED unless marked otherwise.
Architecture
Read from this run's own config.json (shipped in this repo).
- Encoder β torchvision-free ResNet-M,
in_channels 9(3 RGB frames at 100 ms spacing, channel-stacked),image_size 256,base_width 88, blocks(3, 6, 16, 6). - Anchor vocabulary β 128 trajectory anchors by furthest-point sampling from a 4,096-window
pool,
seed 0. Verified a strict, bit-exact prefix of XL's 256 (refc_anchors_base128.pt, same script / source / pool cap / seed;max|A β B[:128]| = 0at load). - Decoder β
d 384, 8 heads, 4 layers,ff_mult 4,aux_hidden 384, 2 truncated-denoise steps,noise_std 0.1. - Measurement encoder β
hidden 128,d_out 128,ego_dropout 0.5. - LAW latent-world-model auxiliary β
hidden 2048. - Strategic-context graft β
hidden 512,d_ctx 64;hierarchy true,graft_maneuver true. - Imagination graft (H15) β OFF by preset design (XL-only). Contributes 0 params here.
- Horizons
(5, 10, 15, 20)steps @ 10 Hz = 0.5 / 1 / 1.5 / 2 s Β·path_dists (2, 5, 10, 20)m Β·speed_bins 4,speed_max 30.0Β·graft_target_latent false,grounded_selector false,refc1 false.
Parameters β measured at instantiation (config.json param_breakdown)
| module | params |
|---|---|
| encoder | 90,458,632 |
| decoder | 8,634,505 |
| law | 2,902,720 |
| strategic | 1,903,680 |
| aux | 274,760 |
| measurement | 17,280 |
| imagination | 0 (graft off) |
| total | 104,191,577 |
n_params_trainable = 104,191,577. The code docstring's "~110 M" was 5.6 % high β the measured
number above is the one to quote. Anchors are buffers, not parameters.
Training
| item | value |
|---|---|
| Corpus | NVIDIA PhysicalAI-AV front-wide, 2,376 episodes / 406,099 windows (train split) β booked in this run's config.json data block |
| Strict-parity build key | physicalai-train-e438721ae894 |
| Corrupt-clip skip-hash | f09e44db (24 corrupt front-wide clips excluded) |
| Steps | 29,999 / 30,000 (metrics.json final.step = 29999, steps = 30000) |
| Optimizer | Adam (DiffusionDrive/TCP convention, not AdamW), lr 1e-4, warmup 2000, cosine |
| Loss weights | traj 1.0 Β· cls 1.0 Β· law 0.5 Β· route 0.1 Β· man 0.1 Β· speed_cls 0.2 |
| Batch / workers | 20 / 6 Β· seed 0 Β· device cuda |
| Route labels | v2.1 (refb_labels.route_from_future_v21, use_net_dyaw=False, ROUTE_UNKNOWN=3 masked out of the 0.1-weight CE, never clamped) |
| Hardware | tanitad-pod3; finished 2026-07-21 04:44 UTC. Evaluated on tanitad-eval 2026-07-21 05:18β05:19 UTC under the refc-base-eval GPU lock |
Exact command:
cd /workspace/TanitAD/stack && PYTHONPATH=/workspace/TanitAD/stack \
PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True nohup python3 scripts/refc_train.py \
--data-root /workspace/pai_epcache \
--out /workspace/experiments/refc-diffusion-base-v21-30k \
--steps 30000 --mode diffusion --config base \
--anchors /workspace/experiments/refc_anchors_base128.pt \
--labels v21 \
--batch 20 --workers 6
Parity with XL: same corpus and parity key, 30 k steps, same optimizer/schedule, same loss weights,
--mode diffusion, --batch 20 --workers 6.
Deliberate differences vs XL: --config base (2.42Γ smaller) Β· 128 anchors (prefix of XL's 256) Β·
H15 imagination OFF Β· route labels v2.1 (see caveat 4).
v2.1 label coverage (4,000-window sample, recorded in config.json): left 0.121 Β· straight 0.5645 Β·
right 0.115 Β· UNKNOWN 0.1995 (masked out) β 80.05 % judgeable, against v1's straight-by-default
target. Unknown reasons: gray_zone 0.1103 Β· no_arc 0.0892 Β· road_following 0.5645 Β· tight_transient 0.236.
Code: stack/scripts/refc_train.py gained --labels {v1,v21} (default v1 = XL-reproducible),
RouteV21Dataset, a fail-loud masked route CE, and 5 k/15 k/20 k/30 k milestone archiving (a gate series
XL lacks). 15/15 tests/test_refc.py pass.
Final training metrics (step 29,999, metrics.json): loss 1.37374 Β· traj 0.18426 Β· cls 1.09855 Β·
law 0.01716 Β· route 0.28667 Β· man 0.5368 Β· anchor_acc 0.55 Β· man_acc 0.90 Β· route_acc 0.85.
Evaluation
TanitEval taniteval.refc_eval, open-loop, on the clean held-out split
physicalai-val-0c5f7dac3b11 β 40 episodes β 881 windows, episode-disjoint from train.
Protocol identical to XL: window 8, stride 8, K = 20 steps @ 10 Hz, waypoints [5, 10, 15, 20],
metric-BEV ego frame, nav = follow, anchored-diffusion decode with 2 truncated-denoise steps
over 128 anchors, argmax-confidence select. ckpt_step in the result JSON = 29999.
Parity with XL's eval proven three ways: the same 881 episode ids, bit-identical ground truth, and a bit-identical CV baseline in every stratum.
Headline β decision-grade (n = 881 windows / 40 episodes)
| metric | REF-C-base (104.2 M) | REF-C-XL (251.9 M) | paired Ξ (base β XL) |
|---|---|---|---|
| ADE@2s (full-set) | 0.4728 Β· CI95 [0.3835, 0.5699] | 0.4714 Β· [0.3896, 0.5556] | +0.0013 [β0.0281, +0.0316] β NOT separated |
| FDE@2s (full-set) | 1.0031 Β· [0.8148, 1.2087] | 1.0061 Β· [0.8301, 1.1875] | β0.0030 [β0.0619, +0.0584] β NOT separated |
| miss@2m (full-set) | 0.1419 Β· [0.0874, 0.2000] | 0.1419 Β· [0.0943, 0.1918] | +0.0000 [β0.0261, +0.0272] β NOT separated |
| TMS-openloop (full-set) | 0.1957 | 0.2135 | β |
| ADE@0.5 / 1 / 1.5 s | 0.0708 / 0.1620 / 0.2960 | 0.0681 / 0.1592 / 0.2932 | β |
Estimator: taniteval/ci.py episode-cluster bootstrap, B = 2000 over the 40 val episodes; the
paired form for the deltas. Per-window ADE correlation 0.789, so the pairing is doing real work β
the non-separation is not a weak test.
Trivial floor on the same split: constant-velocity ADE@2s 0.8377 (full-set) / 0.8248 (heldout). REF-C-base beats CV overall and in every stratum.
Verdict: REF-C-base and REF-C-XL are statistically indistinguishable on everything that ships. All three paired intervals straddle zero and the point deltas are β€ 0.003 m β a 2.42Γ parameter cut and a 2.20Γ encoder cut (90,458,632 vs 199,496,532) cost nothing measurable on this corpus.
Estimator caveat β never quote an interval without its estimator. A legacy row for this arm reads 0.4523 Β± 0.0497. That
Β±isoverlapping_holdout_se(historically mislabelled "8-split episode-disjoint jackknife"); it is neither a jackknife nor a valid SE, measured 1.28β2.06Γ too narrow across 10 arms. It is retained only for continuity with older publications. Use the episode-cluster bootstrap row. Source:Project Steering/CI_RECOMPUTE_2026-07-20.json.
Strata (full-set, step 29,999)
| stratum | base ADE@2s | XL ADE@2s | CV | n |
|---|---|---|---|---|
| speed high | 0.3510 | 0.3243 | 0.6468 | 294 |
| speed med | 0.4483 | 0.4989 | 0.9345 | 293 |
| speed low | 0.6189 | 0.5912 | 0.9322 | 294 |
| curv straight | 0.3866 | 0.3865 | 0.4393 | 634 |
| curv gentle | 0.6778 | 0.6751 | 1.3566 | 125 |
| curv sharp | 0.7105 | 0.7040 | 2.3764 | 122 |
The two arms trade strata (base wins med by 0.051; XL wins high by 0.027 and low by 0.028) β no stratum-level ordering survives as a scale story.
Fan quality β the read the decision actually needs
Raw: taniteval/results/scaleab_refc-base-30k_vs_refc-xl-30k.json.
| base (128 anchors) | XL (256) | XL restricted to its first 128 | |
|---|---|---|---|
| oracle-in-fan | 0.1914 [0.1654, 0.2184] | 0.1640 [0.1414, 0.1902] | 0.2624 [0.2262, 0.3011] |
| sel_gap (selected β oracle) | 0.2813 | 0.3075 | 0.2091 |
frac_sel_2x_worse |
0.4109 | 0.4540 | 0.3190 |
Paired: base β XL(256) +0.0275 [+0.0142, +0.0405] SEPARATED (XL better) Β· base β XL(128) β0.0710 [β0.0965, β0.0502] SEPARATED (base better).
Oracle-in-fan over the first K anchors (base β€ XL at every matched K; XL's entire oracle advantage arrives with anchors 129β256):
| K | 4 | 8 | 16 | 32 | 64 | 128 | 256 |
|---|---|---|---|---|---|---|---|
| base | 3.193 | 1.686 | 0.813 | 0.527 | 0.283 | 0.191 | β |
| XL | 3.535 | 2.274 | 1.226 | 0.806 | 0.437 | 0.262 | 0.164 |
The fan lever is anchor-vocabulary WIDTH, not encoder scale. Anchors are buffers (~0.048 M total), and the decoder is only ~1.7 ms of base's 21.8 ms tick (encoder 90.7 %) β so widening the vocabulary is nearly free while widening the encoder demonstrably bought nothing here.
Efficiency β batch 1, one A40, identical precision flags
| base | XL | ratio | |
|---|---|---|---|
| plan tick p50 fp32 / tf32 / amp16 | 21.78 / 15.81 / 15.88 ms | 44.06 / 27.78 / 21.00 ms | 1.32β2.02Γ faster |
| p99 fp32 | 22.33 ms | 44.44 ms | meets 10 Hz in all 3 precisions (both arms) |
| GFLOPs / peak alloc | 292.5 / 556.7 MB | 702.2 / 1178.4 MB | 0.42Γ / 0.47Γ |
| encoder share of the tick | 90.7 % | 88.7 % | β |
Honest caveats β all recorded in the registry; read them before quoting this model
Selection flaw β REF-C ranks with the UN-refined anchor's score. All anchors are denoised, but selection uses the t=0 classifier score over the original anchors; the denoise passes' own confidences are discarded. Geometry is refined, ranking is not (base sel_gap 0.2813,
frac_sel_2x_worse0.4109).The oracle gap is ~92 % IRREDUCIBLE β it is not available headroom. Settled across 47 trained arms: a learned re-scorer recovers at most 8.4 % of it on its own training data; a hand-written cost re-rank recovers 0.0 %; scoring the refined confidences is 2.9Γ WORSE than baseline (that head is unsupervised at denoise timesteps). A target-speed term in the selection score is REFUTED (a GT-perfect speed matcher scores worse than baseline). Selection is not the productive lever on this architecture.
base β XL is a non-separation, not a proof of equality. All three intervals straddle zero; that is evidence of no measurable difference on this corpus at n = 881 / 40 episodes, not of identity.
CONFOUND β scale, anchor count and labels move together. base trained on route labels v2.1, XL on v1. The matched-K control above removes the anchor-count confound; nothing removes the label one. Calibration from flagship v1.5, the only place the label change was measured end-to-end: ADE +0.025 m (not CI-separated) but oracle β0.058 m β i.e. v2.1 labels improved the proposal set by more than either oracle delta measured here, and base is the arm that had them. Do not present a clean scaling conclusion. What IS separable: on ADE/FDE/miss the arms tie. What is NOT separable: the sign and size of the encoder-scale effect on oracle-in-fan. The clean resolution remains one control run (XL-with-v2.1 or base-with-v1) β not yet run. The clean scale test in this ladder is small-vs-base (both v2.1): there small is SEPARATED-worse by ~0.053 m on selected ADE@2s, but better per-anchor on a matched 64-anchor vocabulary β the knee is anchor count, not encoder size.
Closed-loop numbers are ENV-CONFOUNDED β do NOT read them as a model result. On the AlpaSim NuRec suite (n = 12) this arm scores at-fault collision 33.3 %, off-road 16.7 %, pass 6/12, mean score 0.345, dist-to-GT 1.642 m. But REF-C's open-loop ADE on those same reconstructions is 1.52 m β 3.21Γ its real-footage 0.4728, i.e. the input is ~3Γ off the training distribution. Those rates measure model Γ reconstruction fidelity, not the model. The base β₯ XL ordering survives (both eat the same OOD); the levels do not. n = 12 means one scene = 8.3 pp.
Reproducibility gap (declared, not hidden): the TanitEval harness that produced every number above lives on
tanitad-eval:/root/tanitevaland is not committed to the TanitAD repo. The checkpoint survives; the exact evaluator currently does not travel with it.This is a research reference arm, not a deployable driving system. Open-loop ADE does not predict closed-loop behaviour (measured elsewhere in this program: 0.45 m open-loop β 1.69 m closed-loop for the flagship). Claim strength on all metrics above: open-loop / weak.
Files
| file | contents |
|---|---|
ckpt.pt |
full training checkpoint β {model: state_dict (487 tensors), opt: optimizer state, step: 29999}, 1,250,838,325 B, md5 8f10d6f934f4199e11ddc7352e074939 |
config.json |
the run's own config: architecture, args, optimizer, loss weights, v2.1 label stats, data block, param_breakdown |
metrics.json |
final training + val metrics at step 29,999 |
Verify after download: md5sum ckpt.pt must print 8f10d6f934f4199e11ddc7352e074939.
Intended use / license
Internal research and evaluation only. Derived from NVIDIA PhysicalAI-AV (gated corpus) β not redistributable. Not a deployable driving system.
- Downloads last month
- 2