# Validation status and required gates ## Current evidence boundary | Scope | Status in this source release | |---|---| | Adapter identity and immutable training provenance | Included and hash-locked | | Static package and prohibited-artifact scan | Must pass before archive release | | External Raven foundation identity | Locked; verified by the receiving operator | | Adapter-specific ONNX export | Not included; generated and verified on target | | Adapter-specific TensorRT correctness | Not included; required after target build | | Adapter-specific TeaCache quality/latency selection | Passed on B200; threshold 0.3 selected | | RTX 4090 physical validation | Not completed | | RTX 4090 200 ms / 5 Hz certification | No claim | | Full DAgger observation-to-action validation | Not completed | The acceleration method was exercised previously on the parent adapter and other GPU classes. Those results demonstrate implementation feasibility only. They are not transferable correctness, quality, memory, or latency evidence for the update-5000 adapter on an RTX 4090. This release deliberately omits old engines, old adapter-fused ONNX, and parent-adapter headline numbers. The completed adapter-specific B200 campaign is retained under `validation/b200_rgb030_depth030_20260730/` and summarized in `SUMMARY.json`. It selects threshold `0.3`; it does not certify RTX 4090 latency or memory use. ## Package-release gates Before creating the ZIP: 1. regenerate `config/SOURCE_MANIFEST.sha256`; 2. run `python3 bin/doctor.py --stage package`; 3. run Python compilation, shell syntax, and CPU contract tests; 4. inspect the archive for `.onnx`, `.onnx.data`, `.engine`, `.plan`, timing caches, CUDA Graph artifacts, foundation weights, generated outputs, and stale step-228270 model binaries; 5. test the ZIP with `unzip -t`; 6. calculate the final ZIP SHA-256; and 7. re-extract to a clean directory and rerun package-stage doctor. Passing these checks means the source bundle is intact and thin. It does not certify inference. ## Target export gates `bin/export_onnx.sh` must produce: - an ONNX graph and external-data file; - a PyTorch correctness capture; - an export summary; and - `export_manifest.json`. The export manifest must match: - adapter SHA-256 `c68de259ce5ee8e4df5a45296be8328242b3ebba351990df1d86f4e3a7a2d588`; - the foundation DiT SHA in `config/foundation_models.lock.json`; - the packaged exporter SHA; - the packaged source-manifest SHA; - the exporter Python/PyTorch/ONNX/ONNXScript/SafeTensors versions; and - opset 20, batch 1, `ext1/ext2/wrist`, `416x240`, one-frame, lag-256 reference profile. Reject an export derived from the parent step-228270 adapter. ## TensorRT build gates The target build must: 1. record the exact GPU, compute capability, driver, PyTorch/CUDA, TensorRT, precision, workspace, adapter, and ONNX lineage; 2. produce finite output from the captured input; 3. measure relative RMS error against PyTorch at or below `0.05`; 4. write an engine SHA-256 and target-keyed build manifest; 5. promote `assets/generated/engines/current` only after passing; and 6. pass `doctor --stage runtime`. CUDA Graph remains disabled unless a separate target-local campaign proves correctness, clean teardown, stability, memory behavior, and a positive end-to-end latency delta. ## Functional smoke gates Using `assets/fixtures/lag256/request.json`: - every condition image is RGB `416x240`; - view order is `ext1`, `ext2`, `wrist`; - reference and target source indices differ by exactly 256; - the latent is finite with shape `[1, 16, 1, 90, 52]`; - decoded RGB contains three complete views at the original per-view resolution; - no condition view is silently duplicated, resized, or omitted; and - output hashes and the runtime/engine manifests are retained. The fixture is a regression case, not a representative quality benchmark. ## TeaCache and five-step quality campaign Use matched prompt, conditions, seed, model, and target for: | Variant | Purpose | |---|---| | 50-step PyTorch | High-step visual reference | | 5-step PyTorch, TeaCache off | Coarse-schedule reference | | 5-step TensorRT, TeaCache off | Isolate TensorRT numerical effect | | 5-step TensorRT, thresholds 0.1/0.2/0.3/0.4 | Measure cache speed/quality frontier | Report per-view and aggregate PSNR, SSIM, and a perceptual metric, but do not select a configuration on scalar metrics alone. Review: - fine material texture and color; - object identity and sharpness; - robot/object contacts; - depth and mask adherence; - cross-view geometry; and - prompt/task consistency. The configured threshold `0.3` passed this campaign on B200. It remains provisional on each new deployment target until the matched target-local campaign passes. Record the actual TeaCache block calls/skips rather than assuming a threshold always produces a fixed skip pattern. ## Temporal and closed-loop campaign Evaluate four distinct regimes: 1. frame-0 bootstrap for targets `t=0..255`; 2. teacher-conditioned true lag-256; 3. autoregressive true lag-256 using stored model outputs; and 4. deliberately missing one RGB or depth view, matching the fine-tuning dropout semantics. Do not label an all-three-RGB-none test as in-distribution: this fine-tune drops exactly one RGB view on a triggered sample, not the complete modality. For autoregressive validation, generate at least 600 sequential frames (20 seconds at 30 fps). Examine offsets 256 and 512 for delayed recurrences. Plot quality/drift versus frame index and compare early, middle, and final segments. Include both quantitative metrics where targets exist and a side-by-side video. ## RTX 4090 latency, memory, and endurance On the physical final target, separate: - process/model startup; - prompt-cache miss; - warm exact-prompt latent generation; - warm exact-prompt RGB generation; - sensor/preprocessing and input transfer; - VAE decode; - output handoff/serialization; - DAgger policy computation; and - complete observation-to-action latency. For every reported path retain raw iterations and report mean, median, p90, p99, minimum, maximum, failures, GPU/process memory, system RAM, clocks, temperature, power, throttling, and display/co-resident workload. The primary 5 Hz gate is p90 complete full-RGB application latency at or below 200 ms. A best-case or latent-only iteration cannot satisfy that gate. Endurance must cover the longest intended demonstration and show: - no OOM, NaN/Inf, illegal access, TensorRT error, or missed teardown; - bounded process/device memory growth; - stable thermals and clocks; - no stale-input queue growth; and - acceptable autoregressive image quality. ## Evidence record Retain per target: - package and foundation doctor reports; - `export_manifest.json`; - `engine_build_manifest.json`; - TensorRT correctness and microbenchmark JSON; - raw latent and RGB timing records; - exact prompt, seed, input hashes, and output hashes; - representative comparison grids and closed-loop video; - `nvidia-smi -q` or equivalent power/thermal snapshot; - container digest or bare-metal dependency freeze; and - application-level DAgger timing and task-quality report. Only completed evidence may change `config/deployment.json`, `RELEASE_READY.json`, this document, or the model card from “pending” to “validated.”