roborender_image3f / docs /VALIDATION.md
Ravenh97's picture
Image3F rgb030-depth030 step-5000: adapter + deployment source
dd0a8f2 verified
|
Raw History Blame Contribute Delete
7.37 kB

Validation status and required gates

Current evidence boundary

Scope Status in this source release
Adapter identity and immutable training provenance Included and hash-locked
Static package and prohibited-artifact scan Must pass before archive release
External Raven foundation identity Locked; verified by the receiving operator
Adapter-specific ONNX export Not included; generated and verified on target
Adapter-specific TensorRT correctness Not included; required after target build
Adapter-specific TeaCache quality/latency selection Passed on B200; threshold 0.3 selected
RTX 4090 physical validation Not completed
RTX 4090 200 ms / 5 Hz certification No claim
Full DAgger observation-to-action validation Not completed

The acceleration method was exercised previously on the parent adapter and other GPU classes. Those results demonstrate implementation feasibility only. They are not transferable correctness, quality, memory, or latency evidence for the update-5000 adapter on an RTX 4090. This release deliberately omits old engines, old adapter-fused ONNX, and parent-adapter headline numbers.

The completed adapter-specific B200 campaign is retained under validation/b200_rgb030_depth030_20260730/ and summarized in SUMMARY.json. It selects threshold 0.3; it does not certify RTX 4090 latency or memory use.

Package-release gates

Before creating the ZIP:

  1. regenerate config/SOURCE_MANIFEST.sha256;
  2. run python3 bin/doctor.py --stage package;
  3. run Python compilation, shell syntax, and CPU contract tests;
  4. inspect the archive for .onnx, .onnx.data, .engine, .plan, timing caches, CUDA Graph artifacts, foundation weights, generated outputs, and stale step-228270 model binaries;
  5. test the ZIP with unzip -t;
  6. calculate the final ZIP SHA-256; and
  7. re-extract to a clean directory and rerun package-stage doctor.

Passing these checks means the source bundle is intact and thin. It does not certify inference.

Target export gates

bin/export_onnx.sh must produce:

  • an ONNX graph and external-data file;
  • a PyTorch correctness capture;
  • an export summary; and
  • export_manifest.json.

The export manifest must match:

  • adapter SHA-256 c68de259ce5ee8e4df5a45296be8328242b3ebba351990df1d86f4e3a7a2d588;
  • the foundation DiT SHA in config/foundation_models.lock.json;
  • the packaged exporter SHA;
  • the packaged source-manifest SHA;
  • the exporter Python/PyTorch/ONNX/ONNXScript/SafeTensors versions; and
  • opset 20, batch 1, ext1/ext2/wrist, 416x240, one-frame, lag-256 reference profile.

Reject an export derived from the parent step-228270 adapter.

TensorRT build gates

The target build must:

  1. record the exact GPU, compute capability, driver, PyTorch/CUDA, TensorRT, precision, workspace, adapter, and ONNX lineage;
  2. produce finite output from the captured input;
  3. measure relative RMS error against PyTorch at or below 0.05;
  4. write an engine SHA-256 and target-keyed build manifest;
  5. promote assets/generated/engines/current only after passing; and
  6. pass doctor --stage runtime.

CUDA Graph remains disabled unless a separate target-local campaign proves correctness, clean teardown, stability, memory behavior, and a positive end-to-end latency delta.

Functional smoke gates

Using assets/fixtures/lag256/request.json:

  • every condition image is RGB 416x240;
  • view order is ext1, ext2, wrist;
  • reference and target source indices differ by exactly 256;
  • the latent is finite with shape [1, 16, 1, 90, 52];
  • decoded RGB contains three complete views at the original per-view resolution;
  • no condition view is silently duplicated, resized, or omitted; and
  • output hashes and the runtime/engine manifests are retained.

The fixture is a regression case, not a representative quality benchmark.

TeaCache and five-step quality campaign

Use matched prompt, conditions, seed, model, and target for:

Variant Purpose
50-step PyTorch High-step visual reference
5-step PyTorch, TeaCache off Coarse-schedule reference
5-step TensorRT, TeaCache off Isolate TensorRT numerical effect
5-step TensorRT, thresholds 0.1/0.2/0.3/0.4 Measure cache speed/quality frontier

Report per-view and aggregate PSNR, SSIM, and a perceptual metric, but do not select a configuration on scalar metrics alone. Review:

  • fine material texture and color;
  • object identity and sharpness;
  • robot/object contacts;
  • depth and mask adherence;
  • cross-view geometry; and
  • prompt/task consistency.

The configured threshold 0.3 passed this campaign on B200. It remains provisional on each new deployment target until the matched target-local campaign passes. Record the actual TeaCache block calls/skips rather than assuming a threshold always produces a fixed skip pattern.

Temporal and closed-loop campaign

Evaluate four distinct regimes:

  1. frame-0 bootstrap for targets t=0..255;
  2. teacher-conditioned true lag-256;
  3. autoregressive true lag-256 using stored model outputs; and
  4. deliberately missing one RGB or depth view, matching the fine-tuning dropout semantics.

Do not label an all-three-RGB-none test as in-distribution: this fine-tune drops exactly one RGB view on a triggered sample, not the complete modality.

For autoregressive validation, generate at least 600 sequential frames (20 seconds at 30 fps). Examine offsets 256 and 512 for delayed recurrences. Plot quality/drift versus frame index and compare early, middle, and final segments. Include both quantitative metrics where targets exist and a side-by-side video.

RTX 4090 latency, memory, and endurance

On the physical final target, separate:

  • process/model startup;
  • prompt-cache miss;
  • warm exact-prompt latent generation;
  • warm exact-prompt RGB generation;
  • sensor/preprocessing and input transfer;
  • VAE decode;
  • output handoff/serialization;
  • DAgger policy computation; and
  • complete observation-to-action latency.

For every reported path retain raw iterations and report mean, median, p90, p99, minimum, maximum, failures, GPU/process memory, system RAM, clocks, temperature, power, throttling, and display/co-resident workload.

The primary 5 Hz gate is p90 complete full-RGB application latency at or below 200 ms. A best-case or latent-only iteration cannot satisfy that gate.

Endurance must cover the longest intended demonstration and show:

  • no OOM, NaN/Inf, illegal access, TensorRT error, or missed teardown;
  • bounded process/device memory growth;
  • stable thermals and clocks;
  • no stale-input queue growth; and
  • acceptable autoregressive image quality.

Evidence record

Retain per target:

  • package and foundation doctor reports;
  • export_manifest.json;
  • engine_build_manifest.json;
  • TensorRT correctness and microbenchmark JSON;
  • raw latent and RGB timing records;
  • exact prompt, seed, input hashes, and output hashes;
  • representative comparison grids and closed-loop video;
  • nvidia-smi -q or equivalent power/thermal snapshot;
  • container digest or bare-metal dependency freeze; and
  • application-level DAgger timing and task-quality report.

Only completed evidence may change config/deployment.json, RELEASE_READY.json, this document, or the model card from “pending” to “validated.”