Download docs/VALIDATION.md from Ravenh97/roborender_image3f: direct link, hf CLI and curl.
- Browser
- Download file 7.37 kB
-
https://huggingface.co/Ravenh97/roborender_image3f/resolve/main/docs/VALIDATION.md
- Command line
-
hf download hf://Ravenh97/roborender_image3f/docs/VALIDATION.md
-
curl -L -o VALIDATION.md https://huggingface.co/Ravenh97/roborender_image3f/resolve/main/docs/VALIDATION.md
Validation status and required gates
Current evidence boundary
| Scope | Status in this source release |
|---|---|
| Adapter identity and immutable training provenance | Included and hash-locked |
| Static package and prohibited-artifact scan | Must pass before archive release |
| External Raven foundation identity | Locked; verified by the receiving operator |
| Adapter-specific ONNX export | Not included; generated and verified on target |
| Adapter-specific TensorRT correctness | Not included; required after target build |
| Adapter-specific TeaCache quality/latency selection | Passed on B200; threshold 0.3 selected |
| RTX 4090 physical validation | Not completed |
| RTX 4090 200 ms / 5 Hz certification | No claim |
| Full DAgger observation-to-action validation | Not completed |
The acceleration method was exercised previously on the parent adapter and other GPU classes. Those results demonstrate implementation feasibility only. They are not transferable correctness, quality, memory, or latency evidence for the update-5000 adapter on an RTX 4090. This release deliberately omits old engines, old adapter-fused ONNX, and parent-adapter headline numbers.
The completed adapter-specific B200 campaign is retained under
validation/b200_rgb030_depth030_20260730/ and summarized in
SUMMARY.json. It selects threshold 0.3; it does not certify RTX 4090
latency or memory use.
Package-release gates
Before creating the ZIP:
- regenerate
config/SOURCE_MANIFEST.sha256; - run
python3 bin/doctor.py --stage package; - run Python compilation, shell syntax, and CPU contract tests;
- inspect the archive for
.onnx,.onnx.data,.engine,.plan, timing caches, CUDA Graph artifacts, foundation weights, generated outputs, and stale step-228270 model binaries; - test the ZIP with
unzip -t; - calculate the final ZIP SHA-256; and
- re-extract to a clean directory and rerun package-stage doctor.
Passing these checks means the source bundle is intact and thin. It does not certify inference.
Target export gates
bin/export_onnx.sh must produce:
- an ONNX graph and external-data file;
- a PyTorch correctness capture;
- an export summary; and
export_manifest.json.
The export manifest must match:
- adapter SHA-256
c68de259ce5ee8e4df5a45296be8328242b3ebba351990df1d86f4e3a7a2d588; - the foundation DiT SHA in
config/foundation_models.lock.json; - the packaged exporter SHA;
- the packaged source-manifest SHA;
- the exporter Python/PyTorch/ONNX/ONNXScript/SafeTensors versions; and
- opset 20, batch 1,
ext1/ext2/wrist,416x240, one-frame, lag-256 reference profile.
Reject an export derived from the parent step-228270 adapter.
TensorRT build gates
The target build must:
- record the exact GPU, compute capability, driver, PyTorch/CUDA, TensorRT, precision, workspace, adapter, and ONNX lineage;
- produce finite output from the captured input;
- measure relative RMS error against PyTorch at or below
0.05; - write an engine SHA-256 and target-keyed build manifest;
- promote
assets/generated/engines/currentonly after passing; and - pass
doctor --stage runtime.
CUDA Graph remains disabled unless a separate target-local campaign proves correctness, clean teardown, stability, memory behavior, and a positive end-to-end latency delta.
Functional smoke gates
Using assets/fixtures/lag256/request.json:
- every condition image is RGB
416x240; - view order is
ext1,ext2,wrist; - reference and target source indices differ by exactly 256;
- the latent is finite with shape
[1, 16, 1, 90, 52]; - decoded RGB contains three complete views at the original per-view resolution;
- no condition view is silently duplicated, resized, or omitted; and
- output hashes and the runtime/engine manifests are retained.
The fixture is a regression case, not a representative quality benchmark.
TeaCache and five-step quality campaign
Use matched prompt, conditions, seed, model, and target for:
| Variant | Purpose |
|---|---|
| 50-step PyTorch | High-step visual reference |
| 5-step PyTorch, TeaCache off | Coarse-schedule reference |
| 5-step TensorRT, TeaCache off | Isolate TensorRT numerical effect |
| 5-step TensorRT, thresholds 0.1/0.2/0.3/0.4 | Measure cache speed/quality frontier |
Report per-view and aggregate PSNR, SSIM, and a perceptual metric, but do not select a configuration on scalar metrics alone. Review:
- fine material texture and color;
- object identity and sharpness;
- robot/object contacts;
- depth and mask adherence;
- cross-view geometry; and
- prompt/task consistency.
The configured threshold 0.3 passed this campaign on B200. It remains
provisional on each new deployment target until the matched target-local
campaign passes.
Record the actual TeaCache block calls/skips rather than assuming a threshold
always produces a fixed skip pattern.
Temporal and closed-loop campaign
Evaluate four distinct regimes:
- frame-0 bootstrap for targets
t=0..255; - teacher-conditioned true lag-256;
- autoregressive true lag-256 using stored model outputs; and
- deliberately missing one RGB or depth view, matching the fine-tuning dropout semantics.
Do not label an all-three-RGB-none test as in-distribution: this fine-tune drops exactly one RGB view on a triggered sample, not the complete modality.
For autoregressive validation, generate at least 600 sequential frames (20 seconds at 30 fps). Examine offsets 256 and 512 for delayed recurrences. Plot quality/drift versus frame index and compare early, middle, and final segments. Include both quantitative metrics where targets exist and a side-by-side video.
RTX 4090 latency, memory, and endurance
On the physical final target, separate:
- process/model startup;
- prompt-cache miss;
- warm exact-prompt latent generation;
- warm exact-prompt RGB generation;
- sensor/preprocessing and input transfer;
- VAE decode;
- output handoff/serialization;
- DAgger policy computation; and
- complete observation-to-action latency.
For every reported path retain raw iterations and report mean, median, p90, p99, minimum, maximum, failures, GPU/process memory, system RAM, clocks, temperature, power, throttling, and display/co-resident workload.
The primary 5 Hz gate is p90 complete full-RGB application latency at or below 200 ms. A best-case or latent-only iteration cannot satisfy that gate.
Endurance must cover the longest intended demonstration and show:
- no OOM, NaN/Inf, illegal access, TensorRT error, or missed teardown;
- bounded process/device memory growth;
- stable thermals and clocks;
- no stale-input queue growth; and
- acceptable autoregressive image quality.
Evidence record
Retain per target:
- package and foundation doctor reports;
export_manifest.json;engine_build_manifest.json;- TensorRT correctness and microbenchmark JSON;
- raw latent and RGB timing records;
- exact prompt, seed, input hashes, and output hashes;
- representative comparison grids and closed-loop video;
nvidia-smi -qor equivalent power/thermal snapshot;- container digest or bare-metal dependency freeze; and
- application-level DAgger timing and task-quality report.
Only completed evidence may change config/deployment.json,
RELEASE_READY.json, this document, or the model card from “pending” to
“validated.”