|
Download docs/VALIDATION.md from Ravenh97/roborender_image3f: direct link, hf CLI and curl.
- Browser
- Download file 7.37 kB
-
https://huggingface.co/Ravenh97/roborender_image3f/resolve/main/docs/VALIDATION.md
- Command line
-
hf download hf://Ravenh97/roborender_image3f/docs/VALIDATION.md
-
curl -L -o VALIDATION.md https://huggingface.co/Ravenh97/roborender_image3f/resolve/main/docs/VALIDATION.md
7.37 kB
| # Validation status and required gates | |
| ## Current evidence boundary | |
| | Scope | Status in this source release | | |
| |---|---| | |
| | Adapter identity and immutable training provenance | Included and hash-locked | | |
| | Static package and prohibited-artifact scan | Must pass before archive release | | |
| | External Raven foundation identity | Locked; verified by the receiving operator | | |
| | Adapter-specific ONNX export | Not included; generated and verified on target | | |
| | Adapter-specific TensorRT correctness | Not included; required after target build | | |
| | Adapter-specific TeaCache quality/latency selection | Passed on B200; threshold 0.3 selected | | |
| | RTX 4090 physical validation | Not completed | | |
| | RTX 4090 200 ms / 5 Hz certification | No claim | | |
| | Full DAgger observation-to-action validation | Not completed | | |
| The acceleration method was exercised previously on the parent adapter and | |
| other GPU classes. Those results demonstrate implementation feasibility only. | |
| They are not transferable correctness, quality, memory, or latency evidence | |
| for the update-5000 adapter on an RTX 4090. This release deliberately omits | |
| old engines, old adapter-fused ONNX, and parent-adapter headline numbers. | |
| The completed adapter-specific B200 campaign is retained under | |
| `validation/b200_rgb030_depth030_20260730/` and summarized in | |
| `SUMMARY.json`. It selects threshold `0.3`; it does not certify RTX 4090 | |
| latency or memory use. | |
| ## Package-release gates | |
| Before creating the ZIP: | |
| 1. regenerate `config/SOURCE_MANIFEST.sha256`; | |
| 2. run `python3 bin/doctor.py --stage package`; | |
| 3. run Python compilation, shell syntax, and CPU contract tests; | |
| 4. inspect the archive for `.onnx`, `.onnx.data`, `.engine`, `.plan`, timing | |
| caches, CUDA Graph artifacts, foundation weights, generated outputs, and | |
| stale step-228270 model binaries; | |
| 5. test the ZIP with `unzip -t`; | |
| 6. calculate the final ZIP SHA-256; and | |
| 7. re-extract to a clean directory and rerun package-stage doctor. | |
| Passing these checks means the source bundle is intact and thin. It does not | |
| certify inference. | |
| ## Target export gates | |
| `bin/export_onnx.sh` must produce: | |
| - an ONNX graph and external-data file; | |
| - a PyTorch correctness capture; | |
| - an export summary; and | |
| - `export_manifest.json`. | |
| The export manifest must match: | |
| - adapter SHA-256 | |
| `c68de259ce5ee8e4df5a45296be8328242b3ebba351990df1d86f4e3a7a2d588`; | |
| - the foundation DiT SHA in `config/foundation_models.lock.json`; | |
| - the packaged exporter SHA; | |
| - the packaged source-manifest SHA; | |
| - the exporter Python/PyTorch/ONNX/ONNXScript/SafeTensors versions; and | |
| - opset 20, batch 1, `ext1/ext2/wrist`, `416x240`, one-frame, lag-256 | |
| reference profile. | |
| Reject an export derived from the parent step-228270 adapter. | |
| ## TensorRT build gates | |
| The target build must: | |
| 1. record the exact GPU, compute capability, driver, PyTorch/CUDA, TensorRT, | |
| precision, workspace, adapter, and ONNX lineage; | |
| 2. produce finite output from the captured input; | |
| 3. measure relative RMS error against PyTorch at or below `0.05`; | |
| 4. write an engine SHA-256 and target-keyed build manifest; | |
| 5. promote `assets/generated/engines/current` only after passing; and | |
| 6. pass `doctor --stage runtime`. | |
| CUDA Graph remains disabled unless a separate target-local campaign proves | |
| correctness, clean teardown, stability, memory behavior, and a positive | |
| end-to-end latency delta. | |
| ## Functional smoke gates | |
| Using `assets/fixtures/lag256/request.json`: | |
| - every condition image is RGB `416x240`; | |
| - view order is `ext1`, `ext2`, `wrist`; | |
| - reference and target source indices differ by exactly 256; | |
| - the latent is finite with shape `[1, 16, 1, 90, 52]`; | |
| - decoded RGB contains three complete views at the original per-view | |
| resolution; | |
| - no condition view is silently duplicated, resized, or omitted; and | |
| - output hashes and the runtime/engine manifests are retained. | |
| The fixture is a regression case, not a representative quality benchmark. | |
| ## TeaCache and five-step quality campaign | |
| Use matched prompt, conditions, seed, model, and target for: | |
| | Variant | Purpose | | |
| |---|---| | |
| | 50-step PyTorch | High-step visual reference | | |
| | 5-step PyTorch, TeaCache off | Coarse-schedule reference | | |
| | 5-step TensorRT, TeaCache off | Isolate TensorRT numerical effect | | |
| | 5-step TensorRT, thresholds 0.1/0.2/0.3/0.4 | Measure cache speed/quality frontier | | |
| Report per-view and aggregate PSNR, SSIM, and a perceptual metric, but do not | |
| select a configuration on scalar metrics alone. Review: | |
| - fine material texture and color; | |
| - object identity and sharpness; | |
| - robot/object contacts; | |
| - depth and mask adherence; | |
| - cross-view geometry; and | |
| - prompt/task consistency. | |
| The configured threshold `0.3` passed this campaign on B200. It remains | |
| provisional on each new deployment target until the matched target-local | |
| campaign passes. | |
| Record the actual TeaCache block calls/skips rather than assuming a threshold | |
| always produces a fixed skip pattern. | |
| ## Temporal and closed-loop campaign | |
| Evaluate four distinct regimes: | |
| 1. frame-0 bootstrap for targets `t=0..255`; | |
| 2. teacher-conditioned true lag-256; | |
| 3. autoregressive true lag-256 using stored model outputs; and | |
| 4. deliberately missing one RGB or depth view, matching the fine-tuning | |
| dropout semantics. | |
| Do not label an all-three-RGB-none test as in-distribution: this fine-tune | |
| drops exactly one RGB view on a triggered sample, not the complete modality. | |
| For autoregressive validation, generate at least 600 sequential frames | |
| (20 seconds at 30 fps). Examine offsets 256 and 512 for delayed recurrences. | |
| Plot quality/drift versus frame index and compare early, middle, and final | |
| segments. Include both quantitative metrics where targets exist and a | |
| side-by-side video. | |
| ## RTX 4090 latency, memory, and endurance | |
| On the physical final target, separate: | |
| - process/model startup; | |
| - prompt-cache miss; | |
| - warm exact-prompt latent generation; | |
| - warm exact-prompt RGB generation; | |
| - sensor/preprocessing and input transfer; | |
| - VAE decode; | |
| - output handoff/serialization; | |
| - DAgger policy computation; and | |
| - complete observation-to-action latency. | |
| For every reported path retain raw iterations and report mean, median, p90, | |
| p99, minimum, maximum, failures, GPU/process memory, system RAM, clocks, | |
| temperature, power, throttling, and display/co-resident workload. | |
| The primary 5 Hz gate is p90 complete full-RGB application latency at or below | |
| 200 ms. A best-case or latent-only iteration cannot satisfy that gate. | |
| Endurance must cover the longest intended demonstration and show: | |
| - no OOM, NaN/Inf, illegal access, TensorRT error, or missed teardown; | |
| - bounded process/device memory growth; | |
| - stable thermals and clocks; | |
| - no stale-input queue growth; and | |
| - acceptable autoregressive image quality. | |
| ## Evidence record | |
| Retain per target: | |
| - package and foundation doctor reports; | |
| - `export_manifest.json`; | |
| - `engine_build_manifest.json`; | |
| - TensorRT correctness and microbenchmark JSON; | |
| - raw latent and RGB timing records; | |
| - exact prompt, seed, input hashes, and output hashes; | |
| - representative comparison grids and closed-loop video; | |
| - `nvidia-smi -q` or equivalent power/thermal snapshot; | |
| - container digest or bare-metal dependency freeze; and | |
| - application-level DAgger timing and task-quality report. | |
| Only completed evidence may change `config/deployment.json`, | |
| `RELEASE_READY.json`, this document, or the model card from “pending” to | |
| “validated.” | |