Render-JAM 2B (b1) โ€” RGB and URDF render jointly denoised on one time axis

Post-trained from Cosmos-Predict2.5 2B. The model denoises two streams stacked along the time axis: the camera RGB and a URDF render of the robot arm. The goal is given as the render's last frame, so the target pose enters as a picture rather than as a vector.

Layout

iter_000010000/   torch.distributed.checkpoint (DCP)
iter_000015000/     model/  optim/  scheduler/  trainer/

Each iteration is ~20 GB: model/ 12 GB, optim/ 7.7 GB. Load with torch.distributed.checkpoint; the optim/ shards are only needed to resume training.

Configuration

resolution 432 x 768
frames per stream 45 -> 12 latents (Wan2.1 VAE, temporal /4)
stacked sequence 24 latents (12 RGB + 12 render)
conditioning 3 positions: RGB t=0, render t=0, render t=T-1 (goal)
objective rectified flow
data ActionNet GR1-T1, 2,903 grasp episodes

The RGB stream has no terminal condition โ€” the goal reaches it only through render -> RGB attention.

Measured on 10 held-out episodes

IoU (arm) arm error (px) PSNR LPIPS
iter 10000 0.444 23.76 โ€” 0.160
iter 15000 0.431 22.04 22.04 0.154

IoU saturates between the two while pixel fidelity keeps improving, so 15000 is the better checkpoint despite the slightly lower IoU.

Ablations at iter 15000: removing the goal image drops IoU to 0.338 and inflates the arm area by 1.47x; replacing it with a wrong goal gives 0.350-0.378. Supplying only the goal state vector (no picture) gives 0.365 โ€” the state input is inert.

Downloads last month
-
Video Preview
loading

Model tree for Eurong2/renderjam-2b-b1

Finetuned
(23)
this model