Spaces:
Running on Zero
[Admin] Fix garbled output: bind each transformer block to its own AoTI weights
The Space currently returns pure noise for every generation (verified live: 30 steps, guidance 4.0, seed 42 on the model card's own example returns an 832x1152 image of RGB noise).
Root cause: the AoTI graph runs all 40 blocks with block 0's weights
package.pt2 is exported with the block's weights as runtime constants, so one compiled graph can serve all 40 JoyImageEditPlusTransformerBlocks — but each block still has to be bound to its own weights. The current code builds a single ZeroGPUCompiledModel from _blocks[0].state_dict() and applies that same instance to every block:
_weights = ZeroGPUWeights(_blocks[0].state_dict(), to_cuda=True)
_compiled = ZeroGPUCompiledModel(_pt2_path, _weights)
for _blk in _blocks:
spaces.aoti_apply(_compiled, _blk)
ZeroGPUCompiledModel.__call__ loads self.weights.constants_map once and reuses it, so the transformer effectively becomes block 0 repeated 40 times — no error, just noise.
The supported API does the right thing: given a ModuleList, spaces.aoti_load_from_package_dir patches every block with one shared LazyAOTIModel while keeping that block's own state_dict as the constants to load. (Same mechanism as spaces.aoti_blocks_load; used directly here because the artifact lives in a dataset repo.)
Fallout: the demo was never actually resident on the GPU
spaces.aoti_apply calls drain_module_parameters, which replaced all 40 blocks' parameters with empty tensors — silently dropping ~29GB from the ZeroGPU pack. That is the only reason the Space fit the default 48GB MIG slice. With the weights correctly retained the pack is 50.3GB, which does not fit, so the GPU function now requests size="xlarge" (full 96GB card). Verified: without it the first CUDA op OOMs (MIG 2g.48gb, 47.4GiB).
Two robustness fixes found along the way
requirements.txtpinsdiffusersto a moving branch. The revision the Space is currently running is not the revision a rebuild would install. On the branch tipJoyImageEditPlusPipelinerenamedvae_image_processor->image_processorand changed the block'simage_rotary_embargument from((cos, sin), None)to(cos, sin)— the first breaksapp.py, the second invalidatespackage.pt2. Pinned to3fa516a2, the branch tip the Space and the AoTI artifact were both built from.app.pyno longer callspipe.vae_image_processor.height/widthare left unset; the pipeline derives the same 1024-base bucket from the last reference image internally, so this works on either pipeline revision.
Verification
Tested end-to-end on a duplicate Space with the same inputs, prompt and seed as the examples/ row. Before: RGB noise. After: a coherent generation matching the reference output.
Alternative if xlarge (2x visitor quota) is not wanted: drop the AoTI block and go back to pipe.enable_model_cpu_offload(). Peak VRAM then stays under the 48GB slice, at the cost of the compiled graph and per-call weight transfers. Happy to switch this PR to that shape instead.
Admi