PSHuman / inference.py

Commit History

Reset torch default device to cpu after offload.profile(): fixes device mismatches in econdataset/pymaf/smplx code that assumes CPU-default tensor creation, without affecting mmgps explicit hook-based device management
181be20
verified

Daankular commited on

Add GPU memory diagnostic print to confirm whether size=xlarge is actually being granted
33ea232
verified

Daankular commited on

Try mmgp VerylowRAM_LowVRAM profile: LowRAM_LowVRAM still OOMs during attention
432c498
verified

Daankular commited on

Fix remaining comment typo
3a0919b
verified

Daankular commited on

Fix comment typo (apostrophe escaping) from previous commit
97e5396
verified

Daankular commited on

Exclude image_normalizer from mmgp offload management (its .scale() method bypasses mmgps forward-hook, stranding buffers on CPU); pin it on GPU directly instead
a143759
verified

Daankular commited on

Fix create_mean_pose() to build its numpy array from CPU tensors explicitly, instead of masking it with a global torch default-device reset that broke mmgp's own device handling
7f3d882
verified

Daankular commited on

Reset torch default device to cpu after offload.profile(): mmgp leaves it as cuda, which broke SMPLDataset tensor creation later in the script
fc017ad
verified

Daankular commited on

Use MMGP budget-based offloading/quantization instead of a blanket .to('cuda'): the pipeline genuinely OOMs on this ZeroGPU MIG slice's effective VRAM
9eacf98
verified

Daankular commited on

Disable expandable_segments CUDA allocator: it hits an NVML assertion on MIG-sliced Blackwell GPUs
bb80551
verified

Daankular commited on

Keep PSHuman's custom multiview attention processors but swap their inner xformers.ops.memory_efficient_attention call for torch's native scaled_dot_product_attention (xformers' Hopper kernel crashes on Blackwell/sm_120)
2080275
verified

Daankular commited on

Drop xformers: its Hopper-specific flash-attention kernel crashes on Blackwell (sm_120) with 'CUDA error: invalid argument'. torch 2.8's native SDPA (used automatically by diffusers) replaces it.
4725c84
verified

Daankular commited on

Update inference.py
adb792d
verified

fffiloni commited on

Update inference.py
ab2d09a
verified

fffiloni commited on

Migrated from GitHub
2252f3d
verified

fffiloni commited on