YAML Metadata Warning:empty or missing yaml metadata in repo card
Check out the documentation for more information.
Qwen-Image-2.1-viewpoint-orbit-LoRA
An RGBA viewpoint-orbit LoRA for Qwen/Qwen-Image-2.1. Give it ONE transparent (RGBA) image of an object and a relative camera instruction; it returns the same object from the requested viewpoint β also as a transparent RGBA image.
- Base model: Qwen/Qwen-Image-2.1, bf16, unquantized
- Adapter: rank 32 / alpha 32, trained 2,000 steps at 768 px (~2.2 s/step on one A100-80GB) with ostris/ai-toolkit, arch
qwen_image_2,rgba: true - Training data: synthetic orbit renders of Google Scanned Objects β ML-Intern-lab/gso-orbit-rgba (renders, pose metadata, pair lists and split all included)
Try it in the Space: ML-Intern-lab/Qwen-Image-2.1-viewpoint-orbit-LoRA.
Usage (verified)
This exact path was run end-to-end during evaluation (verify_diffusers_lora.py on the pinned diffusers commit; loads with zero dropped-key warnings, output is RGBA):
import torch
from huggingface_hub import hf_hub_download
from diffusers import QwenImage21Pipeline
pipe = QwenImage21Pipeline.from_pretrained("Qwen/Qwen-Image-2.1", torch_dtype=torch.bfloat16).to("cuda")
lora = hf_hub_download("ML-Intern-lab/Qwen-Image-2.1-viewpoint-orbit-LoRA",
"checkpoints/steps2000res768/orbit_alpha_lora_gate_up_split.safetensors")
pipe.load_lora_weights(lora) # diffusers-keyed file: zero dropped keys
src = ... # your source image, 768x768 RGBA PIL.Image
out = pipe(image=src, width=768, height=768, num_inference_steps=40, true_cfg_scale=1.0,
prompt="<orbit> rotate the camera 90 degrees to the right, eye level. "
"The image has alpha channel and the background is transparent.").images[0]
out.save("orbit.png") # out.mode == "RGBA" β transparency is preserved
Notes:
true_cfg_scale=1.0(no CFG) is how the model was evaluated and how it is meant to be run here.- The output is RGBA because the 2.1 VAE is natively RGBA β no matting post-processing involved.
- The file
orbit_alpha_lora_gate_up_split.safetensorsis the diffusers-keyed adapter. The original trainer output (checkpoints/steps2000res768/orbit_alpha_lora/orbit_alpha_lora.safetensors) uses ComfyUI-style keys and mostly loads in diffusers, but silently drops theimg_mlp.gate_upbranch (a fused SwiGLU projection diffusers stores as two separate linears). Use the_gate_up_splitfile; both are verified to give matching metrics. - Verified against
diffusers@ commit0121a91f9d419ff7234c8a5923f82c244e6f1914withtransformers>=5.17,peft,torch 2.13+.
Instruction grammar
Azimuth is relative to the source view (scanned objects β and user uploads β have no canonical front); elevation is absolute. Every instruction starts with <orbit>. The full training set is exactly these 23 instructions (80β81 pairs each, 1,844 pairs total):
| low angle | eye level | elevated | |
|---|---|---|---|
| rotate 45Β° left | β | β | β |
| rotate 90Β° left | β | β | β |
| rotate 135Β° left | β | β | β |
| rotate 45Β° right | β | β | β |
| rotate 90Β° right | β | β | β |
| rotate 135Β° right | β | β | β |
| rotate 180 degrees (no left/right) | β | β | β |
| keep the camera angle (elevation only) | β (low angle) | β | β (elevated) |
Verbatim forms:
<orbit> rotate the camera {45|90|135|180} degrees to the {left|right}, {low angle|eye level|elevated}(180 omits left/right:<orbit> rotate the camera 180 degrees, eye level)<orbit> keep the camera angle, {low angle|elevated}(20% of training pairs)
At inference we append The image has alpha channel and the background is transparent. β chosen by an A/B on the base model (alpha IoU 0.830 with the suffix vs 0.798 without, on 10 pairs).
Results vs the base model (ground truth)
160 held-out edits (40 real-scanned objects Γ 4 targets), 40 inference steps, 768 px, fixed per-pair seeds, metrics against the true renders (LPIPS/PSNR composited on mid-gray):
| config | alpha IoU β | LPIPS β | PSNR β | DINOv2 sim β |
|---|---|---|---|---|
| base model, plain grammar ("rotate the camera 90 degrees to the right") | 0.727 | 0.097 | 21.74 | 0.714 |
base model, <orbit> grammar |
0.731 | 0.099 | 21.74 | 0.713 |
| this LoRA (step-2000 checkpoint) | 0.794 | 0.090 | 22.34 | 0.734 |
By rotation size (alpha IoU): 45Β° 0.835, 90Β° 0.773 (base: ~0.67 β the hard bucket, where hidden sides must be invented), 180Β° 0.835. The step-2000 checkpoint was chosen by best LPIPS with no alpha-IoU loss over {500, 1000, 1500, 2000} on a fixed 48-edit subset; it won on every metric, so no trade-off arose.
Per-edit records: eval/final/, eval/ckpt_select/, eval/baseline/ in ML-Intern-lab/gso-orbit-rgba.
Out-of-domain check
12 text-to-image-generated transparent subjects (cartoon mascot, sneaker, armchair, robot toy, potted plant, game character, perfume bottle, backpack, headphones, coffee mug, desk lamp, rubber duck), orbited 7 views each at eye level by chaining 45Β° moves β every edit sourced from the original image, never from a generated one. Contact sheets and ping-pong turntable GIFs:
eval/ood/turntable_*.gif in ML-Intern-lab/gso-orbit-rgba β e.g. turntable_sneaker.gif, turntable_robot_toy.gif, turntable_cartoon_mascot.gif.
Where it breaks, honestly:
- Transparency holds β per-view alpha coverage stayed at 0.45β0.49 across all 84 views (no mask collapse, no background bleed-through), and no view dropped below 5% coverage.
- Identity drifts with azimuth. Small 45Β° steps are faithful; by cumulative 180β315Β° (chained as legal 135Β°/90Β°/45Β° instructions) silhouettes stay plausible but fine details (logos, straps, handles) sometimes get re-invented rather than preserved. There is no ground truth for these subjects, so this is a qualitative observation from the contact sheets, not a measured number.
- Style transfer of lighting: the model was trained on clean studio-lit scans; on T2I subjects with baked-in lighting it occasionally carries the source's shading into the new view instead of re-lighting.
Limitations
- Trained exclusively on Google Scanned Objects (household-scale, studio-lit, no ground plane). Performance on other scales/materials is unmeasured; see the OOD section.
- Relative azimuth only β there are no absolute poses ("front view") by design; the camera always moves relative to the input view.
- Elevation is limited to the trained range (low angle β β20Β°, eye level, elevated β +40Β°).
- 90Β° moves are the hardest bucket (hidden sides must be invented); good, but measurably below small rotations.
- Non-commercial β see License.
Training details
- Data: 24 rendered views (8 azimuths Γ 3 elevations) per object, 768Γ768 RGBA, transparent background, three-point lighting, Z-up handled, objects framed to 55β65% of frame; objects whose any-view coverage fell under 8% of the frame were rejected (529 of 1,030 rejected β 501 usable; 461 train / 40 held-out, stratified by category).
- Pairs: 4 per training object (1,844), source always eye-level at a random azimuth; target = source azimuth + one of the 7 relative moves (or same azimuth for the 20% elevation-only pairs) at one of 3 elevations; instruction balance 80/81 across all 23.
- Trainer: ai-toolkit
sd_trainer,control_pathediting,rgba: true,match_target_resdefault, bf16 unquantized (Qwen/Qwen-Image-2.1), adamw8bit, gradient checkpointing, cached text embeddings,timestep_type: weighted, lr 1e-4, batch 1, 2,000 steps, checkpoints every 500. - Loss curve:
checkpoints/steps2000res768/loss_log.jsonlin this repo (2,000 steps, ~0.01β0.2, no divergence).
License
This adapter is a derivative of Qwen-Image-2.1 and is distributed under the Qwen RESEARCH LICENSE AGREEMENT β non-commercial use only. The verbatim license text is included in this repo as LICENSE.
Change notice (per license Β§3b): this repo adds only new LoRA adapter weight files (safetensors) trained on top of Qwen/Qwen-Image-2.1; no base-model files are modified or redistributed.
Built with Qwen. "Qwen is licensed under the Qwen RESEARCH LICENSE AGREEMENT, Copyright (c) 2026 Hangzhou Tongyi Laboratory Technology Co., Ltd. All Rights Reserved."
The repo name uses "Qwen-Image-2.1" descriptively, as a fine-tune of Qwen-Image-2.1, which license Β§4c permits; the adapter itself is called Orbit Alpha (file names keep the orbit_alpha_lora prefix).
Attribution
- Base model: Qwen/Qwen-Image-2.1 (Qwen Research License).
- Training data: Google Scanned Objects (original dataset), obtained via suvadityamuk/google-scanned-objects, CC-BY 4.0. Renders produced with trimesh + pyrender.
- Camera-LoRA idea: credit to fal (fal/Qwen-Image-Edit-2511-Multiple-Angles-LoRA) and dx8152 for the multiple-angles LoRA concept. The relative-azimuth grammar here is deliberately different from fal's absolute poses.
- Tooling: ostris/ai-toolkit for training, diffusers
QwenImage21Pipelinefor inference/eval, LPIPS + DINOv2 for metrics.