[NeurIPS-2026] Thinking in Boxes: 3D Editing in Real Images Made Easy

Pradhaan S Bhat1∗ · Naveen Chandra R1∗ · Rishubh Parihar1 · Vaibhav Vavilala2 · R. Venkatesh Babu1 · D.A. Forsyth · Anand Bhattad4

1 Indian Institute of Science 2 Apple 3 UIUC 4 Johns Hopkins University

∗ Equal Contribution

arXiv Project Page GitHub Hugging Face Demo

This is the trained LoRA used in the paper Thinking In Boxes: 3D Editing in Real Images Made Easy.

Thinking-In-Boxes is an Image-to-Image Generative LoRA for black-forest-labs/FLUX.1-Kontext-dev trained for the task of Geometric Image Editing.

This LoRA is trained on the Thinking-In-Boxes dataset available here. Details of training are available in the supplementary section of the paper.

Usage with Diffusers 🧨

Please refer to the GitHub Repository on setup, inference and training.

Citation

If you find our work useful, please consider citing:

@misc{bhat2026thinkingboxes3dediting,
  title         = {Thinking in Boxes: 3D Editing in Real Images Made Easy},
  author        = {Pradhaan S Bhat and Naveen Chandra R and Rishubh Parihar and Vaibhav Vavilala and R. Venkatesh Babu and D. A. Forsyth and Anand Bhattad},
  year          = {2026},
  eprint        = {2606.20556},
  archivePrefix = {arXiv},
  primaryClass  = {cs.CV},
  url           = {https://arxiv.org/abs/2606.20556}
}

Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for pradhaansbhat/Thinking-In-Boxes

Finetuned
(63)
this model

Dataset used to train pradhaansbhat/Thinking-In-Boxes

Paper for pradhaansbhat/Thinking-In-Boxes