Instructions to use pradhaansbhat/Thinking-In-Boxes with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Diffusers
How to use pradhaansbhat/Thinking-In-Boxes with Diffusers:
pip install -U diffusers transformers accelerate
import torch from diffusers import DiffusionPipeline from diffusers.utils import load_image # switch to "mps" for apple devices pipe = DiffusionPipeline.from_pretrained("pradhaansbhat/Thinking-In-Boxes", dtype=torch.bfloat16, device_map="cuda") prompt = "Turn this cat into a dog" input_image = load_image("https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/diffusers/cat.png") image = pipe(image=input_image, prompt=prompt).images[0] - Notebooks
- Google Colab
- Kaggle
|
Download README.md from pradhaansbhat/Thinking-In-Boxes: direct link, hf CLI and curl.
- Browser
- Download file 2.84 kB
-
https://huggingface.co/pradhaansbhat/Thinking-In-Boxes/resolve/main/README.md
- Command line
-
hf download hf://pradhaansbhat/Thinking-In-Boxes/README.md
-
curl -L -o README.md https://huggingface.co/pradhaansbhat/Thinking-In-Boxes/resolve/main/README.md
2.84 kB
| license: cc-by-4.0 | |
| language: | |
| - en | |
| base_model: | |
| - black-forest-labs/FLUX.1-Kontext-dev | |
| pipeline_tag: image-to-image | |
| tags: | |
| - Generative Modeling | |
| - Image Editing | |
| - Geometric Editing | |
| - 3D Vision | |
| datasets: | |
| - pradhaansbhat/Thinking-In-Boxes | |
| library_name: diffusers | |
| # [NeurIPS-2026] Thinking in Boxes: 3D Editing in Real Images Made Easy | |
| <div> | |
| <img src="assets/thumbnail.png"> | |
| </div> | |
| [**Pradhaan S Bhat**](https://pradhaansbhat.github.io/)<sup>1</sup><sup>∗</sup> · [**Naveen Chandra R**](https://www.linkedin.com/in/naveen-chandra-r-7230aa192)<sup>1</sup><sup>∗</sup> · [**Rishubh Parihar**](https://rishubhpar.github.io/)<sup>1</sup> · [**Vaibhav Vavilala**](https://www.linkedin.com/in/vaibhav-vavilala)<sup>2</sup> · [**R. Venkatesh Babu**](https://cds.iisc.ac.in/faculty/venky/)<sup>1</sup> · [**D.A. Forsyth**](http://luthuli.cs.uiuc.edu/~daf/) · [**Anand Bhattad**](https://anandbhattad.github.io/)<sup>4</sup> | |
| <sup>1</sup> Indian Institute of Science | |
| <sup>2</sup> Apple | |
| <sup>3</sup> UIUC | |
| <sup>4</sup> Johns Hopkins University | |
| <sup>∗</sup> Equal Contribution | |
| [](https://arxiv.org/abs/2606.20556) | |
| [](https://thinking-in-boxes.github.io) | |
| [](https://github.com/PradhaanSBhat/Thinking-In-Boxes) | |
| [](https://huggingface.co/datasets/pradhaansbhat/Thinking-In-Boxes) | |
| []() | |
| <div> | |
| <img src="assets/teaser.png"> | |
| </div> | |
| This is the trained LoRA used in the paper **Thinking In Boxes: 3D Editing in Real Images Made Easy**. | |
| Thinking-In-Boxes is an Image-to-Image Generative LoRA for `black-forest-labs/FLUX.1-Kontext-dev` trained for the task of Geometric Image Editing. | |
| This LoRA is trained on the Thinking-In-Boxes dataset available [here](https://huggingface.co/datasets/pradhaansbhat/Thinking-In-Boxes). Details of training are available in the supplementary section of the paper. | |
| ## Usage with Diffusers 🧨 | |
| Please refer to the [GitHub Repository](https://github.com/PradhaanSBhat/Thinking-In-Boxes) on setup, inference and training. | |
| ## Citation | |
| If you find our work useful, please consider citing: | |
| ```bibtex | |
| @misc{bhat2026thinkingboxes3dediting, | |
| title = {Thinking in Boxes: 3D Editing in Real Images Made Easy}, | |
| author = {Pradhaan S Bhat and Naveen Chandra R and Rishubh Parihar and Vaibhav Vavilala and R. Venkatesh Babu and D. A. Forsyth and Anand Bhattad}, | |
| year = {2026}, | |
| eprint = {2606.20556}, | |
| archivePrefix = {arXiv}, | |
| primaryClass = {cs.CV}, | |
| url = {https://arxiv.org/abs/2606.20556} | |
| } | |
| ``` | |
| --- | |