File size: 2,842 Bytes
8d9dc3f
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
422a51a
b26e73f
422a51a
 
8d9dc3f
 
 
 
 
 
 
 
 
 
 
 
 
 
 
422a51a
b26e73f
422a51a
 
8d9dc3f
 
 
 
32035d8
8d9dc3f
 
 
422a51a
8d9dc3f
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
---
license: cc-by-4.0
language:
- en
base_model:
- black-forest-labs/FLUX.1-Kontext-dev
pipeline_tag: image-to-image
tags:
- Generative Modeling
- Image Editing
- Geometric Editing
- 3D Vision
datasets:
- pradhaansbhat/Thinking-In-Boxes
library_name: diffusers
---

# [NeurIPS-2026] Thinking in Boxes: 3D Editing in Real Images Made Easy

<div>
  <img src="assets/thumbnail.png">
</div>

[**Pradhaan S Bhat**](https://pradhaansbhat.github.io/)<sup>1</sup><sup>∗</sup> · [**Naveen Chandra R**](https://www.linkedin.com/in/naveen-chandra-r-7230aa192)<sup>1</sup><sup>∗</sup> · [**Rishubh Parihar**](https://rishubhpar.github.io/)<sup>1</sup> · [**Vaibhav Vavilala**](https://www.linkedin.com/in/vaibhav-vavilala)<sup>2</sup> · [**R. Venkatesh Babu**](https://cds.iisc.ac.in/faculty/venky/)<sup>1</sup> · [**D.A. Forsyth**](http://luthuli.cs.uiuc.edu/~daf/) · [**Anand Bhattad**](https://anandbhattad.github.io/)<sup>4</sup>

<sup>1</sup> Indian Institute of Science
<sup>2</sup> Apple
<sup>3</sup> UIUC
<sup>4</sup> Johns Hopkins University

<sup>∗</sup> Equal Contribution

[![arXiv](https://img.shields.io/badge/arXiv-2606.20556-b31b1b.svg)](https://arxiv.org/abs/2606.20556)
[![Project Page](https://img.shields.io/badge/Project-Page-1f6feb.svg)](https://thinking-in-boxes.github.io)
[![GitHub](https://img.shields.io/badge/Github-Repository-blue?logo=github)](https://github.com/PradhaanSBhat/Thinking-In-Boxes)
[![Hugging Face](https://img.shields.io/badge/%F0%9F%A4%97%20Hugging%20Face-Dataset-orange)](https://huggingface.co/datasets/pradhaansbhat/Thinking-In-Boxes)
[![Demo](https://img.shields.io/badge/🎬%20Demo-blue.svg)]()

<div>
  <img src="assets/teaser.png">
</div>

This is the trained LoRA used in the paper **Thinking In Boxes: 3D Editing in Real Images Made Easy**.

Thinking-In-Boxes is an Image-to-Image Generative LoRA for `black-forest-labs/FLUX.1-Kontext-dev` trained for the task of Geometric Image Editing.

This LoRA is trained on the Thinking-In-Boxes dataset available [here](https://huggingface.co/datasets/pradhaansbhat/Thinking-In-Boxes). Details of training are available in the supplementary section of the paper.

## Usage with Diffusers 🧨

Please refer to the [GitHub Repository](https://github.com/PradhaanSBhat/Thinking-In-Boxes) on setup, inference and training.

## Citation

If you find our work useful, please consider citing:

```bibtex
@misc{bhat2026thinkingboxes3dediting,
  title         = {Thinking in Boxes: 3D Editing in Real Images Made Easy},
  author        = {Pradhaan S Bhat and Naveen Chandra R and Rishubh Parihar and Vaibhav Vavilala and R. Venkatesh Babu and D. A. Forsyth and Anand Bhattad},
  year          = {2026},
  eprint        = {2606.20556},
  archivePrefix = {arXiv},
  primaryClass  = {cs.CV},
  url           = {https://arxiv.org/abs/2606.20556}
}
```
---