Thinking-In-Boxes / README.md
pradhaansbhat's picture
Update README.md
066e486 verified
|
Raw History Blame Contribute Delete
2.84 kB
metadata
license: cc-by-4.0
language:
  - en
base_model:
  - black-forest-labs/FLUX.1-Kontext-dev
pipeline_tag: image-to-image
tags:
  - Generative Modeling
  - Image Editing
  - Geometric Editing
  - 3D Vision
datasets:
  - pradhaansbhat/Thinking-In-Boxes
library_name: diffusers

[NeurIPS-2026] Thinking in Boxes: 3D Editing in Real Images Made Easy

Pradhaan S Bhat1∗ · Naveen Chandra R1∗ · Rishubh Parihar1 · Vaibhav Vavilala2 · R. Venkatesh Babu1 · D.A. Forsyth · Anand Bhattad4

1 Indian Institute of Science 2 Apple 3 UIUC 4 Johns Hopkins University

∗ Equal Contribution

arXiv Project Page GitHub Hugging Face Demo

This is the trained LoRA used in the paper Thinking In Boxes: 3D Editing in Real Images Made Easy.

Thinking-In-Boxes is an Image-to-Image Generative LoRA for black-forest-labs/FLUX.1-Kontext-dev trained for the task of Geometric Image Editing.

This LoRA is trained on the Thinking-In-Boxes dataset available here. Details of training are available in the supplementary section of the paper.

Usage with Diffusers 🧨

Please refer to the GitHub Repository on setup, inference and training.

Citation

If you find our work useful, please consider citing:

@misc{bhat2026thinkingboxes3dediting,
  title         = {Thinking in Boxes: 3D Editing in Real Images Made Easy},
  author        = {Pradhaan S Bhat and Naveen Chandra R and Rishubh Parihar and Vaibhav Vavilala and R. Venkatesh Babu and D. A. Forsyth and Anand Bhattad},
  year          = {2026},
  eprint        = {2606.20556},
  archivePrefix = {arXiv},
  primaryClass  = {cs.CV},
  url           = {https://arxiv.org/abs/2606.20556}
}