File size: 8,664 Bytes
cdcb732
 
 
367247d
 
cdcb732
eab31d8
367247d
 
 
 
 
eab31d8
367247d
cdcb732
 
 
 
d4644e7
 
 
 
eab31d8
 
cdcb732
367247d
 
eab31d8
39f4a38
367247d
 
fdb3b5d
39f4a38
cdcb732
 
39f4a38
 
 
367247d
6f3f10e
cdcb732
39f4a38
cdcb732
8342dae
43dd849
8342dae
43dd849
 
 
 
 
cdcb732
43dd849
 
cdcb732
6f3f10e
 
59d9b25
cdcb732
6f3f10e
 
 
 
 
 
cdcb732
8de96f9
 
 
 
 
43dd849
 
 
 
6f3f10e
43dd849
 
 
8342dae
 
eab31d8
8342dae
eab31d8
 
 
 
 
 
 
 
 
 
 
cdcb732
 
eab31d8
cdcb732
eab31d8
cdcb732
eab31d8
 
 
 
 
 
cdcb732
eab31d8
cdcb732
 
8342dae
 
cdcb732
43dd849
eab31d8
43dd849
cdcb732
 
eab31d8
 
 
8de96f9
eab31d8
 
 
 
cdcb732
 
367247d
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
---
base_model: stabilityai/stable-diffusion-3-medium-diffusers
language: en
license: cc-by-nc-4.0
pipeline_tag: image-to-image
tags:
- arxiv:2609.06490
- super-resolution
- image-super-resolution
- extreme-zoom
- chain-of-zoom
- diffusion
- self-distillation
- faithfulness
---

# OracleZoom

[Shubhashis Roy Dipta](https://roydipta.com)\*, [Sourajit Saha](https://sourajitcs.github.io/)\*, [Shaswati Saha](https://scholar.google.com/citations?hl=en&user=_pTdzsAAAAAJ&view_op=list_works&sortby=pubdate), [Nobin Sarwar](https://smsnobin77.github.io/) 路 University of Maryland, Baltimore County

\*Equal contribution.

**Reference-constrained recursive super-resolution, inspired by on-policy self-distillation.**  
OracleZoom drives recursive 4x super-resolution out to 256x while staying faithful to the observed scene. Ground truth runs out at 4x, so OracleZoom trains on its own recursive predictions and uses the last available ground-truth image as a reference past that boundary. **This repo is self-contained**: the merged model, the inference code, and the required checkpoints are all here. You only download two public base models (Stable Diffusion 3-medium, Qwen2.5-VL-3B) automatically.

[![Base](https://img.shields.io/badge/Base-SD3%20%2B%20Qwen2.5--VL-blue)](https://huggingface.co/stabilityai/stable-diffusion-3-medium-diffusers)  
[![Method](https://img.shields.io/badge/Method-Chain--of--Zoom-orange)](https://github.com/bryanswkim/Chain-of-Zoom)  
[![Paper](https://img.shields.io/badge/Paper-arXiv%3A2609.06490-red)](https://arxiv.org/abs/2609.06490)  
[![Demo](https://img.shields.io/badge/Demo-%F0%9F%A4%97%20Space-yellow)](https://huggingface.co/spaces/dipta007/OracleZoom)  
[![Code](https://img.shields.io/badge/Code-GitHub-black)](https://github.com/dipta007/OracleZoom)  
[![Project](https://img.shields.io/badge/Project-Page-lightgrey)](https://dipta007.github.io/OracleZoom/)  
[![Data](https://img.shields.io/badge/Data-%F0%9F%A4%97%204KLSDB-yellow)](https://huggingface.co/datasets/dipta007/OracleZoom-4KLSDB-train)  
[![Collection](https://img.shields.io/badge/Collection-%F0%9F%A4%97%20OracleZoom-yellow)](https://huggingface.co/collections/dipta007/oraclezoom)  
[![License](https://img.shields.io/badge/License-CC--BY--NC--4.0-lightgrey)](https://creativecommons.org/licenses/by-nc/4.0/)

**Try it in the browser:** [馃 Space](https://huggingface.co/spaces/dipta007/OracleZoom) (no install, runs on ZeroGPU)

**Paper:** [arXiv:2609.06490](https://arxiv.org/abs/2609.06490) ([HF paper page](https://huggingface.co/papers/2609.06490)) 路 **Code:** https://github.com/dipta007/OracleZoom 路 **Project page:** https://dipta007.github.io/OracleZoom/ 路 **Everything in one place:** [馃 collection](https://huggingface.co/collections/dipta007/oraclezoom)

## Quickstart (one image, all scales)

No GPU? Use the [馃 Space](https://huggingface.co/spaces/dipta007/OracleZoom) instead. To run it yourself you need one NVIDIA GPU (~16 GB) and Python 3.10.

```bash
# 0. One-time: Stable Diffusion 3 is gated, so accept its license on HF, then log in
pip install -U "huggingface_hub[cli]"
hf auth login

# 1. Download this repo (merged model + code + checkpoints)
hf download dipta007/OracleZoom --local-dir OracleZoom
cd OracleZoom

# 2. Install dependencies
pip install -r requirements.txt

# 3. Super-resolve ONE image (4x -> 16x -> 64x -> 256x)
python inference.py --input /path/to/photo.jpg --output ./outputs
```

**Results** in `./outputs/`:
`photo_1x.png` (the 512x512 input crop), `photo_4x.png`, `photo_16x.png`, `photo_64x.png`, `photo_256x.png`.

Stable Diffusion 3-medium and Qwen2.5-VL-3B download automatically on first run.

> **Batching many images:** `inference.py` exposes `zoom_image(sr, model, proc, pvi, image_path, out_dir)`. Build the models once (`build_sr(...)`, `build_vlm(...)`) and call `zoom_image` in a loop over your images.

## Training data
The curated training set is released separately at
[dipta007/OracleZoom-4KLSDB-train](https://huggingface.co/datasets/dipta007/OracleZoom-4KLSDB-train).
The released model uses its `1k` config (1,000 curated 4K images).

## What's in this repo
| Path | What it is |
|---|---|
| `merged_transformer.safetensors` | The OracleZoom super-resolution transformer (SD3 + Chain-of-Zoom's SR module + our distilled adapter, merged), fp32, ~8.35 GB. |
| `inference.py` | Self-contained runner: one image in, all scales out (recursive zoom + VLM prompting). |
| `coz/` | Vendored Chain-of-Zoom inference code (the one-step SR wrapper + helpers). |
| `ckpt/` | Chain-of-Zoom's SR-VAE and VLM-prompt (Qwen LoRA) checkpoints needed by the pipeline. |
| `requirements.txt` | Python dependencies. |

## Method
Recursive SR (Chain-of-Zoom) reuses a 4x backbone step after step to reach 16x-256x. The ground truth needed to supervise those steps grows geometrically: a 256x target would need 52 GB per image. So supervision stops at 4x, and every deeper step is **blind**, seeing only a blurred crop of its own previous output. Errors compound and the invented detail may be hallucinated.

**Training on the model's own predictions, constrained by the last reference.** Training follows the model's own 4x then 16x predictions and backpropagates through both steps. Five objectives separate what the reference can verify from what it cannot:

| Objective | What it does | Weight |
|---|---|---|
| Direct supervision | LPIPS between the decoded 4x prediction and ground truth. | `w_4x` 1.0 |
| Cross-scale consistency | Project the 16x prediction back to the reference resolution, then match the aligned region of the 4x ground truth. | `w_deep` 1.0 |
| Quality guidance | Frozen TOPIQ-NR guides the detail the reference cannot verify. | `beta_reward` 0.4 |
| KL prior | Keep adapted latents close to the pretrained model's prediction on the same input. | `beta_kl` 8.0 |
| EMA consistency | A slowly updated adapter copy (decay 0.95), run on the ground-truth input, stabilizes training at the supervision boundary. | `lambda_ema` 0.1 |

Only a rank-16 adapter is trained (7.1M parameters, 1,000 curated 4K images); the backbone, VAE, and prompter stay frozen. The KL prior is what keeps quality guidance honest: removing it raises 16x CLIPIQA from 0.714 to 0.794, but worsens projected DISTS from 0.215 to 0.330 and raises judged hallucination from 0.303 to 0.907. The released weights have the adapter already merged in.

## Results
Every method runs inside the same zoom loop and is scored by the same code, over seven test sets (4KLSDB, DIV2K, DIV8K, DRealSR, FFHQ, Flickr2K, RealSR) and 4x to 256x.

| Axis | Metric | Ours | Best baseline |
|---|---|---|---|
| Quality (no-reference) | CLIPIQA, mean over scales | **0.713** | 0.621 (Chain-of-Zoom) |
| Quality (no-reference) | CLIPIQA @256x | **0.706** | 0.579 (Chain-of-Zoom) |
| Fidelity @4x (ground truth exists) | LPIPS | **0.199** | 0.215 (Chain-of-Zoom) |
| Fidelity @4x | DISTS | **0.160** | 0.164 (SeeSR) |
| Deeper scales (InternVL3.5-38B judge) | preferred over Chain-of-Zoom | **68% @64x, 78% @256x** | - |
| Deeper scales | hallucination rate | **0.21 @64x, 0.14 @256x** | 0.55, 0.70 (Chain-of-Zoom) |

Past 4x there is no ground truth, so the judge measures consistency with the preceding zooms, not recovery of unseen detail. Ties and abstentions are excluded from the win rate. No-reference quality scores alone do not establish agreement with the observed scene.

## Intended Use
- **In-scope:** research on faithful extreme (recursive) super-resolution of natural photographs.
- **Out-of-scope:** forensic/evidentiary use (detail past 4x is generated, not recovered); real-camera-zoom claims (the benchmark uses synthetic center-crop zoom).

## Acknowledgements & Licensing
The `coz/` code and the checkpoints in `ckpt/` are from [Chain-of-Zoom](https://github.com/bryanswkim/Chain-of-Zoom) and are redistributed here for convenience; please respect their original license and cite them. The pipeline uses Stable Diffusion 3-medium and Qwen2.5-VL-3B under their respective licenses. OracleZoom's own contribution (the trained adapter, merged into `merged_transformer.safetensors`) is released for **research, non-commercial** use (CC-BY-NC-4.0).

## Citation
```bibtex
@misc{dipta2026oraclezoomonpolicyselfdistillationinspired,
  title={OracleZoom: On-Policy Self-Distillation Inspired Reference-Constrained Recursive Image Super Resolution},
  author={Shubhashis Roy Dipta and Sourajit Saha and Shaswati Saha and Nobin Sarwar},
  year={2026},
  eprint={2609.06490},
  archivePrefix={arXiv},
  primaryClass={cs.CV},
  url={https://arxiv.org/abs/2609.06490}
}
```
Please also cite Chain-of-Zoom and OSEDiff.