File size: 4,423 Bytes
774e461
 
 
 
 
 
 
 
 
 
c18ebfc
774e461
 
 
 
 
c18ebfc
a7ffa7e
a916e21
774e461
c18ebfc
774e461
c18ebfc
774e461
c18ebfc
774e461
c18ebfc
774e461
c18ebfc
774e461
c18ebfc
774e461
c18ebfc
 
 
774e461
c18ebfc
774e461
c18ebfc
 
 
 
774e461
c18ebfc
774e461
c18ebfc
 
774e461
c18ebfc
 
 
 
774e461
c18ebfc
 
 
774e461
a7ffa7e
774e461
c18ebfc
774e461
a7ffa7e
774e461
c18ebfc
774e461
 
c18ebfc
 
 
 
 
 
774e461
 
c18ebfc
 
774e461
a7ffa7e
 
 
 
 
 
774e461
c18ebfc
 
 
 
 
 
 
 
 
 
 
 
 
 
774e461
c18ebfc
 
 
 
a7ffa7e
774e461
c18ebfc
774e461
a7ffa7e
 
 
 
 
 
 
 
774e461
c18ebfc
 
a7ffa7e
c18ebfc
 
a7ffa7e
c18ebfc
a7ffa7e
c18ebfc
 
a7ffa7e
c18ebfc
a7ffa7e
c18ebfc
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
---
license: cc-by-nc-4.0
base_model: Qwen/Qwen3-VL-4B-Instruct
pipeline_tag: image-text-to-text
library_name: transformers
language:
  - en
tags:
  - remote-sensing
  - visual-grounding
  - horizontal-bounding-box
  - oriented-bounding-box
  - reinforcement-learning
  - qwen3-vl
---

<div align="center">

<h1 align="center">GeoBox-R1: Curriculum-Guided SFT and Geometric RL for Unified Box-Level Remote Sensing Visual Grounding</h1>

Chenxi Lan\*, Yuchen Wu\*, Minghang Zhou, Tianyu Li, Zhihao Qiu, Guoqing Wang<sup>†</sup>

<sup>\*</sup>Equal contribution &nbsp;&nbsp; <sup>†</sup>Corresponding author

*Under review at AAAI 2027*

[Project page](https://yuchenwu73.github.io/geobox-r1/) · [Code](https://github.com/yuchenwu73/GeoBox-R1)

</div>

## Overview

GeoBox-R1 is a 4B vision-language model for unified remote-sensing visual grounding. Given an
aerial or satellite image and a referring expression, the same model can produce either a
horizontal bounding box (HBB) or an oriented bounding box (OBB).

The model starts from Qwen3-VL-4B-Instruct and is trained in two stages:

1. **Curriculum-guided SFT** orders examples from HBB grounding to OBB grounding and then
   HBB-to-OBB chain-of-thought reasoning.
2. **Geometric RL (GDPO)** improves geometric precision with rotated-IoU and adaptive
   Wasserstein rewards, without a learned reward model.

## Results

Macro averages are shown below. Full comparisons, per-dataset results, and the evaluation
protocol are available on the [project page](https://yuchenwu73.github.io/geobox-r1/).

| Task | Evaluation sets | Acc@0.5 | Acc@0.7 | mIoU / mRIoU |
| --- | :-: | :-: | :-: | :-: |
| HBB | 7 | **58.78** | **42.22** | **50.39** |
| OBB | 3 | **47.32** | **27.55** | **39.85** |

Among the evaluated baselines, GeoBox-R1 achieves the best macro averages while using 4B
parameters. GDPO is trained only on OBB samples, but it also improves HBB performance over the
SFT stage.

## Usage

Install a recent Transformers release together with PyTorch, Pillow, and Accelerate, then run:

````python
from PIL import Image
from transformers import AutoModelForImageTextToText, AutoProcessor

model_id = "yuchenwu73/GeoBox-R1"

model = AutoModelForImageTextToText.from_pretrained(
    model_id,
    dtype="auto",
    device_map="auto",
)
processor = AutoProcessor.from_pretrained(model_id)

image = Image.open("scene.png").convert("RGB")
expression = "the brown SUV on the right"

prompt = f"""Locate the instance that matches the description: [{expression}]. Report oriented bbox coordinates in following JSON format:
```json
[
\t{{"oriented_bbox": [[x1, y1], [x2, y2], [x3, y3], [x4, y4]]}}
]
```"""

messages = [
    {
        "role": "user",
        "content": [
            {"type": "image"},
            {"type": "text", "text": prompt},
        ],
    }
]
text = processor.apply_chat_template(
    messages,
    tokenize=False,
    add_generation_prompt=True,
)
inputs = processor(text=[text], images=[image], return_tensors="pt").to(model.device)

generated = model.generate(**inputs, max_new_tokens=256)
generated = generated[:, inputs.input_ids.shape[1]:]
print(processor.batch_decode(generated, skip_special_tokens=True)[0])
````

For HBB grounding, replace the prompt with:

````python
prompt = f"""Locate the instance that matches the description: [{expression}]. Report horizontal bbox coordinates in following JSON format:
```json
[
\t{{"horizontal_bbox": [x1, y1, x2, y2]}}
]
```"""
````

Coordinates are quantized to `[0, 1000]`. Multiply x coordinates by the image width divided by
1000, and y coordinates by the image height divided by 1000, to recover pixel coordinates.

The repository also provides evaluation scripts, an interactive demo, and the complete training
pipeline: [github.com/yuchenwu73/GeoBox-R1](https://github.com/yuchenwu73/GeoBox-R1).

## License

The model weights are released under the **CC BY-NC 4.0** license. Users must also comply with
the licenses and terms of the underlying Qwen3-VL model and any input datasets they use.

## Citation

```bibtex
@misc{geoboxr1,
  title  = {GeoBox-R1: Curriculum-Guided SFT and Geometric RL for
            Unified Box-Level Remote Sensing Visual Grounding},
  author = {Lan, Chenxi and Wu, Yuchen and Zhou, Minghang and
            Li, Tianyu and Qiu, Zhihao and Wang, Guoqing},
  year   = {2026},
  url    = {https://yuchenwu73.github.io/geobox-r1/},
  note   = {Preprint}
}
```