yuchenwu73 commited on
Commit
c18ebfc
·
verified ·
1 Parent(s): a7ffa7e

docs: prepare concise public model card

Browse files

Remove links to private companion repositories, keep only the project and code links, condense results, and replace machine-local training paths with portable inference metadata.

Files changed (2) hide show
  1. README.md +82 -111
  2. args.json +5 -531
README.md CHANGED
@@ -8,90 +8,75 @@ language:
8
  tags:
9
  - remote-sensing
10
  - visual-grounding
 
11
  - oriented-bounding-box
12
  - reinforcement-learning
13
  - qwen3-vl
14
  ---
15
 
16
- **English** | [简体中文](#geobox-r1-中文)
17
 
18
  # GeoBox-R1
19
 
20
- Unified box-level remote sensing visual grounding — one model producing both horizontal (HBB)
21
- and oriented (OBB) bounding boxes.
22
 
23
- > **GeoBox-R1: Curriculum-Guided SFT and Geometric RL for Unified Box-Level Remote Sensing Visual Grounding**
24
- > Chenxi Lan\*, Yuchen Wu\*, Minghang Zhou, Tianyu Li, Zhihao Qiu, Guoqing Wang
25
- > Under review at AAAI 2027. (\* equal contribution)
26
- >
27
- > [Project page](https://yuchenwu73.github.io/GeoBox-R1/) ·
28
- > [Code](https://github.com/yuchenwu73/GeoBox-R1) ·
29
- > [Training data](https://huggingface.co/datasets/yuchenwu73/GeoBox-R1-Data) ·
30
- > [Stage-1 SFT checkpoint](https://huggingface.co/yuchenwu73/GeoBox-R1-SFT)
31
 
32
- This repository holds the **final model** — merged weights after Stage-1 SFT and Stage-2 GDPO.
33
- For the first stage alone, use [`GeoBox-R1-SFT`](https://huggingface.co/yuchenwu73/GeoBox-R1-SFT).
34
 
35
- ## Training
36
 
37
- Built on **Qwen3-VL-4B-Instruct** in two stages:
38
 
39
- 1. **Curriculum-guided SFT** — training data ordered easy-to-hard: HBB → OBB → HBB-to-OBB CoT.
40
- LoRA (rank 16, alpha 32), vision encoder and merger frozen, lr `1e-4`, 1 epoch, 2× RTX 4090.
41
- 2. **Geometric RL (GDPO)** — refines OBB prediction on top of the SFT checkpoint with two
42
- rule-based geometric rewards: Rotated IoU and an adaptive Wasserstein distance (λ = 0.5 each).
43
- G = 8 rollouts, β = 0.02, τ_c = 8, lr `5e-6`, 1 epoch, 3× A100 40G (one vLLM rollout server
44
- and two GDPO workers).
45
 
46
- ## Results
47
-
48
- Best macro averages on all three metrics, across 7 HBB and 3 OBB evaluation sets.
49
 
50
- ### HBB (7 evaluation sets, Acc@0.5)
 
 
51
 
52
- | Model | Params | DIOR-Test | DIOR-Val | RSVG-Test | RSVG-Val | GeoChat* | VRSBench* | AVVG | **Avg.** |
53
- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |
54
- | Qwen3-VL | 8B | 54.14 | 53.68 | 33.82 | 34.80 | 50.73 | 53.55 | 25.22 | 43.70 |
55
- | GeoGround | 7B | **77.70** | **77.13** | 26.65 | 27.81 | **69.76** | 65.76 | 21.60 | 52.35 |
56
- | InternVL3 (SFT) | 8B | 74.83 | 75.06 | 46.05 | 44.05 | 59.23 | 58.05 | 28.41 | 55.10 |
57
- | GeoBox-R1 (SFT) | 4B | 74.91 | 74.21 | 48.57 | 46.29 | 61.39 | 64.45 | 30.33 | 57.17 |
58
- | **GeoBox-R1** | **4B** | 76.61 | 75.11 | **51.26** | **48.38** | 61.13 | **66.84** | **32.14** | **58.78** |
59
 
60
- Acc@0.7 and mIoU macro averages are **42.22** and **50.39**, also the best.
 
 
 
61
 
62
- ### OBB (3 evaluation sets, Acc@0.5 / Acc@0.7 / mRIoU)
63
 
64
- | Model | Params | GeoChat* | VRSBench* | AVVG | **Avg.** |
65
- | --- | --- | --- | --- | --- | --- |
66
- | InternVL3 (SFT) | 8B | 49.79 / 23.10 / 41.66 | 42.01 / 20.65 / 40.68 | 15.43 / 7.40 / 15.39 | 35.74 / 17.05 / 32.58 |
67
- | GeoGround | 7B | 58.72 / 25.49 / 46.89 | 53.26 / 29.82 / 48.35 | 13.89 / 4.10 / 15.64 | 41.96 / 19.81 / 36.96 |
68
- | GeoBox-R1 (SFT) | 4B | 55.96 / 30.45 / 45.45 | 51.14 / 27.69 / 45.91 | 22.23 / 15.07 / 19.20 | 43.11 / 24.40 / 36.85 |
69
- | **GeoBox-R1** | **4B** | **60.56 / 35.19 / 48.92** | **56.61 / 30.55 / 49.43** | **24.79 / 16.89 / 21.18** | **47.32 / 27.55 / 39.85** |
70
 
71
- At 4B parameters this beats the 7B–8B state of the art by **3.68/4.89/2.61** points on HBB and
72
- **5.36/7.74/2.89** on OBB, with the largest gains at the stricter Acc@0.7 threshold.
 
 
73
 
74
- **OBB-only RL does not cost HBB accuracy.** GDPO trains on OBB samples alone, yet the HBB macro
75
- average rises from 57.17/40.33/48.82 to 58.78/42.22/50.39 — Acc@0.5 improves on 6 of 7 sets,
76
- and Acc@0.7 and mIoU improve on all 7.
77
 
78
  ## Usage
79
 
80
- The prompts below are **byte-for-byte identical to the ones used in training and evaluation**,
81
- including the spaces inside the coordinate lists. Changing the spacing changes tokenization.
82
 
83
  ````python
84
- from transformers import AutoModelForImageTextToText, AutoProcessor
85
  from PIL import Image
 
86
 
87
  model_id = "yuchenwu73/GeoBox-R1"
88
- model = AutoModelForImageTextToText.from_pretrained(model_id, dtype="auto", device_map="auto")
 
 
 
 
 
89
  processor = AutoProcessor.from_pretrained(model_id)
90
 
91
- image = Image.open("scene.png")
92
- expression = "the brown suv on the right"
93
 
94
- # Oriented box (OBB)
95
  prompt = f"""Locate the instance that matches the description: [{expression}]. Report oriented bbox coordinates in following JSON format:
96
  ```json
97
  [
@@ -99,14 +84,28 @@ prompt = f"""Locate the instance that matches the description: [{expression}]. R
99
  ]
100
  ```"""
101
 
102
- messages = [{"role": "user", "content": [{"type": "image"}, {"type": "text", "text": prompt}]}]
103
- text = processor.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
 
 
 
 
 
 
 
 
 
 
 
 
104
  inputs = processor(text=[text], images=[image], return_tensors="pt").to(model.device)
105
- out = model.generate(**inputs, max_new_tokens=256)
106
- print(processor.decode(out[0][inputs.input_ids.shape[1]:], skip_special_tokens=True))
 
 
107
  ````
108
 
109
- For horizontal boxes, use the same call with:
110
 
111
  ````python
112
  prompt = f"""Locate the instance that matches the description: [{expression}]. Report horizontal bbox coordinates in following JSON format:
@@ -117,64 +116,36 @@ prompt = f"""Locate the instance that matches the description: [{expression}]. R
117
  ```"""
118
  ````
119
 
120
- Coordinates are quantized to `[0, 1000]`; scale by image width and height to recover pixels.
121
-
122
- The model can also be served with [ms-swift](https://github.com/modelscope/ms-swift), the
123
- framework used for training.
124
-
125
- ## Citation
126
-
127
- ```bibtex
128
- @inproceedings{geoboxr1,
129
- title = {GeoBox-R1: Curriculum-Guided SFT and Geometric RL for
130
- Unified Box-Level Remote Sensing Visual Grounding},
131
- author = {Lan, Chenxi and Wu, Yuchen and Zhou, Minghang and
132
- Li, Tianyu and Qiu, Zhihao and Wang, Guoqing},
133
- booktitle = {Proceedings of the AAAI Conference on Artificial Intelligence},
134
- year = {2027},
135
- note = {Under review}
136
- }
137
- ```
138
-
139
- ---
140
-
141
- # GeoBox-R1 中文
142
-
143
- [English](#geobox-r1) | **简体中文**
144
-
145
- 统一边界框级遥感视觉定位模型,同时输出水平框(HBB)与旋转框(OBB)。
146
-
147
- 本仓库是**最终模型**(Stage-1 SFT + Stage-2 GDPO 后的合并权重)。
148
- 只需要第一阶段结果请用 [`GeoBox-R1-SFT`](https://huggingface.co/yuchenwu73/GeoBox-R1-SFT)。
149
 
150
- ## 训练流程
 
151
 
152
- 基座 **Qwen3-VL-4B-Instruct**,两阶段:
153
 
154
- 1. **课程式 SFT** — 训练数据按 HBB → OBB → HBB-to-OBB CoT 由易到难排列。
155
- LoRA(rank 16,alpha 32),冻结视觉编码器与 merger,lr `1e-4`,1 epoch,2× RTX 4090。
156
- 2. **几何强化学习(GDPO)** — 在 SFT 检查点上用两个基于规则的几何奖励细化 OBB:
157
- Rotated IoU 与自适应 Wasserstein 距离(λ 各 0.5)。
158
- G=8 rollouts,β=0.02,τ_c=8,lr `5e-6`,1 epoch,3× A100 40G(1 个 vLLM rollout 服务 + 2 个 GDPO worker)。
 
159
 
160
- ## 结果
161
 
162
- 7 个 HBB 与 3 个 OBB 评测集,三项指标的宏平均均取得最佳:
 
163
 
164
- | 任务 | Acc@0.5 | Acc@0.7 | mIoU / mRIoU |
165
- | --- | --- | --- | --- |
166
- | HBB(7 个集) | **58.78** | **42.22** | **50.39** |
167
- | OBB(3 个集) | **47.32** | **27.55** | **39.85** |
168
-
169
- 以 4B 参数超过 7B–8B 的现有最优模型:HBB 领先 **3.68/4.89/2.61** 点,
170
- OBB 领先 **5.36/7.74/2.89** 点,在更严格的 Acc@0.7 上增益最大。
171
- 逐数据集的完整结果见上方英文表格或[项目主页](https://yuchenwu73.github.io/GeoBox-R1/)。
172
-
173
- **OBB-only RL 不牺牲 HBB**:GDPO 只用 OBB 样本训练,HBB 宏平均反而从 57.17/40.33/48.82
174
- 升到 58.78/42.22/50.39 —— Acc@0.5 在 7 个集合中的 6 个上升,Acc@0.7 与 mIoU 全部 7 个上升。
175
-
176
- ## 使用
177
 
178
- 代码见上方英文 [Usage](#usage) 一节。提示词与训练、评测时**逐字节一致**,
179
- 包括坐标列表里的空格 —— 改动空格会改变分词结果。坐标量化到 `[0, 1000]`,
180
- 按图像宽高缩放即可还原到像素。
 
 
 
 
 
 
 
 
 
8
  tags:
9
  - remote-sensing
10
  - visual-grounding
11
+ - horizontal-bounding-box
12
  - oriented-bounding-box
13
  - reinforcement-learning
14
  - qwen3-vl
15
  ---
16
 
17
+ <div align="center">
18
 
19
  # GeoBox-R1
20
 
21
+ **Curriculum-Guided SFT and Geometric RL for Unified Box-Level Remote Sensing Visual Grounding**
 
22
 
23
+ Chenxi Lan\*, Yuchen Wu\*, Minghang Zhou, Tianyu Li, Zhihao Qiu, Guoqing Wang<sup>†</sup>
 
 
 
 
 
 
 
24
 
25
+ <sup>\*</sup>Equal contribution &nbsp;&nbsp; <sup>†</sup>Corresponding author
 
26
 
27
+ *Under review at AAAI 2027*
28
 
29
+ [Project page](https://yuchenwu73.github.io/geobox-r1/) · [Code](https://github.com/yuchenwu73/GeoBox-R1)
30
 
31
+ </div>
 
 
 
 
 
32
 
33
+ ## Overview
 
 
34
 
35
+ GeoBox-R1 is a 4B vision-language model for unified remote-sensing visual grounding. Given an
36
+ aerial or satellite image and a referring expression, the same model can produce either a
37
+ horizontal bounding box (HBB) or an oriented bounding box (OBB).
38
 
39
+ The model starts from Qwen3-VL-4B-Instruct and is trained in two stages:
 
 
 
 
 
 
40
 
41
+ 1. **Curriculum-guided SFT** orders examples from HBB grounding to OBB grounding and then
42
+ HBB-to-OBB chain-of-thought reasoning.
43
+ 2. **Geometric RL (GDPO)** improves geometric precision with rotated-IoU and adaptive
44
+ Wasserstein rewards, without a learned reward model.
45
 
46
+ ## Results
47
 
48
+ Macro averages are shown below. Full comparisons, per-dataset results, and the evaluation
49
+ protocol are available on the [project page](https://yuchenwu73.github.io/geobox-r1/).
 
 
 
 
50
 
51
+ | Task | Evaluation sets | Acc@0.5 | Acc@0.7 | mIoU / mRIoU |
52
+ | --- | :-: | :-: | :-: | :-: |
53
+ | HBB | 7 | **58.78** | **42.22** | **50.39** |
54
+ | OBB | 3 | **47.32** | **27.55** | **39.85** |
55
 
56
+ Among the evaluated baselines, GeoBox-R1 achieves the best macro averages while using 4B
57
+ parameters. GDPO is trained only on OBB samples, but it also improves HBB performance over the
58
+ SFT stage.
59
 
60
  ## Usage
61
 
62
+ Install a recent Transformers release together with PyTorch, Pillow, and Accelerate, then run:
 
63
 
64
  ````python
 
65
  from PIL import Image
66
+ from transformers import AutoModelForImageTextToText, AutoProcessor
67
 
68
  model_id = "yuchenwu73/GeoBox-R1"
69
+
70
+ model = AutoModelForImageTextToText.from_pretrained(
71
+ model_id,
72
+ dtype="auto",
73
+ device_map="auto",
74
+ )
75
  processor = AutoProcessor.from_pretrained(model_id)
76
 
77
+ image = Image.open("scene.png").convert("RGB")
78
+ expression = "the brown SUV on the right"
79
 
 
80
  prompt = f"""Locate the instance that matches the description: [{expression}]. Report oriented bbox coordinates in following JSON format:
81
  ```json
82
  [
 
84
  ]
85
  ```"""
86
 
87
+ messages = [
88
+ {
89
+ "role": "user",
90
+ "content": [
91
+ {"type": "image"},
92
+ {"type": "text", "text": prompt},
93
+ ],
94
+ }
95
+ ]
96
+ text = processor.apply_chat_template(
97
+ messages,
98
+ tokenize=False,
99
+ add_generation_prompt=True,
100
+ )
101
  inputs = processor(text=[text], images=[image], return_tensors="pt").to(model.device)
102
+
103
+ generated = model.generate(**inputs, max_new_tokens=256)
104
+ generated = generated[:, inputs.input_ids.shape[1]:]
105
+ print(processor.batch_decode(generated, skip_special_tokens=True)[0])
106
  ````
107
 
108
+ For HBB grounding, replace the prompt with:
109
 
110
  ````python
111
  prompt = f"""Locate the instance that matches the description: [{expression}]. Report horizontal bbox coordinates in following JSON format:
 
116
  ```"""
117
  ````
118
 
119
+ Coordinates are quantized to `[0, 1000]`. Multiply x coordinates by the image width divided by
120
+ 1000, and y coordinates by the image height divided by 1000, to recover pixel coordinates.
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
121
 
122
+ The repository also provides evaluation scripts, an interactive demo, and the complete training
123
+ pipeline: [github.com/yuchenwu73/GeoBox-R1](https://github.com/yuchenwu73/GeoBox-R1).
124
 
125
+ ## Limitations
126
 
127
+ - The model is designed for single-object visual grounding in remote-sensing imagery; it is not
128
+ a general-purpose detector.
129
+ - Predictions are generated as text and may occasionally be malformed or refer to the wrong
130
+ object, especially for ambiguous expressions or very small targets.
131
+ - Reported results follow the benchmark splits and protocols described on the project page and
132
+ may not transfer directly to other imagery domains.
133
 
134
+ ## License
135
 
136
+ The model weights are released under the **CC BY-NC 4.0** license. Users must also comply with
137
+ the licenses and terms of the underlying Qwen3-VL model and any input datasets they use.
138
 
139
+ ## Citation
 
 
 
 
 
 
 
 
 
 
 
 
140
 
141
+ ```bibtex
142
+ @misc{geoboxr1,
143
+ title = {GeoBox-R1: Curriculum-Guided SFT and Geometric RL for
144
+ Unified Box-Level Remote Sensing Visual Grounding},
145
+ author = {Lan, Chenxi and Wu, Yuchen and Zhou, Minghang and
146
+ Li, Tianyu and Qiu, Zhihao and Wang, Guoqing},
147
+ year = {2026},
148
+ url = {https://yuchenwu73.github.io/geobox-r1/},
149
+ note = {Preprint}
150
+ }
151
+ ```
args.json CHANGED
@@ -1,535 +1,9 @@
1
  {
2
- "output_dir": "/data2/longfeiqi/sutando/MLLM4RSVG/output/GRPO/v64-20260329-180422",
3
- "overwrite_output_dir": false,
4
- "do_train": false,
5
- "do_eval": false,
6
- "do_predict": false,
7
- "eval_strategy": "no",
8
- "prediction_loss_only": false,
9
- "per_device_train_batch_size": 8,
10
- "per_device_eval_batch_size": 1,
11
- "per_gpu_train_batch_size": null,
12
- "per_gpu_eval_batch_size": null,
13
- "gradient_accumulation_steps": 4,
14
- "eval_accumulation_steps": null,
15
- "eval_delay": 0,
16
- "torch_empty_cache_steps": null,
17
- "learning_rate": 5e-06,
18
- "weight_decay": 0.1,
19
- "adam_beta1": 0.9,
20
- "adam_beta2": 0.95,
21
- "adam_epsilon": 1e-08,
22
- "max_grad_norm": 1.0,
23
- "num_train_epochs": 1.0,
24
- "max_steps": -1,
25
- "lr_scheduler_type": "cosine",
26
- "lr_scheduler_kwargs": null,
27
- "warmup_ratio": 0.05,
28
- "warmup_steps": 0,
29
- "log_level": "passive",
30
- "log_level_replica": "warning",
31
- "log_on_each_node": true,
32
- "logging_dir": "/data2/longfeiqi/sutando/MLLM4RSVG/output/GRPO/v64-20260329-180422/runs",
33
- "logging_strategy": "steps",
34
- "logging_first_step": true,
35
- "logging_steps": 1,
36
- "logging_nan_inf_filter": true,
37
- "save_strategy": "steps",
38
- "save_steps": 100.0,
39
- "save_total_limit": 3,
40
- "save_safetensors": true,
41
- "save_on_each_node": false,
42
- "save_only_model": false,
43
- "restore_callback_states_from_checkpoint": false,
44
- "no_cuda": false,
45
- "use_cpu": false,
46
- "use_mps_device": false,
47
- "seed": 42,
48
- "data_seed": 42,
49
- "jit_mode_eval": false,
50
- "bf16": true,
51
- "fp16": false,
52
- "fp16_opt_level": "O1",
53
- "half_precision_backend": "auto",
54
- "bf16_full_eval": false,
55
- "fp16_full_eval": false,
56
- "tf32": null,
57
- "local_rank": 0,
58
- "ddp_backend": null,
59
- "tpu_num_cores": null,
60
- "tpu_metrics_debug": false,
61
- "debug": null,
62
- "dataloader_drop_last": false,
63
- "eval_steps": 100.0,
64
- "dataloader_num_workers": 4,
65
- "dataloader_prefetch_factor": null,
66
- "past_index": -1,
67
- "run_name": "/data2/longfeiqi/sutando/MLLM4RSVG/output/GRPO/v64-20260329-180422",
68
- "disable_tqdm": null,
69
- "remove_unused_columns": false,
70
- "label_names": null,
71
- "load_best_model_at_end": false,
72
- "metric_for_best_model": "loss",
73
- "greater_is_better": false,
74
- "ignore_data_skip": false,
75
- "fsdp": [],
76
- "fsdp_min_num_params": 0,
77
- "fsdp_config": null,
78
- "fsdp_transformer_layer_cls_to_wrap": null,
79
- "accelerator_config": {
80
- "dispatch_batches": false
81
- },
82
- "parallelism_config": null,
83
- "deepspeed": {
84
- "fp16": {
85
- "enabled": "auto",
86
- "loss_scale": 0,
87
- "loss_scale_window": 1000,
88
- "initial_scale_power": 16,
89
- "hysteresis": 2,
90
- "min_loss_scale": 1
91
- },
92
- "bf16": {
93
- "enabled": "auto"
94
- },
95
- "zero_optimization": {
96
- "stage": 2,
97
- "offload_optimizer": {
98
- "device": "none",
99
- "pin_memory": true
100
- },
101
- "allgather_partitions": true,
102
- "allgather_bucket_size": 200000000.0,
103
- "overlap_comm": false,
104
- "reduce_scatter": true,
105
- "reduce_bucket_size": 200000000.0,
106
- "contiguous_gradients": true
107
- },
108
- "gradient_accumulation_steps": "auto",
109
- "gradient_clipping": "auto",
110
- "steps_per_print": 2000,
111
- "train_batch_size": "auto",
112
- "train_micro_batch_size_per_gpu": "auto",
113
- "wall_clock_breakdown": false
114
- },
115
- "label_smoothing_factor": 0.0,
116
- "optim": "adamw_torch_fused",
117
- "optim_args": null,
118
- "adafactor": false,
119
- "group_by_length": false,
120
- "length_column_name": "length",
121
- "report_to": [
122
- "swanlab"
123
- ],
124
- "project": "huggingface",
125
- "trackio_space_id": "trackio",
126
- "ddp_find_unused_parameters": null,
127
- "ddp_bucket_cap_mb": null,
128
- "ddp_broadcast_buffers": null,
129
- "dataloader_pin_memory": true,
130
- "dataloader_persistent_workers": false,
131
- "skip_memory_metrics": true,
132
- "use_legacy_prediction_loop": false,
133
- "push_to_hub": false,
134
- "resume_from_checkpoint": null,
135
- "hub_model_id": null,
136
- "hub_strategy": "every_save",
137
- "hub_token": null,
138
- "hub_private_repo": null,
139
- "hub_always_push": false,
140
- "hub_revision": null,
141
- "gradient_checkpointing": true,
142
- "gradient_checkpointing_kwargs": null,
143
- "include_inputs_for_metrics": false,
144
- "include_for_metrics": [],
145
- "eval_do_concat_batches": true,
146
- "fp16_backend": "auto",
147
- "push_to_hub_model_id": null,
148
- "push_to_hub_organization": null,
149
- "push_to_hub_token": null,
150
- "mp_parameters": "",
151
- "auto_find_batch_size": false,
152
- "full_determinism": false,
153
- "torchdynamo": null,
154
- "ray_scope": "last",
155
- "ddp_timeout": 18000000,
156
- "torch_compile": false,
157
- "torch_compile_backend": null,
158
- "torch_compile_mode": null,
159
- "include_tokens_per_second": false,
160
- "include_num_input_tokens_seen": false,
161
- "neftune_noise_alpha": null,
162
- "optim_target_modules": null,
163
- "batch_eval_metrics": false,
164
- "eval_on_start": false,
165
- "use_liger_kernel": false,
166
- "liger_kernel_config": null,
167
- "eval_use_gather_object": false,
168
- "average_tokens_across_devices": true,
169
- "sortish_sampler": false,
170
- "predict_with_generate": false,
171
- "generation_max_length": null,
172
- "generation_num_beams": null,
173
- "generation_config": null,
174
- "tuner_backend": "peft",
175
- "vit_gradient_checkpointing": null,
176
- "router_aux_loss_coef": 0.0,
177
- "enable_dft_loss": false,
178
- "enable_channel_loss": false,
179
- "check_model": true,
180
- "acc_strategy": "token",
181
- "train_dataloader_shuffle": true,
182
- "max_epochs": null,
183
- "aligner_lr": null,
184
- "vit_lr": null,
185
- "use_logits_to_keep": null,
186
- "ds3_gather_for_generation": true,
187
- "resume_only_model": false,
188
- "optimizer": null,
189
- "loss_type": "grpo",
190
- "eval_metric": null,
191
- "callbacks": [],
192
- "early_stop_interval": null,
193
- "eval_use_evalscope": false,
194
- "eval_dataset": [],
195
- "eval_dataset_args": null,
196
- "eval_limit": null,
197
- "eval_generation_config": null,
198
- "extra_eval_args": null,
199
- "tuner_type": "lora",
200
- "use_galore": false,
201
- "galore_target_modules": null,
202
- "galore_rank": 128,
203
- "galore_update_proj_gap": 50,
204
- "galore_scale": 1.0,
205
- "galore_proj_type": "std",
206
- "galore_optim_per_parameter": false,
207
- "galore_with_embedding": false,
208
- "galore_quantization": false,
209
- "galore_proj_quant": false,
210
- "galore_proj_bits": 4,
211
- "galore_proj_group_size": 256,
212
- "galore_cos_threshold": 0.4,
213
- "galore_gamma_proj": 2,
214
- "galore_queue_size": 5,
215
- "lisa_activated_layers": 0,
216
- "lisa_step_interval": 20,
217
- "use_flash_ckpt": false,
218
- "use_ray": false,
219
- "ray_exp_name": null,
220
- "device_groups": null,
221
- "model": "../checkpoints/checkpoint-sft_no-joint_cot",
222
  "model_type": "qwen3_vl",
223
- "model_revision": null,
224
- "task_type": "causal_lm",
225
- "torch_dtype": "bfloat16",
226
- "attn_impl": "flash_attention_2",
227
- "experts_impl": null,
228
- "new_special_tokens": [],
229
- "num_labels": null,
230
- "problem_type": null,
231
- "rope_scaling": null,
232
- "device_map": null,
233
- "max_memory": {},
234
- "max_model_len": null,
235
- "local_repo_path": null,
236
- "init_strategy": null,
237
  "template": "qwen3_vl",
238
- "system": null,
239
- "max_length": 2048,
240
- "truncation_strategy": "left",
241
- "max_pixels": null,
242
- "agent_template": null,
243
- "norm_bbox": null,
244
- "use_chat_template": true,
245
- "padding_side": "right",
246
- "padding_free": false,
247
- "loss_scale": "last_round",
248
- "sequence_parallel_size": 1,
249
- "template_backend": "swift",
250
- "response_prefix": null,
251
- "enable_thinking": null,
252
- "add_non_thinking_prefix": true,
253
- "dataset": [
254
- "../refGeo*/RL/rl_obb_train_20.0%.jsonl"
255
- ],
256
- "val_dataset": [],
257
- "cached_dataset": [],
258
- "cached_val_dataset": [],
259
- "split_dataset_ratio": 0.0,
260
- "dataset_num_proc": 1,
261
- "load_from_cache_file": true,
262
- "dataset_shuffle": true,
263
- "val_dataset_shuffle": false,
264
- "streaming": false,
265
- "interleave_prob": null,
266
- "stopping_strategy": "first_exhausted",
267
- "shuffle_buffer_size": 1000,
268
- "download_mode": "reuse_dataset_if_exists",
269
- "columns": {},
270
- "strict": false,
271
- "model_name": null,
272
- "model_author": null,
273
- "custom_dataset_info": [],
274
- "quant_method": null,
275
- "quant_bits": null,
276
- "hqq_axis": null,
277
- "bnb_4bit_compute_dtype": "bfloat16",
278
- "bnb_4bit_quant_type": "nf4",
279
- "bnb_4bit_use_double_quant": true,
280
- "bnb_4bit_quant_storage": null,
281
- "max_new_tokens": 256,
282
- "temperature": 0.9,
283
- "top_k": 50,
284
- "top_p": 0.9,
285
- "repetition_penalty": 1.0,
286
- "num_beams": 1,
287
- "stream": false,
288
- "stop_words": [],
289
- "logprobs": false,
290
- "top_logprobs": null,
291
- "structured_outputs_regex": null,
292
- "train_type": null,
293
- "adapters": [],
294
- "external_plugins": [
295
- "../rl_func_plugin.py"
296
- ],
297
- "custom_register_path": [],
298
- "model_kwargs": {},
299
- "load_args": false,
300
- "load_data_args": false,
301
- "packing": false,
302
- "packing_length": null,
303
- "packing_num_proc": 1,
304
- "lazy_tokenize": true,
305
- "use_hf": false,
306
- "ignore_args_error": false,
307
- "use_swift_lora": false,
308
- "freeze_parameters": [],
309
- "freeze_parameters_regex": null,
310
- "freeze_parameters_ratio": 0.0,
311
- "trainable_parameters": [],
312
- "trainable_parameters_regex": null,
313
- "freeze_llm": false,
314
- "freeze_vit": true,
315
- "freeze_aligner": true,
316
- "target_modules": [
317
- "all-linear"
318
- ],
319
- "target_regex": null,
320
- "target_parameters": null,
321
- "modules_to_save": [],
322
- "lora_rank": 16,
323
- "lora_alpha": 32,
324
- "lora_dropout": 0.05,
325
- "lora_bias": "none",
326
- "lora_dtype": null,
327
- "lorap_lr_ratio": null,
328
- "use_rslora": false,
329
- "use_dora": false,
330
- "lora_ga_batch_size": 2,
331
- "lora_ga_iters": 2,
332
- "lora_ga_max_length": 1024,
333
- "lora_ga_direction": "ArB2r",
334
- "lora_ga_scale": "stable",
335
- "lora_ga_stable_gamma": 16,
336
- "init_weights": true,
337
- "fourier_n_frequency": 2000,
338
- "fourier_scaling": 300.0,
339
- "boft_block_size": 4,
340
- "boft_block_num": 0,
341
- "boft_n_butterfly_factor": 1,
342
- "boft_dropout": 0.0,
343
- "vera_rank": 256,
344
- "vera_projection_prng_key": 0,
345
- "vera_dropout": 0.0,
346
- "vera_d_initial": 0.1,
347
- "adapter_act": "gelu",
348
- "adapter_length": 128,
349
- "adalora_target_r": 8,
350
- "adalora_init_r": 12,
351
- "adalora_tinit": 0,
352
- "adalora_tfinal": 0,
353
- "adalora_deltaT": 1,
354
- "adalora_beta1": 0.85,
355
- "adalora_beta2": 0.85,
356
- "adalora_orth_reg_weight": 0.5,
357
- "llamapro_num_new_blocks": 4,
358
- "llamapro_num_groups": null,
359
- "reft_layer_key": null,
360
- "reft_layers": null,
361
- "reft_rank": 4,
362
- "reft_intervention_type": "LoreftIntervention",
363
- "reft_args": null,
364
- "swanlab_token": null,
365
- "swanlab_project": "RL",
366
- "swanlab_workspace": null,
367
- "swanlab_exp_name": "GRPO@[20%Data IoU(0.5) + Adaptive_WD(0.5) GDPO tau=8]",
368
- "swanlab_notification_method": null,
369
- "swanlab_webhook_url": null,
370
- "swanlab_secret": null,
371
- "swanlab_sender_email": null,
372
- "swanlab_receiver_email": null,
373
- "swanlab_smtp_server": null,
374
- "swanlab_smtp_port": null,
375
- "swanlab_email_language": "zh",
376
- "swanlab_mode": "cloud",
377
- "add_version": true,
378
- "create_checkpoint_symlink": false,
379
- "zero_hpz_partition_size": null,
380
- "deepspeed_autotp_size": null,
381
- "reward_model": null,
382
- "reward_adapters": [],
383
- "reward_model_type": null,
384
- "reward_model_revision": null,
385
- "num_ppo_epochs": 4,
386
- "whiten_rewards": false,
387
- "kl_coef": 0.05,
388
- "cliprange": 0.2,
389
- "vf_coef": 0.1,
390
- "cliprange_value": 0.2,
391
- "gamma": 1.0,
392
- "lam": 0.95,
393
- "num_mini_batches": 1,
394
- "local_rollout_forward_batch_size": 64,
395
- "num_sample_generations": 10,
396
- "response_length": 256,
397
- "missing_eos_penalty": null,
398
- "vllm_gpu_memory_utilization": 0.9,
399
- "vllm_tensor_parallel_size": 1,
400
- "vllm_pipeline_parallel_size": 1,
401
- "vllm_enable_expert_parallel": false,
402
- "vllm_max_num_seqs": null,
403
- "vllm_max_model_len": null,
404
- "vllm_disable_custom_all_reduce": true,
405
- "vllm_enforce_eager": false,
406
- "vllm_limit_mm_per_prompt": null,
407
- "vllm_max_lora_rank": 16,
408
- "vllm_enable_prefix_caching": true,
409
- "vllm_use_async_engine": null,
410
- "vllm_quantization": null,
411
- "vllm_reasoning_parser": null,
412
- "vllm_disable_cascade_attn": false,
413
- "vllm_mm_processor_cache_gb": null,
414
- "vllm_speculative_config": null,
415
- "vllm_engine_kwargs": {},
416
- "vllm_data_parallel_size": 1,
417
- "use_vllm": true,
418
- "vllm_mode": "server",
419
- "vllm_enable_lora": false,
420
- "vllm_server_base_url": null,
421
- "vllm_server_host": [
422
- "127.0.0.1"
423
- ],
424
- "vllm_server_port": [
425
- 8897
426
- ],
427
- "vllm_server_timeout": 240.0,
428
- "vllm_server_group_port": [
429
- 51226
430
- ],
431
- "enable_flattened_weight_sync": true,
432
- "async_generate": false,
433
- "sleep_level": 0,
434
- "move_model_batches": null,
435
- "offload_optimizer": false,
436
- "offload_model": false,
437
- "wandb_log_unique_prompts": null,
438
- "epsilon": 0.2,
439
- "epsilon_high": null,
440
- "delta": null,
441
- "cosine_min_len_value_wrong": -0.5,
442
- "cosine_max_len_value_wrong": 0.0,
443
- "cosine_min_len_value_correct": 1.0,
444
- "cosine_max_len_value_correct": 0.5,
445
- "cosine_max_len": null,
446
- "repetition_n_grams": 3,
447
- "repetition_max_penalty": -1.0,
448
- "reward_model_plugin": null,
449
- "chord_sft_dataset": [],
450
- "chord_sft_per_device_train_batch_size": null,
451
- "chord_enable_phi_function": false,
452
- "chord_mu_warmup_steps": null,
453
- "chord_mu_decay_steps": null,
454
- "chord_mu_peak": null,
455
- "chord_mu_valley": null,
456
- "sync_ref_model": false,
457
- "ref_model_sync_steps": 512,
458
- "ref_model_mixup_alpha": 0.6,
459
- "multi_turn_scheduler": null,
460
- "max_turns": null,
461
- "completion_length_limit_scope": "per_round",
462
- "vllm_server_pass_dataset": false,
463
- "dynamic_sample": false,
464
- "max_resample_times": 3,
465
- "overlong_filter": false,
466
- "soft_max_length": null,
467
- "soft_cache_length": null,
468
- "scale_rewards": "gdpo",
469
- "log_entropy": false,
470
- "top_entropy_quantile": 1.0,
471
- "importance_sampling_level": "token",
472
- "tau_pos": 1.0,
473
- "tau_neg": 1.05,
474
- "advantage_estimator": "grpo",
475
- "kl_in_reward": false,
476
- "generation_batch_size": null,
477
- "steps_per_generation": null,
478
- "num_generations_eval": null,
479
- "rollout_importance_sampling_mode": null,
480
- "rollout_importance_sampling_threshold": 2.0,
481
- "log_rollout_offpolicy_metrics": false,
482
- "off_policy_sequence_mask_delta": null,
483
- "num_generations": 8,
484
- "reward_funcs": [
485
- "external_vg_iou",
486
- "external_vg_wd_adaptive"
487
- ],
488
- "reward_weights": [
489
- 0.5,
490
- 0.5
491
- ],
492
- "log_completions": true,
493
- "num_iterations": 1,
494
- "teacher_model": null,
495
- "teacher_adapters": [],
496
- "teacher_model_type": null,
497
- "teacher_model_revision": null,
498
- "teacher_deepspeed": null,
499
- "rlhf_type": "grpo",
500
- "ref_model": null,
501
- "ref_adapters": [],
502
- "ref_model_type": null,
503
- "ref_model_revision": null,
504
- "beta": 0.02,
505
- "label_smoothing": 0,
506
- "max_completion_length": 256,
507
- "rpo_alpha": null,
508
- "ld_alpha": null,
509
- "discopop_tau": 0.05,
510
- "loss_weights": null,
511
- "cpo_alpha": 1.0,
512
- "simpo_gamma": 1,
513
- "desirable_weight": 1.0,
514
- "undesirable_weight": 1.0,
515
- "center_rewards_coefficient": null,
516
- "sft_alpha": 0,
517
- "lmbda": 0.5,
518
- "seq_kd": false,
519
- "offload_teacher_model": false,
520
- "vllm_client": "<swift.rlhf_trainers.vllm_client.VLLMClient object at 0x7498522122c0>",
521
- "swift_version": "4.0.0.dev0",
522
- "ckpt_dir": "../checkpoints/checkpoint-sft_no-joint_cot",
523
- "rank": 0,
524
- "global_world_size": 4,
525
- "local_world_size": 4,
526
- "model_suffix": "checkpoint-sft_no-joint_cot",
527
- "model_info": "ModelInfo(model_type='qwen3_vl', model_dir='/data2/longfeiqi/sutando/MLLM4RSVG/checkpoints/checkpoint-sft_no-joint_cot', torch_dtype=torch.bfloat16, max_model_len=262144, quant_method=None, quant_bits=None, rope_scaling={'mrope_interleaved': True, 'mrope_section': [24, 20, 20], 'rope_type': 'default'}, is_moe_model=False, is_multimodal=True, config=None, task_type='causal_lm', num_labels=None)",
528
- "model_meta": "ModelMeta(model_type='qwen3_vl', model_groups=[ModelGroup(models=[Model(ms_model_id='Qwen/Qwen3-VL-2B-Instruct', hf_model_id='Qwen/Qwen3-VL-2B-Instruct', model_path=None, ms_revision=None, hf_revision=None), Model(ms_model_id='Qwen/Qwen3-VL-2B-Thinking', hf_model_id='Qwen/Qwen3-VL-2B-Thinking', model_path=None, ms_revision=None, hf_revision=None), Model(ms_model_id='Qwen/Qwen3-VL-2B-Instruct-FP8', hf_model_id='Qwen/Qwen3-VL-2B-Instruct-FP8', model_path=None, ms_revision=None, hf_revision=None), Model(ms_model_id='Qwen/Qwen3-VL-2B-Thinking-FP8', hf_model_id='Qwen/Qwen3-VL-2B-Thinking-FP8', model_path=None, ms_revision=None, hf_revision=None), Model(ms_model_id='Qwen/Qwen3-VL-4B-Instruct', hf_model_id='Qwen/Qwen3-VL-4B-Instruct', model_path=None, ms_revision=None, hf_revision=None), Model(ms_model_id='Qwen/Qwen3-VL-4B-Thinking', hf_model_id='Qwen/Qwen3-VL-4B-Thinking', model_path=None, ms_revision=None, hf_revision=None), Model(ms_model_id='Qwen/Qwen3-VL-4B-Instruct-FP8', hf_model_id='Qwen/Qwen3-VL-4B-Instruct-FP8', model_path=None, ms_revision=None, hf_revision=None), Model(ms_model_id='Qwen/Qwen3-VL-4B-Thinking-FP8', hf_model_id='Qwen/Qwen3-VL-4B-Thinking-FP8', model_path=None, ms_revision=None, hf_revision=None), Model(ms_model_id='Qwen/Qwen3-VL-8B-Instruct', hf_model_id='Qwen/Qwen3-VL-8B-Instruct', model_path=None, ms_revision=None, hf_revision=None), Model(ms_model_id='Qwen/Qwen3-VL-8B-Thinking', hf_model_id='Qwen/Qwen3-VL-8B-Thinking', model_path=None, ms_revision=None, hf_revision=None), Model(ms_model_id='Qwen/Qwen3-VL-8B-Instruct-FP8', hf_model_id='Qwen/Qwen3-VL-8B-Instruct-FP8', model_path=None, ms_revision=None, hf_revision=None), Model(ms_model_id='Qwen/Qwen3-VL-8B-Thinking-FP8', hf_model_id='Qwen/Qwen3-VL-8B-Thinking-FP8', model_path=None, ms_revision=None, hf_revision=None), Model(ms_model_id='Qwen/Qwen3-VL-32B-Instruct', hf_model_id='Qwen/Qwen3-VL-32B-Instruct', model_path=None, ms_revision=None, hf_revision=None), Model(ms_model_id='Qwen/Qwen3-VL-32B-Thinking', hf_model_id='Qwen/Qwen3-VL-32B-Thinking', model_path=None, ms_revision=None, hf_revision=None), Model(ms_model_id='Qwen/Qwen3-VL-32B-Instruct-FP8', hf_model_id='Qwen/Qwen3-VL-32B-Instruct-FP8', model_path=None, ms_revision=None, hf_revision=None), Model(ms_model_id='Qwen/Qwen3-VL-32B-Thinking-FP8', hf_model_id='Qwen/Qwen3-VL-32B-Thinking-FP8', model_path=None, ms_revision=None, hf_revision=None)], template='qwen3_vl', ignore_patterns=None, requires=None, tags=[])], loader=<class 'swift.model.models.qwen.Qwen3VLLoader'>, template=None, model_arch=MultiModelKeys(arch_name='qwen3_vl', embedding=None, module_list=None, lm_head=None, q_proj=None, k_proj=None, v_proj=None, o_proj=None, attention=None, mlp=None, down_proj=None, qkv_proj=None, qk_proj=None, qa_proj=None, qb_proj=None, kv_proj=None, kva_proj=None, kvb_proj=None, language_model=['model.language_model', 'lm_head'], aligner=['model.visual.merger', 'model.visual.deepstack_merger_list'], vision_tower=['model.visual'], generator=[]), architectures=['Qwen3VLForConditionalGeneration'], additional_saved_files=[], torch_dtype=None, is_multimodal=True, is_reward=False, task_type=None, ignore_patterns=None, requires=['transformers>=4.57', 'qwen_vl_utils>=0.0.14', 'decord'], tags=['vision', 'video'])",
529
- "model_dir": "/data2/longfeiqi/sutando/MLLM4RSVG/checkpoints/checkpoint-sft_no-joint_cot",
530
- "template_meta": "QwenTemplateMeta(template_type='qwen3_vl', prefix=[], prompt=['<|im_start|>user\\n{{QUERY}}<|im_end|>\\n<|im_start|>assistant\\n'], chat_sep=['<|im_end|>\\n'], suffix=['<|im_end|>\\n'], template_cls=<class 'swift.template.templates.qwen.Qwen3VLTemplate'>, system_prefix=['<|im_start|>system\\n{{SYSTEM}}<|im_end|>\\n'], default_system=None, auto_add_bos=False, stop_words=['<|endoftext|>'], agent_template='hermes', is_thinking=False, thinking_prefix='<think>\\n', non_thinking_prefix='', history_thinking_prefix='')",
531
- "_val_dataset_exists": false,
532
- "hub": "<class 'swift.hub.hub.MSHub'>",
533
- "evaluation_strategy": "steps",
534
- "training_args": "GRPOConfig(output_dir='/data2/longfeiqi/sutando/MLLM4RSVG/output/GRPO/v64-20260329-180422', overwrite_output_dir=False, do_train=False, do_eval=False, do_predict=False, eval_strategy=<IntervalStrategy.NO: 'no'>, prediction_loss_only=False, per_device_train_batch_size=8, per_device_eval_batch_size=1, per_gpu_train_batch_size=None, per_gpu_eval_batch_size=None, gradient_accumulation_steps=4, eval_accumulation_steps=None, eval_delay=0, torch_empty_cache_steps=None, learning_rate=5e-06, weight_decay=0.1, adam_beta1=0.9, adam_beta2=0.95, adam_epsilon=1e-08, max_grad_norm=1.0, num_train_epochs=1.0, max_steps=-1, lr_scheduler_type=<SchedulerType.COSINE: 'cosine'>, lr_scheduler_kwargs=None, warmup_ratio=0.05, warmup_steps=0, log_level='passive', log_level_replica='warning', log_on_each_node=True, logging_dir='/data2/longfeiqi/sutando/MLLM4RSVG/output/GRPO/v64-20260329-180422/runs', logging_strategy=<IntervalStrategy.STEPS: 'steps'>, logging_first_step=True, logging_steps=1, logging_nan_inf_filter=True, save_strategy=<SaveStrategy.STEPS: 'steps'>, save_steps=100, save_total_limit=3, save_safetensors=True, save_on_each_node=False, save_only_model=False, restore_callback_states_from_checkpoint=False, no_cuda=False, use_cpu=False, use_mps_device=False, seed=42, data_seed=42, jit_mode_eval=False, bf16=True, fp16=False, fp16_opt_level='O1', half_precision_backend='auto', bf16_full_eval=False, fp16_full_eval=False, tf32=None, local_rank=0, ddp_backend=None, tpu_num_cores=None, tpu_metrics_debug=False, debug=[], dataloader_drop_last=True, eval_steps=100.0, dataloader_num_workers=4, dataloader_prefetch_factor=2, past_index=-1, run_name='/data2/longfeiqi/sutando/MLLM4RSVG/output/GRPO/v64-20260329-180422', disable_tqdm=False, remove_unused_columns=False, label_names=None, load_best_model_at_end=False, metric_for_best_model='loss', greater_is_better=False, ignore_data_skip=False, fsdp=[], fsdp_min_num_params=0, fsdp_config={'min_num_params': 0, 'xla': False, 'xla_fsdp_v2': False, 'xla_fsdp_grad_ckpt': False}, fsdp_transformer_layer_cls_to_wrap=None, accelerator_config=AcceleratorConfig(split_batches=False, dispatch_batches=False, even_batches=True, use_seedable_sampler=True, non_blocking=False, gradient_accumulation_kwargs=None, use_configured_state=False), parallelism_config=None, deepspeed={'fp16': {'enabled': 'auto', 'loss_scale': 0, 'loss_scale_window': 1000, 'initial_scale_power': 16, 'hysteresis': 2, 'min_loss_scale': 1}, 'bf16': {'enabled': 'auto'}, 'zero_optimization': {'stage': 2, 'offload_optimizer': {'device': 'none', 'pin_memory': True}, 'allgather_partitions': True, 'allgather_bucket_size': 200000000.0, 'overlap_comm': False, 'reduce_scatter': True, 'reduce_bucket_size': 200000000.0, 'contiguous_gradients': True}, 'gradient_accumulation_steps': 'auto', 'gradient_clipping': 'auto', 'steps_per_print': 2000, 'train_batch_size': 'auto', 'train_micro_batch_size_per_gpu': 'auto', 'wall_clock_breakdown': False}, label_smoothing_factor=0.0, optim=<OptimizerNames.ADAMW_TORCH_FUSED: 'adamw_torch_fused'>, optim_args=None, adafactor=False, group_by_length=False, length_column_name='length', report_to=['swanlab'], project='huggingface', trackio_space_id='trackio', ddp_find_unused_parameters=None, ddp_bucket_cap_mb=None, ddp_broadcast_buffers=None, dataloader_pin_memory=True, dataloader_persistent_workers=False, skip_memory_metrics=True, use_legacy_prediction_loop=False, push_to_hub=False, resume_from_checkpoint=None, hub_model_id=None, hub_strategy=<HubStrategy.EVERY_SAVE: 'every_save'>, hub_token=None, hub_private_repo=None, hub_always_push=False, hub_revision=None, gradient_checkpointing=True, gradient_checkpointing_kwargs=None, include_inputs_for_metrics=False, include_for_metrics=[], eval_do_concat_batches=True, fp16_backend='auto', push_to_hub_model_id=None, push_to_hub_organization=None, push_to_hub_token=None, mp_parameters='', auto_find_batch_size=False, full_determinism=False, torchdynamo=None, ray_scope='last', ddp_timeout=18000000, torch_compile=False, torch_compile_backend=None, torch_compile_mode=None, include_tokens_per_second=None, include_num_input_tokens_seen=None, neftune_noise_alpha=None, optim_target_modules=None, batch_eval_metrics=False, eval_on_start=False, use_liger_kernel=False, liger_kernel_config=None, eval_use_gather_object=False, average_tokens_across_devices=None, model_init_kwargs=None, disable_dropout=False, max_prompt_length=512, num_generations=8, max_completion_length=256, ds3_gather_for_generation=True, shuffle_dataset=True, generation_batch_size=128, steps_per_generation=4, temperature=0.9, top_p=0.9, top_k=50, min_p=None, generation_kwargs=None, repetition_penalty=1.0, use_transformers_paged=False, cache_implementation=None, use_vllm=True, vllm_mode='server', vllm_model_impl='vllm', vllm_enable_sleep_mode=False, vllm_guided_decoding_regex=None, vllm_server_base_url=None, vllm_server_host=['127.0.0.1'], vllm_server_port=[8897], vllm_server_timeout=240.0, vllm_gpu_memory_utilization=0.9, vllm_tensor_parallel_size=1, beta=0.02, num_iterations=1, epsilon=0.2, delta=None, epsilon_high=None, importance_sampling_level='token', reward_weights=[0.5, 0.5], scale_rewards='gdpo', loss_type='grpo', mask_truncated_completions=False, sync_ref_model=False, ref_model_mixup_alpha=0.6, ref_model_sync_steps=512, top_entropy_quantile=1.0, use_liger_loss=False, vllm_importance_sampling_correction=True, vllm_importance_sampling_cap=2.0, log_completions=True, num_completions_to_print=None, wandb_log_unique_prompts=None, tuner_backend='peft', vit_gradient_checkpointing=True, router_aux_loss_coef=0.0, enable_dft_loss=False, enable_channel_loss=False, check_model=True, acc_strategy='token', train_dataloader_shuffle=True, max_epochs=None, aligner_lr=None, vit_lr=None, use_logits_to_keep=None, resume_only_model=False, optimizer=None, eval_metric=None, callbacks=[], early_stop_interval=None, eval_use_evalscope=False, eval_dataset=[], eval_dataset_args=None, eval_limit=None, eval_generation_config=None, extra_eval_args=None, tuner_type='lora', use_galore=False, galore_target_modules=None, galore_rank=128, galore_update_proj_gap=50, galore_scale=1.0, galore_proj_type='std', galore_optim_per_parameter=False, galore_with_embedding=False, galore_quantization=False, galore_proj_quant=False, galore_proj_bits=4, galore_proj_group_size=256, galore_cos_threshold=0.4, galore_gamma_proj=2, galore_queue_size=5, lisa_activated_layers=0, lisa_step_interval=20, use_flash_ckpt=False, vllm_pipeline_parallel_size=1, vllm_enable_expert_parallel=False, vllm_max_num_seqs=None, vllm_max_model_len=None, vllm_disable_custom_all_reduce=True, vllm_enforce_eager=False, vllm_limit_mm_per_prompt=None, vllm_max_lora_rank=16, vllm_enable_prefix_caching=True, vllm_use_async_engine=None, vllm_quantization=None, vllm_reasoning_parser=None, vllm_disable_cascade_attn=False, vllm_mm_processor_cache_gb=None, vllm_speculative_config=None, vllm_engine_kwargs={}, vllm_data_parallel_size=1, stop_words=[], vllm_enable_lora=False, lora_rank=16, vllm_server_group_port=[51226], enable_flattened_weight_sync=True, async_generate=False, structured_outputs_regex=None, sleep_level=0, move_model_batches=None, offload_optimizer=False, offload_model=False, cosine_min_len_value_wrong=-0.5, cosine_max_len_value_wrong=0.0, cosine_min_len_value_correct=1.0, cosine_max_len_value_correct=0.5, cosine_max_len=256, repetition_n_grams=3, repetition_max_penalty=-1.0, reward_model=None, reward_model_plugin=None, chord_sft_dataset=[], chord_sft_per_device_train_batch_size=None, chord_enable_phi_function=False, chord_mu_warmup_steps=None, chord_mu_decay_steps=None, chord_mu_peak=None, chord_mu_valley=None, multi_turn_scheduler=None, max_turns=None, completion_length_limit_scope='per_round', vllm_server_pass_dataset=False, dynamic_sample=False, max_resample_times=3, overlong_filter=False, soft_max_length=None, soft_cache_length=None, log_entropy=False, tau_pos=1.0, tau_neg=1.05, advantage_estimator='grpo', kl_in_reward=False, num_generations_eval=None, dataset_shuffle=True, rollout_importance_sampling_mode=None, rollout_importance_sampling_threshold=2.0, log_rollout_offpolicy_metrics=False, off_policy_sequence_mask_delta=None)"
535
  }
 
1
  {
2
+ "swift_version": "4.0.0.dev0",
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
3
  "model_type": "qwen3_vl",
 
 
 
 
 
 
 
 
 
 
 
 
 
 
4
  "template": "qwen3_vl",
5
+ "norm_bbox": "norm1000",
6
+ "task_type": "causal_lm",
7
+ "tuner_type": "full",
8
+ "torch_dtype": "bfloat16"
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
9
  }