File size: 6,363 Bytes
4d33727
7aa0415
 
 
 
 
 
fa2e2a1
 
 
 
 
 
4d33727
7aa0415
fa2e2a1
7aa0415
fa2e2a1
 
 
c5d1608
7aa0415
 
fa2e2a1
7aa0415
fa2e2a1
7aa0415
fa2e2a1
7aa0415
fa2e2a1
7aa0415
fa2e2a1
 
 
7aa0415
fa2e2a1
7aa0415
fa2e2a1
7aa0415
fa2e2a1
7aa0415
fa2e2a1
7aa0415
fa2e2a1
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
7aa0415
fa2e2a1
7aa0415
fa2e2a1
7aa0415
fa2e2a1
 
 
 
7aa0415
fa2e2a1
7aa0415
fa2e2a1
7aa0415
 
 
 
fa2e2a1
 
 
7aa0415
 
fa2e2a1
7aa0415
 
fa2e2a1
7aa0415
fa2e2a1
 
 
7aa0415
 
fa2e2a1
7aa0415
 
fa2e2a1
7aa0415
fa2e2a1
 
7aa0415
fa2e2a1
7aa0415
 
 
fa2e2a1
7aa0415
fa2e2a1
7aa0415
fa2e2a1
7aa0415
fa2e2a1
7aa0415
fa2e2a1
7aa0415
 
fa2e2a1
 
 
 
 
7aa0415
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
---
base_model:
- Qwen/Qwen3.6-35B-A3B
library_name: transformers
license: apache-2.0
license_link: https://huggingface.co/Qwen/Qwen3.6-35B-A3B/blob/main/LICENSE
pipeline_tag: image-text-to-text
tags:
- agent
- tool-use
- reinforcement-learning
- sft
- environment-scaling
---

# CompoWorld: Compositional Environment Scaling for General Agents

<p align="center">
  📄 <a href="https://huggingface.co/papers/2609.33665">Paper</a> &nbsp;|&nbsp;
  📚 <a href="https://arxiv.org/abs/2609.33665">arXiv</a> &nbsp;|&nbsp;
  💻 <a href="https://github.com/AllSpark-Research/AgentEnv">GitHub</a>
</p>

This repository contains **Qwen3.6-35B-A3B agent trained with CompoWorld**, a compositional environment scaling method for training general agents.

## TL;DR

Automatically generated environments provide a scalable source of interaction data for training general agents. However, existing approaches mainly generate tasks within a single environment, while real-world workflows require agents to connect information and actions across multiple services.

**CompoWorld** expands the task space by composing a finite library of reusable services:

- **Verified services**: coding agents turn tool specifications into verified services with typed states and shared interfaces, while a world model handles tools that cannot be reliably implemented.
- **Compositional task generation**: a random-walk procedure connects services through dependency graphs, enabling the generation and verification of tasks that require information to flow across services.
- **SFT + RL**: verified trajectories support supervised fine-tuning (SFT), while a **Completion-Focused Rubric Reward** guides reinforcement learning (RL) toward full task completion by emphasizing criteria with lower pass rates within each rollout group.

We construct **448 services exposing 10,130 tools** and use **3K SFT trajectories** and **1K RL tasks** to train Qwen3.6-35B-A3B.

## Results

CompoWorld improves on its backbone by **+9.17 points on average** across eight agent benchmarks. On AutomationBench, it surpasses frontier models such as Claude Opus 4.6 and leads all compared agent-specialized 35B-A3B models.

**Main results on eight challenging agent benchmarks** (from the [paper](https://huggingface.co/papers/2609.33665)):

| Model | τ³-Banking | DeepPlanning | VitaBench | VitaBench 2.0 | AutomationBench | WildClawBench | SkillsBench | ALE |
|---|---|---|---|---|---|---|---|---|
| **Frontier Closed-Source Models** | | | | | | | | |
| GPT-5.4 | 28.52 | 53.96 | 47.09 | 45.99 | 27.67 | 58.00 | 51.70 | 22.50 |
| Claude Opus 4.6 | 20.27 | 55.21 | 38.31 | 37.71 | 25.50 | 54.60 | 50.20 | 13.30 |
| Gemini-3.1 Pro | 23.71 | 47.08 | 52.69 | 49.53 | 28.17 | 38.70 | 60.80 | 17.10 |
| **Open-Weight Models** | | | | | | | | |
| DeepSeek-V4-Flash | 30.34 | 54.79 | 56.31 | 41.12 | 36.33 | 47.67 | 53.75 | 17.48 |
| GLM-5.2 | 28.87 | 50.00 | 50.26 | 45.32 | 26.33 | 54.20 | 62.10 | 22.33 |
| Kimi-K2.6 | 19.93 | 43.12 | 43.94 | 44.30 | 15.17 | 29.40 | 54.00 | 8.74 |
| Qwen3.8-27B | 33.68 | 52.92 | 41.84 | 47.23 | 38.50 | 56.20 | 35.57 | 20.40 |
| Qwen3.5-397B-A17B | 16.15 | 35.83 | 42.09 | 38.41 | 5.50 | 40.37 | 36.50 | 9.71 |
| Qwen3.6-35B-A3B | 10.65 | 26.04 | 38.94 | 34.47 | 10.33 | 44.28 | 32.52 | 7.77 |
| **Agent-Specialized Models (35B-A3B)** | | | | | | | | |
| Apodex 1.1 Mini | 12.71 | 34.17 | 44.06 | 35.29 | 14.17 | - | 32.29 | - |
| Occamy-1.0 | 37.10 | - | 41.75 | - | 27.60 | 49.16 | - | - |
| Agents-A1 | 7.20 | - | 37.00 | - | 2.20 | 30.73 | - | - |
| Nex-N2-mini | 25.80 | - | 26.25 | - | 5.70 | 30.31 | - | - |
| BigBang-1.0 | 10.30 | - | 46.00 | - | 14.80 | 32.87 | - | - |
| Ornith-1.5-35B | 21.70 | - | 40.25 | - | 18.50 | 45.91 | - | - |
| **CompoWorld (Ours)** | 16.49 | 35.21 | 49.44 | 36.09 | **32.33** | 47.46 | 47.71 | 13.59 |
| Δ vs. backbone | +5.84 | +9.17 | +10.50 | +1.62 | +22.00 | +3.18 | +15.19 | +5.82 |

**Evaluation protocol.** We use OpenHands for SkillsBench with skills enabled, Claude Code for ALE, and OpenClaw for WildClawBench. We report the average score or reward for WildClawBench and both VitaBench versions, Avg@3 for SkillsBench, the pass rate for τ³-Banking, AutomationBench, and ALE, and the average accuracy for DeepPlanning.

## Model Description

- **Base model:** [Qwen/Qwen3.6-35B-A3B](https://huggingface.co/Qwen/Qwen3.6-35B-A3B) (35B total parameters, 3B activated, MoE with vision encoder, 262,144-token context)
- **Training data:** 3K verified SFT trajectories + 1K RL tasks generated from 448 composed services (10,130 tools)
- **RL reward:** Completion-Focused Rubric Reward — emphasizes criteria with lower pass rates within each rollout group
- **Format:** Hugging Face Transformers (compatible with Transformers, vLLM, SGLang, KTransformers, etc.)

## Quickstart

### vLLM

```shell
uv pip install vllm --torch-backend=auto

vllm serve AllSpark-Research/CompoWorld --port 8000 --tp-size 8 \
  --tool-call-parser qwen3_coder --reasoning-parser qwen3 \
  --context-length 262144
```

### SGLang

```shell
uv pip install sglang[all]

python -m sglang.launch_server --model-path AllSpark-Research/CompoWorld \
  --port 8000 --tp-size 8 --mem-fraction-static 0.8 \
  --context-length 262144 --reasoning-parser qwen3 --tool-call-parser qwen3_coder
```

### Transformers

```python
from transformers import AutoModelForImageTextToText, AutoProcessor

model = AutoModelForImageTextToText.from_pretrained(
    "AllSpark-Research/CompoWorld", torch_dtype="auto", device_map="auto"
)
processor = AutoProcessor.from_pretrained("AllSpark-Research/CompoWorld")
```

> [!Note]
> The model has a default context length of 262,144 tokens. We advise maintaining a context length of at least 128K tokens to preserve thinking capabilities.

## License

This model is released under the [Apache 2.0 license](https://huggingface.co/Qwen/Qwen3.6-35B-A3B/blob/main/LICENSE), inherited from the base model.

## Citation

If you find this work useful, please cite:

```bibtex
@article{yang2026compoworld,
  title={CompoWorld: Compositional Environment Scaling for General Agents},
  author={Yang, Xiao-Wen and Xu, Weiyi and Da, Wen and Xu, Hang and Li, Canwei and You, Hong-Jie and Dong, Pusen and Zeng, Yucheng and Luo, Zhaokai and Li, Yu-Feng and Hu, Yao and Chuan, Mu},
  journal={arXiv preprint arXiv:2609.33665},
  year={2026}
}
```