diff --git "a/index.html" "b/index.html" --- "a/index.html" +++ "b/index.html" @@ -1,19 +1,107 @@ - -
- - -You can modify this app directly by editing index.html in the Files and versions tab.
-- Also don't forget to check the - Spaces documentation. -
-190 Unreal Engine cases built from human-assembled scenes. Coding agents write and run code that builds a scene from text, or +repairs one from reference images, and Code4Scene scores the engine-native scene they save, not their code or a rendered view.
+ +1University of California San Diego 2University of California, Berkeley +*Equal contribution †Corresponding author
+ +Coding agents can now operate 3D engines: they write and run code, inspect the result and revise the scene. Code4Scene is a +benchmark of 190 Unreal Engine cases built from human-assembled scenes that evaluates them on two complementary settings under a shared execution +interface. Construction starts from an empty level and an open-ended text specification. Editing gives the agent a corrupted copy of a +scene and reference images of the original, and asks it to restore the intended state while leaving everything else untouched.
+Rather than scoring code or rendered views, Code4Scene evaluates the engine-native scene the agent saved, on task fulfilment, artifact integrity +and static physical validity, and compares edits against withheld ground truth. Across 14 agent configurations, construction and editing rank the +agents differently, spatial relations remain the weakest requirement family, and a third of complete repairs still change the surrounding scene.
From an empty level and the content pack's asset catalog, the agent builds the scene a prompt describes. Many realizations are valid, so the +evaluator judges requirements rather than matching a reference. 20 public cases.
+The agent gets a corrupted copy of a human-assembled scene and reference views of the original. It must restore every target actor (within +5 cm, 5° and 5%) and change nothing else. 75 public cases: 25 indoor, 50 outdoor.
+Overall = ½ Text-to-Scene + ½ Image-to-Scene, over the 95 public cases. Scores are from the paper's Table 2.
Cost is the mean reported cost of one run, estimated from its token usage at the provider's list price (paper Figure 1b; provisional). +Some runs did not report usage; hover a point for how many of the 95 cases are priced.
| # | Agent configuration | Score | Detailed | Overview | Physical | Cost / case |
|---|---|---|---|---|---|---|
| 1 | Claude Fable 5.1 (max)ParetoAnthropic | 0.725 | 0.772 | 0.898 | $18.16 | |
| 2 | GPT-6 Astra (max)OpenAI | 0.707 | 0.733 | 0.712 | $22.47 | |
| 3 | Claude Opus 5 (max)Anthropic | 0.631 | 0.716 | 0.811 | $34.50 | |
| 4 | GPT-5.6 Sol (high)ParetoOpenAI | 0.632 | 0.710 | 0.772 | $4.07 | |
| 5 | Gemini 3.8 Flash (high)ParetoGoogle | 0.637 | 0.664 | 0.655 | $2.49 | |
| 6 | Muse Spark 1.3 (medium)ParetoMeta | 0.598 | 0.661 | 0.651 | $0.10 | |
| 7 | Grok 4.6 (high)xAI | 0.522 | 0.576 | 0.585 | $4.33 | |
| 8 | Qwen 3.8 27B (thinking off)open weightsAlibaba | 0.507 | 0.530 | 0.688 | $0.49 | |
| 9 | Qwen 3.8 27B (thinking on)open weightsAlibaba | 0.505 | 0.460 | 0.694 | $0.33 | |
| 10 | GLM-5.3 Flash (max)open weightsZ.ai | 0.403 | 0.537 | 0.533 | $0.71 | |
| 11 | Gemma 4 31B (thinking on)open weightsParetoGoogle | 0.412 | 0.416 | 0.694 | $0.03 | |
| 12 | Gemma 4 31B (thinking off)open weightsParetoGoogle | 0.438 | 0.370 | 0.725 | $0.02 | |
| 13 | +Inkling (high)Thinking Machines | 0.393 | 0.337 | 0.722 | $0.20 | |
| 14 | DeepSeek V4.1 Flash (high)open weightsDeepSeek | 0.217 | 0.161 | 0.513 | $0.19 |
| # | Agent configuration | Score | Repair F1 | Physical | Indoor | Outdoor | Cost / case |
|---|---|---|---|---|---|---|---|
| 1 | Gemini 3.8 Flash (high)ParetoGoogle | 0.527 | 0.796 | 0.717 | 0.513 | $1.35 | |
| 2 | GPT-6 Astra (max)OpenAI | 0.445 | 0.796 | 0.785 | 0.380 | $3.32 | |
| 3 | Claude Opus 5 (max)Anthropic | 0.389 | 0.786 | 0.618 | 0.393 | $4.46 | |
| 4 | Claude Fable 5.1 (max)Anthropic | 0.332 | 0.795 | 0.535 | 0.369 | $3.09 | |
| 5 | GPT-5.6 Sol (high)OpenAI | 0.319 | 0.689 | 0.497 | 0.341 | $1.88 | |
| 6 | Muse Spark 1.3 (medium)ParetoMeta | 0.246 | 0.807 | 0.633 | 0.221 | $0.07 | |
| 7 | GLM-5.3 Flash (max)open weightsZ.ai | 0.188 | 0.848 | 0.465 | 0.247 | $0.09 | |
| 8 | Qwen 3.8 27B (thinking on)open weightsAlibaba | 0.118 | 0.660 | 0.297 | 0.191 | $0.26 | |
| 9 | DeepSeek V4.1 Flash (high)open weightsParetoDeepSeek | 0.066 | 0.787 | 0.189 | 0.221 | $0.05 | |
| 10 | Qwen 3.8 27B (thinking off)open weightsAlibaba | 0.083 | 0.667 | 0.266 | 0.166 | $0.28 | |
| 11 | +Inkling (high)Thinking Machines | 0.095 | 0.609 | 0.169 | 0.213 | $0.75 | |
| 12 | Grok 4.6 (high)xAI | 0.044 | 0.744 | 0.197 | 0.178 | $0.26 | |
| 13 | Gemma 4 31B (thinking on)open weightsParetoGoogle | 0.003 | 0.615 | 0.127 | 0.125 | $0.03 | |
| 14 | Gemma 4 31B (thinking off)open weightsParetoGoogle | 0.001 | 0.572 | 0.112 | 0.116 | $0.02 |
What the evaluation shows across the 14 agents. Every chart below is drawn from the paper's numbers; hover for values.
Despite nearly identical overall scores, Astra is stronger at construction and Gemini at editing.
Errors persist even when the required objects are present: generating the right objects does not ensure that their relationships satisfy the specification.

Astra performs best on Transform repairs, while Gemini is stronger on Lifecycle and Layout operations.
Across all 14 configurations, 35.8% of complete recoveries still contain unintended changes to the surrounding scene.
For each selected case, every agent's saved scene is exported from Unreal Engine and opens in the browser in 3D, beside the evaluator's scores for it.
@article{code4scene2026,
+ title = {Code4Scene: Benchmarking Coding Agents for Constructing and Editing 3D Scenes},
+ author = {Ye, Xiaokang and Mantri, Siddhant Hitesh and Chen, Zimeng and Zhang, Edward and Zheng, Zhaoxu and Li, Yuanheng and Chen, Yizhao and Huang, Tianyang and Qin, Lianhui},
+ year = {2026}
+}