190 Unreal Engine cases built from human-assembled scenes. Coding agents write and run code that builds a scene from text, or repairs one from reference images, and Code4Scene scores the engine-native scene they save, not their code or a rendered view.
1University of California San Diego 2University of California, Berkeley *Equal contribution †Corresponding author
Coding agents can now operate 3D engines: they write and run code, inspect the result and revise the scene. Code4Scene is a benchmark of 190 Unreal Engine cases built from human-assembled scenes that evaluates them on two complementary settings under a shared execution interface. Construction starts from an empty level and an open-ended text specification. Editing gives the agent a corrupted copy of a scene and reference images of the original, and asks it to restore the intended state while leaving everything else untouched.
Rather than scoring code or rendered views, Code4Scene evaluates the engine-native scene the agent saved, on task fulfilment, artifact integrity and static physical validity, and compares edits against withheld ground truth. Across 14 agent configurations, construction and editing rank the agents differently, spatial relations remain the weakest requirement family, and a third of complete repairs still change the surrounding scene.
From an empty level and the content pack's asset catalog, the agent builds the scene a prompt describes. Many realizations are valid, so the evaluator judges requirements rather than matching a reference. 20 public cases.
The agent gets a corrupted copy of a human-assembled scene and reference views of the original. It must restore every target actor (within 5 cm, 5° and 5%) and change nothing else. 75 public cases: 25 indoor, 50 outdoor.
Scores are from the paper's Table 2. Pick a setting; the chart and the table follow.
Cost is the mean reported cost of one run, estimated from its token usage at the provider's list price (paper Figure 1b; provisional).
Overall = ½ Text-to-Scene + ½ Image-to-Scene over the 95 public cases; its cost is weighted the same way.
| # | Agent configuration | Overall | Text-to-Scene | Image-to-Scene | Cost / case |
|---|---|---|---|---|---|
| 1 | GPT-6 Astra (max)OpenAI | 0.724 | 0.515 | $12.90 | |
| 2 | Gemini 3.8 Flash (high)ParetoGoogle | 0.657 | 0.581 | $1.92 | |
| 3 | Claude Fable 5.1 (max)Anthropic | 0.788 | 0.424 | $10.62 | |
| 4 | Claude Opus 5 (max)Anthropic | 0.718 | 0.468 | $19.48 | |
| 5 | GPT-5.6 Sol (high)OpenAI | 0.707 | 0.393 | $2.97 | |
| 6 | Muse Spark 1.3 (medium)ParetoMeta | 0.646 | 0.358 | $0.09 | |
| 7 | GLM-5.3 Flash (max)open weightsZ.ai | 0.509 | 0.320 | $0.40 | |
| 8 | Qwen 3.8 27B (thinking off)open weightsAlibaba | 0.557 | 0.200 | $0.38 | |
| 9 | Grok 4.6 (high)xAI | 0.567 | 0.184 | $2.29 | |
| 10 | Qwen 3.8 27B (thinking on)open weightsAlibaba | 0.516 | 0.226 | $0.30 | |
| 11 | Inkling (high)Thinking Machines | 0.425 | 0.198 | $0.48 | |
| 12 | Gemma 4 31B (thinking on)open weightsParetoGoogle | 0.471 | 0.126 | $0.03 | |
| 13 | Gemma 4 31B (thinking off)open weightsParetoGoogle | 0.455 | 0.115 | $0.02 | |
| 14 | DeepSeek V4.1 Flash (high)open weightsDeepSeek | 0.243 | 0.210 | $0.12 |
Click a column to sort. Pareto: on the score-against-cost frontier.
What the evaluation shows across the 14 agents. Every chart below is drawn from the paper's numbers; hover for values.
Despite nearly identical overall scores, Astra is stronger at construction and Gemini at editing.
Errors persist even when the required objects are present: generating the right objects does not ensure that their relationships satisfy the specification.

Astra performs best on Transform repairs, while Gemini is stronger on Lifecycle and Layout operations.
Across all 14 configurations, 35.8% of complete recoveries still contain unintended changes to the surrounding scene.
For each selected case, every agent's saved scene is exported from Unreal Engine and opens in the browser in 3D, beside the evaluator's scores for it.
@article{code4scene2026,
title = {Code4Scene: Benchmarking Coding Agents for Constructing and Editing 3D Scenes},
author = {Ye, Xiaokang and Mantri, Siddhant Hitesh and Chen, Zimeng and Zhang, Edward and Zheng, Zhaoxu and Li, Yuanheng and Chen, Yizhao and Huang, Tianyang and Qin, Lianhui},
year = {2026}
}