-
-Code4Scene: Benchmarking Coding Agents for Constructing and Editing 3D Scenes
-
-
-
-
+
+Code4Scene: Benchmarking Coding Agents for Constructing and Editing 3D Scenes
+
+
+
+
-
-
-
Code4Scene: Benchmarking Coding Agents for Constructing and Editing 3D Scenes
-
190 Unreal Engine cases built from human-assembled scenes. Coding agents write and run code that builds a scene from text, or
-repairs one from reference images, and Code4Scene scores the engine-native scene they save, not their code or a rendered view.
Benchmarking coding agents that turn text and reference images into engine-native 3D scenes. Measuring spatial reasoning, task fulfillment and precise control of scene state.
Task fulfillment, artifact integrity and static physical validity.
+ Full task instructionTEXT-TO-SCENE
Build me a compact, sunlit Copenhagen-style canal district with narrow cobblestone streets and rectangular waterways enclosing dense urban blocks. Arrange roughly thirty attached four- to six-storey townhouses in continuous street and canal-front rows, using slender façades, repetitive white-framed windows, arched doors and carriage entrances, modest cornices, and steep tiled roofs packed with dormers and brick chimneys. Include occasional exposed party walls between staggered building rows. Finish the façades in weathered muted red, mustard yellow, dusty blue, warm grey, beige, and exposed brick, with red, orange, and charcoal roof tiles. Run a principal waterfront street along a stone quay, crossed by smaller streets and linked across the canals by short, gently arched stone-and-metal bridges, including one opening onto a broad paved promenade with rounded landings. Form the canal edges from stepped pale stone, dark timber retaining walls, and iron railings around reflective rippling water. Add regularly spaced leafy trees, ornate black lamps, stone bollards, benches, European road signs, zebra crossings, puddles, grime, and sparse weeds.
+
Setting 1 · construction
Text-to-Scene
+
From an empty level and the content pack’s asset catalog, the agent builds the scene a prompt describes. Many realizations are valid; the evaluator judges requirements rather than matching a reference.
Given a corrupted scene and reference views of the original, restore every target actor within 5 cm, 5° and 5%, while preserving the surrounding scene.
+
case score = 0.8 × Repair F1 + 0.2 × Physical Safety
320 full benchmark target · 160 construction + 160 editing, including private cases201 public cases · 129 construction + 72 editing (22 indoor, 50 outdoor)
+
Leaderboard
Leaderboard
+
14 coding-agent configurations · paper results on the original 95 public cases (20 construction + 75 editing). These results predate the 201-case public release.
+
Code4Scene
Paper evaluation · score out of 100
Swipe to view all 14 configurations →
Sorted by score. Bar labels round scores ×100; the table preserves three decimal places. One color per model provider.
+
Score against cost
+
Each agent's score against what one case cost it. The line is the frontier: no agent left of it scores higher. Hover a point for its numbers.
+
OpenAIGoogleAnthropicMetaZ.aiAlibabaxAIThinking MachinesDeepSeekon the frontier
+
Cost is the mean reported cost of one run, estimated from its token usage at the provider's list price (paper Figure 1b; provisional).
+
Ranking
Overall = ½ Text-to-Scene + ½ Image-to-Scene over the 95 public cases; its cost is weighted the same way.
75 public cases (25 indoor, 50 outdoor). Case score = 0.8 × Repair F1 + 0.2 × Physical. Repair F1 counts target actors restored to the withheld ground truth.
Text-to-Scene (left) and Image-to-Scene (right) score per agent, in overall order (paper Table 2). Astra and Gemini tie at 0.619 overall.
Takeaway 2
Spatial Composition remains the weakest requirement family for all 14 agents.
Errors persist even when the required objects are present: generating the right objects does not ensure that their relationships satisfy the specification.
Each row is one agent's four family scores; the orange dot, Spatial Composition, is the lowest in every row (paper Figure 14b).
The objects are there; the relation is not
✓ slab roof 0.95✓ stone coping 0.82✗ slab roof enclosed by the coping 0.78
The judge's own frame of GPT-5.6 Sol on Egyptian Temple. It finds the slab roof and the stone coping, but not the roof enclosed by the coping: the coping is separate low wall segments with visible gaps (ringed).
Mismatch rate when the required objects are present
Across the benchmark (paper Figure 6a).
Takeaway 3
Editing ability varies substantially across repair types.
Astra performs best on Transform repairs, while Gemini is stronger on Lifecycle and Layout operations.
Mean Image-to-Scene score by repair type for four agents; the outlined cell is the best of the four in its row (paper Figure 6b).
Takeaway 4
Recovering the target does not guarantee precise editing.
Across all 14 configurations, 35.8% of complete recoveries still contain unintended changes to the surrounding scene.
1University of California San Diego 2University of California, Berkeley
+*Equal contribution †Corresponding author
Paper abstract 190-case benchmark · 95-case public evaluation
Frontier coding agents can now write and execute code that authors 3D environments, but whether they reliably
understand 3D structure and precisely control scene state remains unclear. The generated 3D scene is a persistent, executable artifact: a convincing
render can hide incorrect spatial relations, intersecting objects, or unintended modifications. We introduce Code4Scene, a benchmark of 190 Unreal
Engine cases built from human-assembled scenes that evaluates coding agents on two complementary settings under a shared execution interface.
@@ -58,39 +69,7 @@ set, construction and editing performance are strongly correlated but not interc
Gemini 3.8 Flash leads editing, and GPT-6 Astra narrowly leads overall. Spatial Composition is the weakest construction category for every agent,
while editing remains imprecise: the best Repair F1 is only 0.527, and 35.8% of edits that fully recover the target still introduce unintended
changes elsewhere in the scene. These results expose a gap between plausible 3D generation and reliable spatial reasoning and state control.
1 · InputPrompt or referenceAn open-ended scene description, or a corrupted scene with reference images of the original.
2 · AgentCoding agentA frontier or open-weight model in its own CLI harness, with the Unreal Engine editor as a tool.
3 · CodeCode in the engineIt writes and runs Python against the editor: spawns, moves and edits actors, inspects, revises.
4 · SceneEngine-native sceneThe saved level (.umap): every actor, transform and material, as the engine holds it.
5 · EvaluatorScored in the engineRequirements located and judged, overview views judged, physics measured, edits diffed against the withheld ground truth.
-
-
Setting 1 · construction
Text-to-Scene
-
From an empty level and the content pack's asset catalog, the agent builds the scene a prompt describes. Many realizations are valid, so the
-evaluator judges requirements rather than matching a reference. 20 public cases.
The agent gets a corrupted copy of a human-assembled scene and reference views of the original. It must restore every target actor (within
-5 cm, 5° and 5%) and change nothing else. 75 public cases: 25 indoor, 50 outdoor.
-
case score = 0.8 × Repair F1 + 0.2 × Physical Safety
-
-
Leaderboard
14 coding agents on the public set
-
Scores are from the paper's Table 2. Pick a setting; the chart and the table follow.
-
-
Score against cost
-
Each agent's score against what one case cost it. The line is the frontier: no agent left of it scores higher. Hover a point for its numbers.
-
OpenAIGoogleAnthropicMetaZ.aiAlibabaxAIThinking MachinesDeepSeekon the frontier
-
Cost is the mean reported cost of one run, estimated from its token usage at the provider's list price (paper Figure 1b; provisional).
-
Ranking
Overall = ½ Text-to-Scene + ½ Image-to-Scene over the 95 public cases; its cost is weighted the same way.
75 public cases (25 indoor, 50 outdoor). Case score = 0.8 × Repair F1 + 0.2 × Physical. Repair F1 counts target actors restored to the withheld ground truth.
Text-to-Scene (left) and Image-to-Scene (right) score per agent, in overall order (paper Table 2). Astra and Gemini tie at 0.619 overall.
Takeaway 2
Spatial Composition remains the weakest requirement family for all 14 agents.
Errors persist even when the required objects are present: generating the right objects does not ensure that their relationships satisfy the specification.
Each row is one agent's four family scores; the orange dot, Spatial Composition, is the lowest in every row (paper Figure 14b).
The objects are there; the relation is not
✓ slab roof 0.95✓ stone coping 0.82✗ slab roof enclosed by the coping 0.78
The judge's own frame of GPT-5.6 Sol on Egyptian Temple. It finds the slab roof and the stone coping, but not the roof enclosed by the coping: the coping is separate low wall segments with visible gaps (ringed).
Mismatch rate when the required objects are present
Across the benchmark (paper Figure 6a).
Takeaway 3
Editing ability varies substantially across repair types.
Astra performs best on Transform repairs, while Gemini is stronger on Lifecycle and Layout operations.
Mean Image-to-Scene score by repair type for four agents; the outlined cell is the best of the four in its row (paper Figure 6b).
Takeaway 4
Recovering the target does not guarantee precise editing.
Across all 14 configurations, 35.8% of complete recoveries still contain unintended changes to the surrounding scene.
1 · InputPrompt or referenceAn open-ended scene description, or a corrupted scene with reference images of the original.
2 · AgentCoding agentA frontier or open-weight model in its own CLI harness, with the Unreal Engine editor as a tool.
3 · CodeCode in the engineIt writes and runs Python against the editor: spawns, moves and edits actors, inspects, revises.
4 · SceneEngine-native sceneThe saved level (.umap): every actor, transform and material, as the engine holds it.
5 · EvaluatorScored in the engineRequirements located and judged, overview views judged, physics measured, edits diffed against the withheld ground truth.
Citation
Cite This Work
@misc{ye2026code4scene,
title = {Code4Scene: Benchmarking Coding Agents for Constructing and Editing 3D Scenes},
@@ -102,12 +81,13 @@ evaluator judges requirements rather than matching a reference. 20 public cases.
url = {https://arxiv.org/abs/2609.36777},
}