-Code4Scene: Benchmarking Coding Agents for Constructing and Editing 3D Scenes
+Code4Scene: Benchmarking Coding Agents for Constructing 3D ScenesSkip to benchmark explorer
-
Benchmarking coding agents that turn text and reference images into engine-native 3D scenes. Measuring spatial reasoning, task fulfillment and precise control of scene state.
Task fulfillment, artifact integrity and static physical validity.
0.2 × Detailed + 0.6 × Overview + 0.2 × Physical
+ Full task instructionTEXT-TO-SCENE
Build me a compact, sunlit Copenhagen-style canal district with narrow cobblestone streets and rectangular waterways enclosing dense urban blocks. Arrange roughly thirty attached four- to six-storey townhouses in continuous street and canal-front rows, using slender façades, repetitive white-framed windows, arched doors and carriage entrances, modest cornices, and steep tiled roofs packed with dormers and brick chimneys. Include occasional exposed party walls between staggered building rows. Finish the façades in weathered muted red, mustard yellow, dusty blue, warm grey, beige, and exposed brick, with red, orange, and charcoal roof tiles. Run a principal waterfront street along a stone quay, crossed by smaller streets and linked across the canals by short, gently arched stone-and-metal bridges, including one opening onto a broad paved promenade with rounded landings. Form the canal edges from stepped pale stone, dark timber retaining walls, and iron railings around reflective rippling water. Add regularly spaced leafy trees, ornate black lamps, stone bollards, benches, European road signs, zebra crossings, puddles, grime, and sparse weeds.
Old Industrial Pallet Bay · Image-to-Scene
Repair the corrupted scene using its reference image; preserve all unrelated actors.
Restore each target within 5 cm, 5° and 5%, without unintended edits.
0.8 × Repair F1 + 0.2 × Physical
+ Full task instructionIMAGE-TO-SCENE
Repair the current outdoor scene so it matches the provided reference image. Make only the minimum changes needed, and preserve all unrelated actors and properties.
320 full benchmark target · 160 construction + 160 editing, including private cases201 public cases · 129 construction + 72 editing (22 indoor, 50 outdoor)
Benchmarking coding agents that turn open-ended scene descriptions into engine-native 3D scenes in Unreal Engine. Measuring spatial reasoning, task fulfillment and physical validity.
Task fulfillment, artifact integrity and static physical validity.
0.2 × Detailed + 0.6 × Overview + 0.2 × Physical
+ Full task instructionTEXT-TO-SCENE
Build me a compact, sunlit Copenhagen-style canal district with narrow cobblestone streets and rectangular waterways enclosing dense urban blocks. Arrange roughly thirty attached four- to six-storey townhouses in continuous street and canal-front rows, using slender façades, repetitive white-framed windows, arched doors and carriage entrances, modest cornices, and steep tiled roofs packed with dormers and brick chimneys. Include occasional exposed party walls between staggered building rows. Finish the façades in weathered muted red, mustard yellow, dusty blue, warm grey, beige, and exposed brick, with red, orange, and charcoal roof tiles. Run a principal waterfront street along a stone quay, crossed by smaller streets and linked across the canals by short, gently arched stone-and-metal bridges, including one opening onto a broad paved promenade with rounded landings. Form the canal edges from stepped pale stone, dark timber retaining walls, and iron railings around reflective rippling water. Add regularly spaced leafy trees, ornate black lamps, stone bollards, benches, European road signs, zebra crossings, puddles, grime, and sparse weeds.
160 construction cases in the full benchmark target, including private cases129 public construction cases
Leaderboard
Leaderboard
-
14 coding-agent configurations · paper results on the original 95 public cases (20 construction + 75 editing). These results predate the 201-case public release.
-
Code4Scene
Paper evaluation · score out of 100
Swipe to view all 14 configurations →
Sorted by score. Bar labels round scores ×100; the table preserves three decimal places. One color per model provider.
-
Code4Scene Pareto Frontier
Quality against the cost of one case · original 95-case public evaluation
Swipe to explore all configurations →
Solid: the frontier — nothing cheaper scores higher. One point per configuration, at its evaluated reasoning effort. Cost is the mean USD per case; overall averages the two task means.
-
14 configurations
Ranking
Overall = ½ Text-to-Scene + ½ Image-to-Scene over the 95 public cases; its cost is weighted the same way.
75 public cases (25 indoor, 50 outdoor). Case score = 0.8 × Repair F1 + 0.2 × Physical. Repair F1 counts target actors restored to the withheld ground truth.
Text-to-Scene (left) and Image-to-Scene (right) score per agent, in overall order (paper Table 2). Astra and Gemini both round to 0.619 overall.
Takeaway 2
Spatial Composition remains the weakest requirement family for all 14 agents.
Errors persist even when the required objects are present: generating the right objects does not ensure that their relationships satisfy the specification.
Each row is one agent's four family scores; the orange dot, Spatial Composition, is the lowest in every row (paper Figure 14b).
The objects are there; the relation is not
Loading 3D scene…Drag to orbit · Scroll to zoom
✓ slab roof 0.95✓ stone coping 0.82✗ slab roof enclosed by the coping 0.78
GPT-5.6 Sol’s saved Egyptian Temple scene. The judge finds the roof and stone coping, but flags their spatial relationship. View the judge’s evidence ↗
Mismatch rate when the required objects are present
Across the benchmark (paper Figure 6a).
Takeaway 3
Editing ability varies substantially across repair types.
Astra performs best on Transform repairs, while Gemini is stronger on Lifecycle and Layout operations.
Mean Image-to-Scene score by repair type for four agents; the outlined cell is the best of the four in its row (paper Figure 6b).
Takeaway 4
Recovering the target does not guarantee precise editing.
Across all 14 configurations, 35.8% of complete recoveries still contain unintended changes to the surrounding scene.
One dot per Image-to-Scene evaluation (paper Figure 6c).
+
Takeaway
Spatial Composition remains the weakest requirement family for all 14 agents.
Errors persist even when the required objects are present: generating the right objects does not ensure that their relationships satisfy the specification.
Each row is one agent's four family scores; the orange dot, Spatial Composition, is the lowest in every row (paper Figure 14b).
The objects are there; the relation is not
Loading 3D scene…Drag to orbit · Scroll to zoom
✓ slab roof 0.95✓ stone coping 0.82✗ slab roof enclosed by the coping 0.78
GPT-5.6 Sol’s saved Egyptian Temple scene. The judge finds the roof and stone coping, but flags their spatial relationship. View the judge’s evidence ↗
Mismatch rate when the required objects are present
Across the benchmark (paper Figure 6a).
Citation
Cite This Work
@misc{ye2026code4scene,
title = {Code4Scene: Benchmarking Coding Agents for Constructing and Editing 3D Scenes},
@@ -50,12 +48,12 @@ svg.ch .dim{opacity:.42}svg.ch g.row:hover{opacity:1}svg.ch .hot{fill:var(--c4s-
url = {https://arxiv.org/abs/2609.36777},
}