Benchmarking coding agents that turn text and reference images into engine-native 3D scenes. Measuring spatial reasoning, task fulfillment and precise control of scene state.
Task fulfillment, artifact integrity and static physical validity.
+ Full task instructionTEXT-TO-SCENE
Build me a compact, sunlit Copenhagen-style canal district with narrow cobblestone streets and rectangular waterways enclosing dense urban blocks. Arrange roughly thirty attached four- to six-storey townhouses in continuous street and canal-front rows, using slender façades, repetitive white-framed windows, arched doors and carriage entrances, modest cornices, and steep tiled roofs packed with dormers and brick chimneys. Include occasional exposed party walls between staggered building rows. Finish the façades in weathered muted red, mustard yellow, dusty blue, warm grey, beige, and exposed brick, with red, orange, and charcoal roof tiles. Run a principal waterfront street along a stone quay, crossed by smaller streets and linked across the canals by short, gently arched stone-and-metal bridges, including one opening onto a broad paved promenade with rounded landings. Form the canal edges from stepped pale stone, dark timber retaining walls, and iron railings around reflective rippling water. Add regularly spaced leafy trees, ornate black lamps, stone bollards, benches, European road signs, zebra crossings, puddles, grime, and sparse weeds.
-
Setting 1 · construction
Text-to-Scene
-
From an empty level and the content pack’s asset catalog, the agent builds the scene a prompt describes. Many realizations are valid; the evaluator judges requirements rather than matching a reference.
Given a corrupted scene and reference views of the original, restore every target actor within 5 cm, 5° and 5%, while preserving the surrounding scene.
-
case score = 0.8 × Repair F1 + 0.2 × Physical Safety
320 full benchmark target · 160 construction + 160 editing, including private cases201 public cases · 129 construction + 72 editing (22 indoor, 50 outdoor)
Benchmarking coding agents that turn text and reference images into engine-native 3D scenes. Measuring spatial reasoning, task fulfillment and precise control of scene state.
Task fulfillment, artifact integrity and static physical validity.
0.2 × Detailed + 0.6 × Overview + 0.2 × Physical
+ Full task instructionTEXT-TO-SCENE
Build me a compact, sunlit Copenhagen-style canal district with narrow cobblestone streets and rectangular waterways enclosing dense urban blocks. Arrange roughly thirty attached four- to six-storey townhouses in continuous street and canal-front rows, using slender façades, repetitive white-framed windows, arched doors and carriage entrances, modest cornices, and steep tiled roofs packed with dormers and brick chimneys. Include occasional exposed party walls between staggered building rows. Finish the façades in weathered muted red, mustard yellow, dusty blue, warm grey, beige, and exposed brick, with red, orange, and charcoal roof tiles. Run a principal waterfront street along a stone quay, crossed by smaller streets and linked across the canals by short, gently arched stone-and-metal bridges, including one opening onto a broad paved promenade with rounded landings. Form the canal edges from stepped pale stone, dark timber retaining walls, and iron railings around reflective rippling water. Add regularly spaced leafy trees, ornate black lamps, stone bollards, benches, European road signs, zebra crossings, puddles, grime, and sparse weeds.
Old Industrial Pallet Bay · Image-to-Scene
Repair the corrupted scene using its reference image; preserve all unrelated actors.
Restore each target within 5 cm, 5° and 5%, without unintended edits.
0.8 × Repair F1 + 0.2 × Physical
+ Full task instructionIMAGE-TO-SCENE
Repair the current outdoor scene so it matches the provided reference image. Make only the minimum changes needed, and preserve all unrelated actors and properties.
320 full benchmark target · 160 construction + 160 editing, including private cases201 public cases · 129 construction + 72 editing (22 indoor, 50 outdoor)
Leaderboard
Leaderboard
14 coding-agent configurations · paper results on the original 95 public cases (20 construction + 75 editing). These results predate the 201-case public release.
-
Code4Scene
Paper evaluation · score out of 100
Swipe to view all 14 configurations →
Sorted by score. Bar labels round scores ×100; the table preserves three decimal places. One color per model provider.
+
Code4Scene
Paper evaluation · score out of 100
Swipe to view all 14 configurations →
Sorted by score. Bar labels round scores ×100; the table preserves three decimal places. One color per model provider.
Code4Scene Pareto Frontier
Quality against the cost of one case · original 95-case public evaluation
Swipe to explore all configurations →
Solid: the frontier — nothing cheaper scores higher. One point per configuration, at its evaluated reasoning effort. Cost is the mean USD per case; overall averages the two task means.
-
Ranking
Overall = ½ Text-to-Scene + ½ Image-to-Scene over the 95 public cases; its cost is weighted the same way.
75 public cases (25 indoor, 50 outdoor). Case score = 0.8 × Repair F1 + 0.2 × Physical. Repair F1 counts target actors restored to the withheld ground truth.
75 public cases (25 indoor, 50 outdoor). Case score = 0.8 × Repair F1 + 0.2 × Physical. Repair F1 counts target actors restored to the withheld ground truth.
1University of California San Diego 2University of California, Berkeley
-*Equal contribution †Corresponding author
Paper abstract 190-case benchmark · 95-case public evaluation
Frontier coding agents can now write and execute code that authors 3D environments, but whether they reliably
-understand 3D structure and precisely control scene state remains unclear. The generated 3D scene is a persistent, executable artifact: a convincing
-render can hide incorrect spatial relations, intersecting objects, or unintended modifications. We introduce Code4Scene, a benchmark of 190 Unreal
-Engine cases built from human-assembled scenes that evaluates coding agents on two complementary settings under a shared execution interface.
-Construction tests scene-level spatial reasoning from open-ended language specifications, where many realizations are valid; editing tests precise
-control of scene state, where the agent must recover the target scene from reference images while preserving everything else.
-
Rather than scoring code or rendered views, Code4Scene evaluates the generated engine-native scene for task fulfillment, artifact integrity, and
-static physical validity, with edits additionally compared against withheld ground truth. Across 14 coding-agent configurations on the 95-case public
-set, construction and editing performance are strongly correlated but not interchangeable (Spearman ρ = 0.78): Claude Fable 5.1 leads construction,
-Gemini 3.8 Flash leads editing, and GPT-6 Astra narrowly leads overall. Spatial Composition is the weakest construction category for every agent,
-while editing remains imprecise: the best Repair F1 is only 0.527, and 35.8% of edits that fully recover the target still introduce unintended
-changes elsewhere in the scene. These results expose a gap between plausible 3D generation and reliable spatial reasoning and state control.
1 · InputPrompt or referenceAn open-ended scene description, or a corrupted scene with reference images of the original.
2 · AgentCoding agentA frontier or open-weight model in its own CLI harness, with the Unreal Engine editor as a tool.
3 · CodeCode in the engineIt writes and runs Python against the editor: spawns, moves and edits actors, inspects, revises.
4 · SceneEngine-native sceneThe saved level (.umap): every actor, transform and material, as the engine holds it.
5 · EvaluatorScored in the engineRequirements located and judged, overview views judged, physics measured, edits diffed against the withheld ground truth.
+
Citation
Cite This Work
@misc{ye2026code4scene,
title = {Code4Scene: Benchmarking Coding Agents for Constructing and Editing 3D Scenes},
@@ -84,4 +62,4 @@ changes elsewhere in the scene. These results expose a gap between plausible 3D
SimWorld Home