diff --git "a/index.html" "b/index.html" --- "a/index.html" +++ "b/index.html" @@ -1,13 +1,14 @@ - + + - -Code4Scene: Benchmarking Coding Agents for Constructing and Editing 3D Scenes - - - - + +Code4Scene: Benchmarking Coding Agents for Constructing and Editing 3D Scenes + + + + - -
-

Code4Scene:
Benchmarking Coding Agents for
Constructing and Editing 3D Scenes

-

190 Unreal Engine cases built from human-assembled scenes. Coding agents write and run code that builds a scene from text, or -repairs one from reference images, and Code4Scene scores the engine-native scene they save, not their code or a rendered view.

-

Xiaokang Ye1,*, Siddhant Hitesh Mantri1,*, -Zimeng Chen1,*, Edward Zhang1, Zhaoxu Zheng1,
-Yuanheng Li2, Yizhao Chen1, Tianyang Huang1, -Lianhui Qin1,†

-

1University of California San Diego    2University of California, Berkeley    -*Equal contribution    †Corresponding author

- -
-
190
Unreal Engine cases
95 public, scored here
-
14
coding-agent configurations
frontier and open weights
-
0.619
best overall score
Astra and Gemini, tied
-
35.8%
of complete repairs
still change the rest of the scene
-
-

How Code4Scene scores a scene

+ +

Constructing and Editing
3D Scenes with Code

Benchmarking coding agents that turn text and reference images into engine-native 3D scenes. Measuring spatial reasoning, task fulfillment and precise control of scene state.

Task Format

Nordic Harbour · Text-to-Scene

An open-ended scene description, an asset catalog and the Unreal Engine editor.

Agent inputText + assets
Task instruction

Build me a compact, sunlit Copenhagen-style canal district with narrow cobblestone streets and rectangular waterways enclosing dense urban blocks.

Townhouses, canals, bridges and a waterfront promenade — with detailed requirements for materials, layout and street furniture.

Given: an empty level and the content pack’s asset catalog.

EvaluationEngine-native scene
CASE SCORE0.877
Detailed Alignment
0.757
Overview Alignment
0.909
Physical Safety
0.900

Task fulfillment, artifact integrity and static physical validity.

+ Full task instructionTEXT-TO-SCENE

Build me a compact, sunlit Copenhagen-style canal district with narrow cobblestone streets and rectangular waterways enclosing dense urban blocks. Arrange roughly thirty attached four- to six-storey townhouses in continuous street and canal-front rows, using slender façades, repetitive white-framed windows, arched doors and carriage entrances, modest cornices, and steep tiled roofs packed with dormers and brick chimneys. Include occasional exposed party walls between staggered building rows. Finish the façades in weathered muted red, mustard yellow, dusty blue, warm grey, beige, and exposed brick, with red, orange, and charcoal roof tiles. Run a principal waterfront street along a stone quay, crossed by smaller streets and linked across the canals by short, gently arched stone-and-metal bridges, including one opening onto a broad paved promenade with rounded landings. Form the canal edges from stepped pale stone, dark timber retaining walls, and iron railings around reflective rippling water. Add regularly spaced leafy trees, ornate black lamps, stone bollards, benches, European road signs, zebra crossings, puddles, grime, and sparse weeds.

+
Setting 1 · construction

Text-to-Scene

+

From an empty level and the content pack’s asset catalog, the agent builds the scene a prompt describes. Many realizations are valid; the evaluator judges requirements rather than matching a reference.

+
case score = 0.2 × Detailed Alignment + 0.6 × Overview Alignment + 0.2 × Physical Safety
+
Setting 2 · editing

Image-to-Scene

+

Given a corrupted scene and reference views of the original, restore every target actor within 5 cm, 5° and 5%, while preserving the surrounding scene.

+
case score = 0.8 × Repair F1 + 0.2 × Physical Safety
320 full benchmark target · 160 construction + 160 editing, including private cases201 public cases · 129 construction + 72 editing (22 indoor, 50 outdoor)
+

Leaderboard

+

14 coding-agent configurations · paper results on the original 95 public cases (20 construction + 75 editing). These results predate the 201-case public release.

+

Code4Scene

Paper evaluation · score out of 100

Swipe to view all 14 configurations →

Sorted by score. Bar labels round scores ×100; the table preserves three decimal places. One color per model provider.

+
Score against cost
+
Each agent's score against what one case cost it. The line is the frontier: no agent left of it scores higher. Hover a point for its numbers.
+
OpenAIGoogleAnthropicMetaZ.aiAlibabaxAIThinking MachinesDeepSeekon the frontier
0.200.300.400.500.600.70$0.03$0.1$0.3$1$3$10$30GeminiMuseGemma onGemma offAstraFableOpusSolGLMQwen offGrokQwen onInklingDeepSeek0.200.300.400.500.600.700.80$0.01$0.03$0.1$0.3$1$3$10$30FableSolGeminiMuseGemma onGemma offAstraOpusGrokQwen offQwen onGLMInklingDeepSeek0.100.200.300.400.500.60$0.03$0.1$0.3$1$3GeminiMuseDeepSeekGemma onGemma offAstraOpusFableSolGLMQwen onQwen offInklingGrokestimated cost per case (US$, log scale)score +

Cost is the mean reported cost of one run, estimated from its token usage at the provider's list price (paper Figure 1b; provisional).

+
Ranking

Overall = ½ Text-to-Scene + ½ Image-to-Scene over the 95 public cases; its cost is weighted the same way.

#Agent configurationOverallText-to-SceneImage-to-SceneCost / case
1GPT-6 Astra (max)OpenAI
0.619
0.7240.515$12.90
2Gemini 3.8 Flash (high)ParetoGoogle
0.619
0.6570.581$1.92
3Claude Fable 5.1 (max)Anthropic
0.606
0.7880.424$10.62
4Claude Opus 5 (max)Anthropic
0.593
0.7180.468$19.48
5GPT-5.6 Sol (high)OpenAI
0.550
0.7070.393$2.97
6Muse Spark 1.3 (medium)ParetoMeta
0.502
0.6460.358$0.09
7GLM-5.3 Flash (max)open weightsZ.ai
0.415
0.5090.320$0.40
8Qwen 3.8 27B (thinking off)open weightsAlibaba
0.378
0.5570.200$0.38
9Grok 4.6 (high)xAI
0.376
0.5670.184$2.29
10Qwen 3.8 27B (thinking on)open weightsAlibaba
0.371
0.5160.226$0.30
11Inkling (high)Thinking Machines
0.312
0.4250.198$0.48
12Gemma 4 31B (thinking on)open weightsParetoGoogle
0.298
0.4710.126$0.03
13Gemma 4 31B (thinking off)open weightsParetoGoogle
0.285
0.4550.115$0.02
14DeepSeek V4.1 Flash (high)open weightsDeepSeek
0.227
0.2430.210$0.12

20 public cases. Case score = 0.2 × Detailed + 0.6 × Overview + 0.2 × Physical. Detailed and Overview are means of the per-case scores.

#Agent configurationScoreDetailedOverviewPhysicalCost / case
1Claude Fable 5.1 (max)ParetoAnthropic
0.788
0.7250.7720.898$18.16
2GPT-6 Astra (max)OpenAI
0.724
0.7070.7330.712$22.47
3Claude Opus 5 (max)Anthropic
0.718
0.6310.7160.811$34.50
4GPT-5.6 Sol (high)ParetoOpenAI
0.707
0.6320.7100.772$4.07
5Gemini 3.8 Flash (high)ParetoGoogle
0.657
0.6370.6640.655$2.49
6Muse Spark 1.3 (medium)ParetoMeta
0.646
0.5980.6610.651$0.10
7Grok 4.6 (high)xAI
0.567
0.5220.5760.585$4.33
8Qwen 3.8 27B (thinking off)open weightsAlibaba
0.557
0.5070.5300.688$0.49
9Qwen 3.8 27B (thinking on)open weightsAlibaba
0.516
0.5050.4600.694$0.33
10GLM-5.3 Flash (max)open weightsZ.ai
0.509
0.4030.5370.533$0.71
11Gemma 4 31B (thinking on)open weightsParetoGoogle
0.471
0.4120.4160.694$0.03
12Gemma 4 31B (thinking off)open weightsParetoGoogle
0.455
0.4380.3700.725$0.02
13Inkling (high)Thinking Machines
0.425
0.3930.3370.722$0.20
14DeepSeek V4.1 Flash (high)open weightsDeepSeek
0.243
0.2170.1610.513$0.19

75 public cases (25 indoor, 50 outdoor). Case score = 0.8 × Repair F1 + 0.2 × Physical. Repair F1 counts target actors restored to the withheld ground truth.

#Agent configurationScoreRepair F1PhysicalIndoorOutdoorCost / case
1Gemini 3.8 Flash (high)ParetoGoogle
0.581
0.5270.7960.7170.513$1.35
2GPT-6 Astra (max)OpenAI
0.515
0.4450.7960.7850.380$3.32
3Claude Opus 5 (max)Anthropic
0.468
0.3890.7860.6180.393$4.46
4Claude Fable 5.1 (max)Anthropic
0.424
0.3320.7950.5350.369$3.09
5GPT-5.6 Sol (high)OpenAI
0.393
0.3190.6890.4970.341$1.88
6Muse Spark 1.3 (medium)ParetoMeta
0.358
0.2460.8070.6330.221$0.07
7GLM-5.3 Flash (max)open weightsZ.ai
0.320
0.1880.8480.4650.247$0.09
8Qwen 3.8 27B (thinking on)open weightsAlibaba
0.226
0.1180.6600.2970.191$0.26
9DeepSeek V4.1 Flash (high)open weightsParetoDeepSeek
0.210
0.0660.7870.1890.221$0.05
10Qwen 3.8 27B (thinking off)open weightsAlibaba
0.200
0.0830.6670.2660.166$0.28
11Inkling (high)Thinking Machines
0.198
0.0950.6090.1690.213$0.75
12Grok 4.6 (high)xAI
0.184
0.0440.7440.1970.178$0.26
13Gemma 4 31B (thinking on)open weightsParetoGoogle
0.126
0.0030.6150.1270.125$0.03
14Gemma 4 31B (thinking off)open weightsParetoGoogle
0.115
0.0010.5720.1120.116$0.02
+

Click a column to sort. Pareto: on the score-against-cost frontier.

+
+

Takeaways

+

What the evaluation shows across the 14 agents. Every chart below is drawn from the paper's numbers; hover for values.

+
Takeaway 1

Construction and editing probe different capabilities: scene-level spatial reasoning and precise control of scene state.

Despite nearly identical overall scores, Astra is stronger at construction and Gemini at editing.

Text-to-Scene (construction)Image-to-Scene (editing)
← Text-to-SceneImage-to-Scene →0.7240.515GPT-6 Astra0.6570.581Gemini 3.8 Flash0.7880.424Claude Fable 5.10.7180.468Claude Opus 50.7070.393GPT-5.6 Sol0.6460.358Muse Spark 1.30.5090.320GLM-5.3 Flash0.5570.200Qwen 3.8 27B · off0.5670.184Grok 4.60.5160.226Qwen 3.8 27B · on0.4250.198Inkling0.4710.126Gemma 4 31B · on0.4550.115Gemma 4 31B · off0.2430.210DeepSeek V4.1 Flash
Text-to-Scene (left) and Image-to-Scene (right) score per agent, in overall order (paper Table 2). Astra and Gemini tie at 0.619 overall.
Takeaway 2

Spatial Composition remains the weakest requirement family for all 14 agents.

Errors persist even when the required objects are present: generating the right objects does not ensure that their relationships satisfy the specification.

Requirement-family score per agent
Identity & EnvironmentContent & QuantityAttributes & MaterialsSpatial Composition
0.20.40.60.81.0GPT-6 AstraGemini 3.8 FlashClaude Fable 5.1Claude Opus 5GPT-5.6 SolMuse Spark 1.3GLM-5.3 FlashQwen 3.8 27B · offGrok 4.6Qwen 3.8 27B · onInklingGemma 4 31B · onGemma 4 31B · offDeepSeek V4.1 FlashMean of 140.360.67
Each row is one agent's four family scores; the orange dot, Spatial Composition, is the lowest in every row (paper Figure 14b).
The objects are there; the relation is not
The judge's view of GPT-5.6 Sol's Egyptian Temple roof
✓ slab roof 0.95✓ stone coping 0.82✗ slab roof enclosed by the coping 0.78
The judge's own frame of GPT-5.6 Sol on Egyptian Temple. It finds the slab roof and the stone coping, but not the roof enclosed by the coping: the coping is separate low wall segments with visible gaps (ringed).
Mismatch rate when the required objects are present
Spatial relation32.0%Distribution19.1%Composition4.3%
Across the benchmark (paper Figure 6a).
Takeaway 3

Editing ability varies substantially across repair types.

Astra performs best on Transform repairs, while Gemini is stronger on Lifecycle and Layout operations.

GPT-6 AstraGemini 3.8 FlashClaude Fable 5.1Claude Opus 5Restore0.3240.3760.2440.388Lifecycle0.5770.7150.5540.593Layout0.5960.7030.4640.512Symmetry0.4660.5440.5550.359Transform0.6130.5320.3530.406
Mean Image-to-Scene score by repair type for four agents; the outlined cell is the best of the four in its row (paper Figure 6b).
Takeaway 4

Recovering the target does not guarantee precise editing.

Across all 14 configurations, 35.8% of complete recoveries still contain unintended changes to the surrounding scene.

1,050Image-to-Scene evaluations (14 agents × 75 cases)
212recovered every target completely
76 · 35.8%of those still changed the surrounding scene
unintended editsclean complete recoveryno complete recovery
One dot per Image-to-Scene evaluation (paper Figure 6c).
+

How Code4Scene scores a scene

The leaderboard, one scene through the evaluator, two scenes compared, and what the benchmark finds, in 90 seconds.

-
-
-

Score the scene, not the code or a render

-

Frontier coding agents can now write and execute code that authors 3D environments, but whether they reliably +

+

About Code4Scene

Benchmarking Coding Agents for Constructing and Editing 3D Scenes

Xiaokang Ye1,*, Siddhant Hitesh Mantri1,*, +Zimeng Chen1,*, Edward Zhang1, Zhaoxu Zheng1,
+Yuanheng Li2, Yizhao Chen1, Tianyang Huang1, +Lianhui Qin1,†

1University of California San Diego    2University of California, Berkeley    +*Equal contribution    †Corresponding author

Paper abstract 190-case benchmark · 95-case public evaluation

Frontier coding agents can now write and execute code that authors 3D environments, but whether they reliably understand 3D structure and precisely control scene state remains unclear. The generated 3D scene is a persistent, executable artifact: a convincing render can hide incorrect spatial relations, intersecting objects, or unintended modifications. We introduce Code4Scene, a benchmark of 190 Unreal Engine cases built from human-assembled scenes that evaluates coding agents on two complementary settings under a shared execution interface. @@ -58,39 +69,7 @@ set, construction and editing performance are strongly correlated but not interc Gemini 3.8 Flash leads editing, and GPT-6 Astra narrowly leads overall. Spatial Composition is the weakest construction category for every agent, while editing remains imprecise: the best Repair F1 is only 0.527, and 35.8% of edits that fully recover the target still introduce unintended changes elsewhere in the scene. These results expose a gap between plausible 3D generation and reliable spatial reasoning and state control.

-

Read the paper on arXiv →

-
1 · InputPrompt or referenceAn open-ended scene description, or a corrupted scene with reference images of the original.
2 · AgentCoding agentA frontier or open-weight model in its own CLI harness, with the Unreal Engine editor as a tool.
3 · CodeCode in the engineIt writes and runs Python against the editor: spawns, moves and edits actors, inspects, revises.
4 · SceneEngine-native sceneThe saved level (.umap): every actor, transform and material, as the engine holds it.
5 · EvaluatorScored in the engineRequirements located and judged, overview views judged, physics measured, edits diffed against the withheld ground truth.
-
-
Setting 1 · construction

Text-to-Scene

-

From an empty level and the content pack's asset catalog, the agent builds the scene a prompt describes. Many realizations are valid, so the -evaluator judges requirements rather than matching a reference. 20 public cases.

-
case score = 0.2 × Detailed Alignment + 0.6 × Overview Alignment + 0.2 × Physical Safety
-
Setting 2 · editing

Image-to-Scene

-

The agent gets a corrupted copy of a human-assembled scene and reference views of the original. It must restore every target actor (within -5 cm, 5° and 5%) and change nothing else. 75 public cases: 25 indoor, 50 outdoor.

-
case score = 0.8 × Repair F1 + 0.2 × Physical Safety
-
-

14 coding agents on the public set

-

Scores are from the paper's Table 2. Pick a setting; the chart and the table follow.

-
-
Score against cost
-
Each agent's score against what one case cost it. The line is the frontier: no agent left of it scores higher. Hover a point for its numbers.
-
OpenAIGoogleAnthropicMetaZ.aiAlibabaxAIThinking MachinesDeepSeekon the frontier
0.200.300.400.500.600.70$0.03$0.1$0.3$1$3$10$30GeminiMuseGemma onGemma offAstraFableOpusSolGLMQwen offGrokQwen onInklingDeepSeek0.200.300.400.500.600.700.80$0.01$0.03$0.1$0.3$1$3$10$30FableSolGeminiMuseGemma onGemma offAstraOpusGrokQwen offQwen onGLMInklingDeepSeek0.100.200.300.400.500.60$0.03$0.1$0.3$1$3GeminiMuseDeepSeekGemma onGemma offAstraOpusFableSolGLMQwen onQwen offInklingGrokestimated cost per case (US$, log scale)score -

Cost is the mean reported cost of one run, estimated from its token usage at the provider's list price (paper Figure 1b; provisional).

-
Ranking

Overall = ½ Text-to-Scene + ½ Image-to-Scene over the 95 public cases; its cost is weighted the same way.

#Agent configurationOverallText-to-SceneImage-to-SceneCost / case
1GPT-6 Astra (max)OpenAI
0.619
0.7240.515$12.90
2Gemini 3.8 Flash (high)ParetoGoogle
0.619
0.6570.581$1.92
3Claude Fable 5.1 (max)Anthropic
0.606
0.7880.424$10.62
4Claude Opus 5 (max)Anthropic
0.593
0.7180.468$19.48
5GPT-5.6 Sol (high)OpenAI
0.550
0.7070.393$2.97
6Muse Spark 1.3 (medium)ParetoMeta
0.502
0.6460.358$0.09
7GLM-5.3 Flash (max)open weightsZ.ai
0.415
0.5090.320$0.40
8Qwen 3.8 27B (thinking off)open weightsAlibaba
0.378
0.5570.200$0.38
9Grok 4.6 (high)xAI
0.376
0.5670.184$2.29
10Qwen 3.8 27B (thinking on)open weightsAlibaba
0.371
0.5160.226$0.30
11Inkling (high)Thinking Machines
0.312
0.4250.198$0.48
12Gemma 4 31B (thinking on)open weightsParetoGoogle
0.298
0.4710.126$0.03
13Gemma 4 31B (thinking off)open weightsParetoGoogle
0.285
0.4550.115$0.02
14DeepSeek V4.1 Flash (high)open weightsDeepSeek
0.227
0.2430.210$0.12

20 public cases. Case score = 0.2 × Detailed + 0.6 × Overview + 0.2 × Physical. Detailed and Overview are means of the per-case scores.

#Agent configurationScoreDetailedOverviewPhysicalCost / case
1Claude Fable 5.1 (max)ParetoAnthropic
0.788
0.7250.7720.898$18.16
2GPT-6 Astra (max)OpenAI
0.724
0.7070.7330.712$22.47
3Claude Opus 5 (max)Anthropic
0.718
0.6310.7160.811$34.50
4GPT-5.6 Sol (high)ParetoOpenAI
0.707
0.6320.7100.772$4.07
5Gemini 3.8 Flash (high)ParetoGoogle
0.657
0.6370.6640.655$2.49
6Muse Spark 1.3 (medium)ParetoMeta
0.646
0.5980.6610.651$0.10
7Grok 4.6 (high)xAI
0.567
0.5220.5760.585$4.33
8Qwen 3.8 27B (thinking off)open weightsAlibaba
0.557
0.5070.5300.688$0.49
9Qwen 3.8 27B (thinking on)open weightsAlibaba
0.516
0.5050.4600.694$0.33
10GLM-5.3 Flash (max)open weightsZ.ai
0.509
0.4030.5370.533$0.71
11Gemma 4 31B (thinking on)open weightsParetoGoogle
0.471
0.4120.4160.694$0.03
12Gemma 4 31B (thinking off)open weightsParetoGoogle
0.455
0.4380.3700.725$0.02
13Inkling (high)Thinking Machines
0.425
0.3930.3370.722$0.20
14DeepSeek V4.1 Flash (high)open weightsDeepSeek
0.243
0.2170.1610.513$0.19

75 public cases (25 indoor, 50 outdoor). Case score = 0.8 × Repair F1 + 0.2 × Physical. Repair F1 counts target actors restored to the withheld ground truth.

#Agent configurationScoreRepair F1PhysicalIndoorOutdoorCost / case
1Gemini 3.8 Flash (high)ParetoGoogle
0.581
0.5270.7960.7170.513$1.35
2GPT-6 Astra (max)OpenAI
0.515
0.4450.7960.7850.380$3.32
3Claude Opus 5 (max)Anthropic
0.468
0.3890.7860.6180.393$4.46
4Claude Fable 5.1 (max)Anthropic
0.424
0.3320.7950.5350.369$3.09
5GPT-5.6 Sol (high)OpenAI
0.393
0.3190.6890.4970.341$1.88
6Muse Spark 1.3 (medium)ParetoMeta
0.358
0.2460.8070.6330.221$0.07
7GLM-5.3 Flash (max)open weightsZ.ai
0.320
0.1880.8480.4650.247$0.09
8Qwen 3.8 27B (thinking on)open weightsAlibaba
0.226
0.1180.6600.2970.191$0.26
9DeepSeek V4.1 Flash (high)open weightsParetoDeepSeek
0.210
0.0660.7870.1890.221$0.05
10Qwen 3.8 27B (thinking off)open weightsAlibaba
0.200
0.0830.6670.2660.166$0.28
11Inkling (high)Thinking Machines
0.198
0.0950.6090.1690.213$0.75
12Grok 4.6 (high)xAI
0.184
0.0440.7440.1970.178$0.26
13Gemma 4 31B (thinking on)open weightsParetoGoogle
0.126
0.0030.6150.1270.125$0.03
14Gemma 4 31B (thinking off)open weightsParetoGoogle
0.115
0.0010.5720.1120.116$0.02
-

Click a column to sort. Pareto: on the score-against-cost frontier.

-
-

Takeaways

-

What the evaluation shows across the 14 agents. Every chart below is drawn from the paper's numbers; hover for values.

-
Takeaway 1

Construction and editing probe different capabilities: scene-level spatial reasoning and precise control of scene state.

Despite nearly identical overall scores, Astra is stronger at construction and Gemini at editing.

Text-to-Scene (construction)Image-to-Scene (editing)
← Text-to-SceneImage-to-Scene →0.7240.515GPT-6 Astra0.6570.581Gemini 3.8 Flash0.7880.424Claude Fable 5.10.7180.468Claude Opus 50.7070.393GPT-5.6 Sol0.6460.358Muse Spark 1.30.5090.320GLM-5.3 Flash0.5570.200Qwen 3.8 27B · off0.5670.184Grok 4.60.5160.226Qwen 3.8 27B · on0.4250.198Inkling0.4710.126Gemma 4 31B · on0.4550.115Gemma 4 31B · off0.2430.210DeepSeek V4.1 Flash
Text-to-Scene (left) and Image-to-Scene (right) score per agent, in overall order (paper Table 2). Astra and Gemini tie at 0.619 overall.
Takeaway 2

Spatial Composition remains the weakest requirement family for all 14 agents.

Errors persist even when the required objects are present: generating the right objects does not ensure that their relationships satisfy the specification.

Requirement-family score per agent
Identity & EnvironmentContent & QuantityAttributes & MaterialsSpatial Composition
0.20.40.60.81.0GPT-6 AstraGemini 3.8 FlashClaude Fable 5.1Claude Opus 5GPT-5.6 SolMuse Spark 1.3GLM-5.3 FlashQwen 3.8 27B · offGrok 4.6Qwen 3.8 27B · onInklingGemma 4 31B · onGemma 4 31B · offDeepSeek V4.1 FlashMean of 140.360.67
Each row is one agent's four family scores; the orange dot, Spatial Composition, is the lowest in every row (paper Figure 14b).
The objects are there; the relation is not
The judge's view of GPT-5.6 Sol's Egyptian Temple roof
✓ slab roof 0.95✓ stone coping 0.82✗ slab roof enclosed by the coping 0.78
The judge's own frame of GPT-5.6 Sol on Egyptian Temple. It finds the slab roof and the stone coping, but not the roof enclosed by the coping: the coping is separate low wall segments with visible gaps (ringed).
Mismatch rate when the required objects are present
Spatial relation32.0%Distribution19.1%Composition4.3%
Across the benchmark (paper Figure 6a).
Takeaway 3

Editing ability varies substantially across repair types.

Astra performs best on Transform repairs, while Gemini is stronger on Lifecycle and Layout operations.

GPT-6 AstraGemini 3.8 FlashClaude Fable 5.1Claude Opus 5Restore0.3240.3760.2440.388Lifecycle0.5770.7150.5540.593Layout0.5960.7030.4640.512Symmetry0.4660.5440.5550.359Transform0.6130.5320.3530.406
Mean Image-to-Scene score by repair type for four agents; the outlined cell is the best of the four in its row (paper Figure 6b).
Takeaway 4

Recovering the target does not guarantee precise editing.

Across all 14 configurations, 35.8% of complete recoveries still contain unintended changes to the surrounding scene.

1,050Image-to-Scene evaluations (14 agents × 75 cases)
212recovered every target completely
76 · 35.8%of those still changed the surrounding scene
unintended editsclean complete recoveryno complete recovery
One dot per Image-to-Scene evaluation (paper Figure 6c).
1 · InputPrompt or referenceAn open-ended scene description, or a corrupted scene with reference images of the original.
2 · AgentCoding agentA frontier or open-weight model in its own CLI harness, with the Unreal Engine editor as a tool.
3 · CodeCode in the engineIt writes and runs Python against the editor: spawns, moves and edits actors, inspects, revises.
4 · SceneEngine-native sceneThe saved level (.umap): every actor, transform and material, as the engine holds it.
5 · EvaluatorScored in the engineRequirements located and judged, overview views judged, physics measured, edits diffed against the withheld ground truth.

Cite This Work

@misc{ye2026code4scene,
   title         = {Code4Scene: Benchmarking Coding Agents for Constructing and Editing 3D Scenes},
@@ -102,12 +81,13 @@ evaluator judges requirements rather than matching a reference. 20 public cases.
   url           = {https://arxiv.org/abs/2609.36777},
 }
- arXiv - GitHub Repository -SimWorld Main Site
-
-