Code4Scene:
Benchmarking Coding Agents for
Constructing and Editing 3D Scenes

190 Unreal Engine cases built from human-assembled scenes. Coding agents write and run code that builds a scene from text, or repairs one from reference images, and Code4Scene scores the engine-native scene they save, not their code or a rendered view.

Xiaokang Ye1,*, Siddhant Hitesh Mantri1,*, Zimeng Chen1,*, Edward Zhang1, Zhaoxu Zheng1,
Yuanheng Li2, Yizhao Chen1, Tianyang Huang1, Lianhui Qin1,†

1University of California San Diego    2University of California, Berkeley    *Equal contribution    †Corresponding author

190
Unreal Engine cases
95 public, scored here
14
coding-agent configurations
frontier and open weights
0.619
best overall score
Astra and Gemini, tied
35.8%
of complete repairs
still change the rest of the scene

How Code4Scene scores a scene

The leaderboard, one scene through the evaluator, two scenes compared, and what the benchmark finds, in 90 seconds.

Score the scene, not the code or a render

Coding agents can now operate 3D engines: they write and run code, inspect the result and revise the scene. Code4Scene is a benchmark of 190 Unreal Engine cases built from human-assembled scenes that evaluates them on two complementary settings under a shared execution interface. Construction starts from an empty level and an open-ended text specification. Editing gives the agent a corrupted copy of a scene and reference images of the original, and asks it to restore the intended state while leaving everything else untouched.

Rather than scoring code or rendered views, Code4Scene evaluates the engine-native scene the agent saved, on task fulfilment, artifact integrity and static physical validity, and compares edits against withheld ground truth. Across 14 agent configurations, construction and editing rank the agents differently, spatial relations remain the weakest requirement family, and a third of complete repairs still change the surrounding scene.

1 · InputPrompt or referenceAn open-ended scene description, or a corrupted scene with reference images of the original.
2 · AgentCoding agentA frontier or open-weight model in its own CLI harness, with the Unreal Engine editor as a tool.
3 · CodeCode in the engineIt writes and runs Python against the editor: spawns, moves and edits actors, inspects, revises.
4 · SceneEngine-native sceneThe saved level (.umap): every actor, transform and material, as the engine holds it.
5 · EvaluatorScored in the engineRequirements located and judged, overview views judged, physics measured, edits diffed against the withheld ground truth.
Setting 1 · construction

Text-to-Scene

From an empty level and the content pack's asset catalog, the agent builds the scene a prompt describes. Many realizations are valid, so the evaluator judges requirements rather than matching a reference. 20 public cases.

case score = 0.2 × Detailed Alignment + 0.6 × Overview Alignment + 0.2 × Physical Safety
Setting 2 · editing

Image-to-Scene

The agent gets a corrupted copy of a human-assembled scene and reference views of the original. It must restore every target actor (within 5 cm, 5° and 5%) and change nothing else. 75 public cases: 25 indoor, 50 outdoor.

case score = 0.8 × Repair F1 + 0.2 × Physical Safety

14 coding agents on the public set

Scores are from the paper's Table 2. Pick a setting; the chart and the table follow.

Score against cost
Each agent's score against what one case cost it. The line is the frontier: no agent left of it scores higher. Hover a point for its numbers.
OpenAIGoogleAnthropicMetaZ.aiAlibabaxAIThinking MachinesDeepSeekon the frontier
0.200.300.400.500.600.70$0.03$0.1$0.3$1$3$10$30GeminiMuseGemma onGemma offAstraFableOpusSolGLMQwen offGrokQwen onInklingDeepSeek0.200.300.400.500.600.700.80$0.01$0.03$0.1$0.3$1$3$10$30FableSolGeminiMuseGemma onGemma offAstraOpusGrokQwen offQwen onGLMInklingDeepSeek0.100.200.300.400.500.60$0.03$0.1$0.3$1$3GeminiMuseDeepSeekGemma onGemma offAstraOpusFableSolGLMQwen onQwen offInklingGrokestimated cost per case (US$, log scale)score

Cost is the mean reported cost of one run, estimated from its token usage at the provider's list price (paper Figure 1b; provisional).

Ranking

Overall = ½ Text-to-Scene + ½ Image-to-Scene over the 95 public cases; its cost is weighted the same way.

#Agent configurationOverallText-to-SceneImage-to-SceneCost / case
1GPT-6 Astra (max)OpenAI
0.619
0.7240.515$12.90
2Gemini 3.8 Flash (high)ParetoGoogle
0.619
0.6570.581$1.92
3Claude Fable 5.1 (max)Anthropic
0.606
0.7880.424$10.62
4Claude Opus 5 (max)Anthropic
0.593
0.7180.468$19.48
5GPT-5.6 Sol (high)OpenAI
0.550
0.7070.393$2.97
6Muse Spark 1.3 (medium)ParetoMeta
0.502
0.6460.358$0.09
7GLM-5.3 Flash (max)open weightsZ.ai
0.415
0.5090.320$0.40
8Qwen 3.8 27B (thinking off)open weightsAlibaba
0.378
0.5570.200$0.38
9Grok 4.6 (high)xAI
0.376
0.5670.184$2.29
10Qwen 3.8 27B (thinking on)open weightsAlibaba
0.371
0.5160.226$0.30
11Inkling (high)Thinking Machines
0.312
0.4250.198$0.48
12Gemma 4 31B (thinking on)open weightsParetoGoogle
0.298
0.4710.126$0.03
13Gemma 4 31B (thinking off)open weightsParetoGoogle
0.285
0.4550.115$0.02
14DeepSeek V4.1 Flash (high)open weightsDeepSeek
0.227
0.2430.210$0.12

20 public cases. Case score = 0.2 × Detailed + 0.6 × Overview + 0.2 × Physical. Detailed and Overview are means of the per-case scores.

#Agent configurationScoreDetailedOverviewPhysicalCost / case
1Claude Fable 5.1 (max)ParetoAnthropic
0.788
0.7250.7720.898$18.16
2GPT-6 Astra (max)OpenAI
0.724
0.7070.7330.712$22.47
3Claude Opus 5 (max)Anthropic
0.718
0.6310.7160.811$34.50
4GPT-5.6 Sol (high)ParetoOpenAI
0.707
0.6320.7100.772$4.07
5Gemini 3.8 Flash (high)ParetoGoogle
0.657
0.6370.6640.655$2.49
6Muse Spark 1.3 (medium)ParetoMeta
0.646
0.5980.6610.651$0.10
7Grok 4.6 (high)xAI
0.567
0.5220.5760.585$4.33
8Qwen 3.8 27B (thinking off)open weightsAlibaba
0.557
0.5070.5300.688$0.49
9Qwen 3.8 27B (thinking on)open weightsAlibaba
0.516
0.5050.4600.694$0.33
10GLM-5.3 Flash (max)open weightsZ.ai
0.509
0.4030.5370.533$0.71
11Gemma 4 31B (thinking on)open weightsParetoGoogle
0.471
0.4120.4160.694$0.03
12Gemma 4 31B (thinking off)open weightsParetoGoogle
0.455
0.4380.3700.725$0.02
13Inkling (high)Thinking Machines
0.425
0.3930.3370.722$0.20
14DeepSeek V4.1 Flash (high)open weightsDeepSeek
0.243
0.2170.1610.513$0.19

75 public cases (25 indoor, 50 outdoor). Case score = 0.8 × Repair F1 + 0.2 × Physical. Repair F1 counts target actors restored to the withheld ground truth.

#Agent configurationScoreRepair F1PhysicalIndoorOutdoorCost / case
1Gemini 3.8 Flash (high)ParetoGoogle
0.581
0.5270.7960.7170.513$1.35
2GPT-6 Astra (max)OpenAI
0.515
0.4450.7960.7850.380$3.32
3Claude Opus 5 (max)Anthropic
0.468
0.3890.7860.6180.393$4.46
4Claude Fable 5.1 (max)Anthropic
0.424
0.3320.7950.5350.369$3.09
5GPT-5.6 Sol (high)OpenAI
0.393
0.3190.6890.4970.341$1.88
6Muse Spark 1.3 (medium)ParetoMeta
0.358
0.2460.8070.6330.221$0.07
7GLM-5.3 Flash (max)open weightsZ.ai
0.320
0.1880.8480.4650.247$0.09
8Qwen 3.8 27B (thinking on)open weightsAlibaba
0.226
0.1180.6600.2970.191$0.26
9DeepSeek V4.1 Flash (high)open weightsParetoDeepSeek
0.210
0.0660.7870.1890.221$0.05
10Qwen 3.8 27B (thinking off)open weightsAlibaba
0.200
0.0830.6670.2660.166$0.28
11Inkling (high)Thinking Machines
0.198
0.0950.6090.1690.213$0.75
12Grok 4.6 (high)xAI
0.184
0.0440.7440.1970.178$0.26
13Gemma 4 31B (thinking on)open weightsParetoGoogle
0.126
0.0030.6150.1270.125$0.03
14Gemma 4 31B (thinking off)open weightsParetoGoogle
0.115
0.0010.5720.1120.116$0.02

Click a column to sort. Pareto: on the score-against-cost frontier.

Takeaways

What the evaluation shows across the 14 agents. Every chart below is drawn from the paper's numbers; hover for values.

Takeaway 1

Construction and editing probe different capabilities: scene-level spatial reasoning and precise control of scene state.

Despite nearly identical overall scores, Astra is stronger at construction and Gemini at editing.

Text-to-Scene (construction)Image-to-Scene (editing)
← Text-to-SceneImage-to-Scene →0.7240.515GPT-6 Astra0.6570.581Gemini 3.8 Flash0.7880.424Claude Fable 5.10.7180.468Claude Opus 50.7070.393GPT-5.6 Sol0.6460.358Muse Spark 1.30.5090.320GLM-5.3 Flash0.5570.200Qwen 3.8 27B · off0.5670.184Grok 4.60.5160.226Qwen 3.8 27B · on0.4250.198Inkling0.4710.126Gemma 4 31B · on0.4550.115Gemma 4 31B · off0.2430.210DeepSeek V4.1 Flash
Text-to-Scene (left) and Image-to-Scene (right) score per agent, in overall order (paper Table 2). Astra and Gemini tie at 0.619 overall.
Takeaway 2

Spatial Composition remains the weakest requirement family for all 14 agents.

Errors persist even when the required objects are present: generating the right objects does not ensure that their relationships satisfy the specification.

Requirement-family score per agent
Identity & EnvironmentContent & QuantityAttributes & MaterialsSpatial Composition
0.20.40.60.81.0GPT-6 AstraGemini 3.8 FlashClaude Fable 5.1Claude Opus 5GPT-5.6 SolMuse Spark 1.3GLM-5.3 FlashQwen 3.8 27B · offGrok 4.6Qwen 3.8 27B · onInklingGemma 4 31B · onGemma 4 31B · offDeepSeek V4.1 FlashMean of 140.360.67
Each row is one agent's four family scores; the orange dot, Spatial Composition, is the lowest in every row (paper Figure 14b).
The objects are there; the relation is not
The judge's view of GPT-5.6 Sol's Egyptian Temple roof
✓ slab roof 0.95✓ stone coping 0.82✗ slab roof enclosed by the coping 0.78
The judge's own frame of GPT-5.6 Sol on Egyptian Temple. It finds the slab roof and the stone coping, but not the roof enclosed by the coping: the coping is separate low wall segments with visible gaps (ringed).
Mismatch rate when the required objects are present
Spatial relation32.0%Distribution19.1%Composition4.3%
Across the benchmark (paper Figure 6a).
Takeaway 3

Editing ability varies substantially across repair types.

Astra performs best on Transform repairs, while Gemini is stronger on Lifecycle and Layout operations.

GPT-6 AstraGemini 3.8 FlashClaude Fable 5.1Claude Opus 5Restore0.3240.3760.2440.388Lifecycle0.5770.7150.5540.593Layout0.5960.7030.4640.512Symmetry0.4660.5440.5550.359Transform0.6130.5320.3530.406
Mean Image-to-Scene score by repair type for four agents; the outlined cell is the best of the four in its row (paper Figure 6b).
Takeaway 4

Recovering the target does not guarantee precise editing.

Across all 14 configurations, 35.8% of complete recoveries still contain unintended changes to the surrounding scene.

1,050Image-to-Scene evaluations (14 agents × 75 cases)
212recovered every target completely
76 · 35.8%of those still changed the surrounding scene
unintended editsclean complete recoveryno complete recovery
One dot per Image-to-Scene evaluation (paper Figure 6c).

Turn the scenes yourself

For each selected case, every agent's saved scene is exported from Unreal Engine and opens in the browser in 3D, beside the evaluator's scores for it.

Open the case explorer

Cite This Work

@article{code4scene2026,
  title   = {Code4Scene: Benchmarking Coding Agents for Constructing and Editing 3D Scenes},
  author  = {Ye, Xiaokang and Mantri, Siddhant Hitesh and Chen, Zimeng and Zhang, Edward and Zheng, Zhaoxu and Li, Yuanheng and Chen, Yizhao and Huang, Tianyang and Qin, Lianhui},
  year    = {2026}
}
GitHub Repository SimWorld Main Site