diff --git "a/index.html" "b/index.html" --- "a/index.html" +++ "b/index.html" @@ -1,19 +1,107 @@ - - - - - My static Space - - - -
-

Welcome to your static Space!

-

You can modify this app directly by editing index.html in the Files and versions tab.

-

- Also don't forget to check the - Spaces documentation. -

-
- - + + +Code4Scene: Benchmarking Coding Agents for Constructing and Editing 3D Scenes + + + + + + +
+

Code4Scene:
Benchmarking Coding Agents for
Constructing and Editing 3D Scenes

+

190 Unreal Engine cases built from human-assembled scenes. Coding agents write and run code that builds a scene from text, or +repairs one from reference images, and Code4Scene scores the engine-native scene they save, not their code or a rendered view.

+

Xiaokang Ye1,*, Siddhant Hitesh Mantri1,*, +Zimeng Chen1,*, Edward Zhang1, Zhaoxu Zheng1,
+Yuanheng Li2, Yizhao Chen1, Tianyang Huang1, +Lianhui Qin1,†

+

1University of California San Diego    2University of California, Berkeley    +*Equal contribution    †Corresponding author

+ +
+
190
Unreal Engine cases
95 public, scored here
+
14
coding-agent configurations
frontier and open weights
+
0.619
best overall score
Astra and Gemini, tied
+
35.8%
of complete repairs
still change the rest of the scene
+
+

How Code4Scene scores a scene

+

The leaderboard, one scene through the evaluator, two scenes compared, and what the benchmark finds, in 90 seconds.

+
+
+

Score the scene, not the code or a render

+

Coding agents can now operate 3D engines: they write and run code, inspect the result and revise the scene. Code4Scene is a +benchmark of 190 Unreal Engine cases built from human-assembled scenes that evaluates them on two complementary settings under a shared execution +interface. Construction starts from an empty level and an open-ended text specification. Editing gives the agent a corrupted copy of a +scene and reference images of the original, and asks it to restore the intended state while leaving everything else untouched.

+

Rather than scoring code or rendered views, Code4Scene evaluates the engine-native scene the agent saved, on task fulfilment, artifact integrity +and static physical validity, and compares edits against withheld ground truth. Across 14 agent configurations, construction and editing rank the +agents differently, spatial relations remain the weakest requirement family, and a third of complete repairs still change the surrounding scene.

+
1 · InputPrompt or referenceAn open-ended scene description, or a corrupted scene with reference images of the original.
2 · AgentCoding agentA frontier or open-weight model in its own CLI harness, with the Unreal Engine editor as a tool.
3 · CodeCode in the engineIt writes and runs Python against the editor: spawns, moves and edits actors, inspects, revises.
4 · SceneEngine-native sceneThe saved level (.umap): every actor, transform and material, as the engine holds it.
5 · EvaluatorScored in the engineRequirements located and judged, overview views judged, physics measured, edits diffed against the withheld ground truth.
+
+
Setting 1 · construction

Text-to-Scene

+

From an empty level and the content pack's asset catalog, the agent builds the scene a prompt describes. Many realizations are valid, so the +evaluator judges requirements rather than matching a reference. 20 public cases.

+
case score = 0.2 × Detailed Alignment + 0.6 × Overview Alignment + 0.2 × Physical Safety
+
Setting 2 · editing

Image-to-Scene

+

The agent gets a corrupted copy of a human-assembled scene and reference views of the original. It must restore every target actor (within +5 cm, 5° and 5%) and change nothing else. 75 public cases: 25 indoor, 50 outdoor.

+
case score = 0.8 × Repair F1 + 0.2 × Physical Safety
+
+

14 coding agents on the public set

+

Overall = ½ Text-to-Scene + ½ Image-to-Scene, over the 95 public cases. Scores are from the paper's Table 2.

+
Score against cost
+
Each agent's score against what one case cost it. The line is the frontier: no agent to its left scores higher.
+
+0.000.150.300.450.600.750.90$0.01$0.03$0.1$0.3$1$3$10$30estimated cost per case (US$, log scale)scoreGeminiMuseGemma onGemma offAstraFableOpusSolGLMQwen offGrokQwen onInklingDeepSeekFableSolGeminiMuseGemma onGemma offAstraOpusGrokQwen offQwen onGLMInklingDeepSeekGeminiMuseDeepSeekGemma onGemma offAstraOpusFableSolGLMQwen onQwen offInklingGrok +

Overall

Gemini 3.8 Flash (high)0.619 · $1.59
Muse Spark 1.3 (medium)0.502 · $0.08
Gemma 4 31B (thinking on)0.298 · $0.03
Gemma 4 31B (thinking off)0.285 · $0.02

Text-to-Scene

Claude Fable 5.1 (max)0.788 · $18.16
GPT-5.6 Sol (high)0.707 · $4.07
Gemini 3.8 Flash (high)0.657 · $2.49
Muse Spark 1.3 (medium)0.646 · $0.10
Gemma 4 31B (thinking on)0.471 · $0.03
Gemma 4 31B (thinking off)0.455 · $0.02

Image-to-Scene

Gemini 3.8 Flash (high)0.581 · $1.35
Muse Spark 1.3 (medium)0.358 · $0.07
DeepSeek V4.1 Flash (high)0.210 · $0.05
Gemma 4 31B (thinking on)0.126 · $0.03
Gemma 4 31B (thinking off)0.115 · $0.02
+

Cost is the mean reported cost of one run, estimated from its token usage at the provider's list price (paper Figure 1b; provisional). +Some runs did not report usage; hover a point for how many of the 95 cases are priced.

+
Text-to-Scene · construction
+
20 public cases. Detailed and Overview are the means of the per-case scores; Score and Physical are from Table 2. Click a column to sort.
+
#Agent configurationScoreDetailedOverviewPhysicalCost / case
1Claude Fable 5.1 (max)ParetoAnthropic
0.788
0.7250.7720.898$18.16
2GPT-6 Astra (max)OpenAI
0.724
0.7070.7330.712$22.47
3Claude Opus 5 (max)Anthropic
0.718
0.6310.7160.811$34.50
4GPT-5.6 Sol (high)ParetoOpenAI
0.707
0.6320.7100.772$4.07
5Gemini 3.8 Flash (high)ParetoGoogle
0.657
0.6370.6640.655$2.49
6Muse Spark 1.3 (medium)ParetoMeta
0.646
0.5980.6610.651$0.10
7Grok 4.6 (high)xAI
0.567
0.5220.5760.585$4.33
8Qwen 3.8 27B (thinking off)open weightsAlibaba
0.557
0.5070.5300.688$0.49
9Qwen 3.8 27B (thinking on)open weightsAlibaba
0.516
0.5050.4600.694$0.33
10GLM-5.3 Flash (max)open weightsZ.ai
0.509
0.4030.5370.533$0.71
11Gemma 4 31B (thinking on)open weightsParetoGoogle
0.471
0.4120.4160.694$0.03
12Gemma 4 31B (thinking off)open weightsParetoGoogle
0.455
0.4380.3700.725$0.02
13Inkling (high)Thinking Machines
0.425
0.3930.3370.722$0.20
14DeepSeek V4.1 Flash (high)open weightsDeepSeek
0.243
0.2170.1610.513$0.19
+
Image-to-Scene · editing
+
75 public cases (25 indoor, 50 outdoor). Repair F1 counts the target actors restored to the withheld ground truth.
+
#Agent configurationScoreRepair F1PhysicalIndoorOutdoorCost / case
1Gemini 3.8 Flash (high)ParetoGoogle
0.581
0.5270.7960.7170.513$1.35
2GPT-6 Astra (max)OpenAI
0.515
0.4450.7960.7850.380$3.32
3Claude Opus 5 (max)Anthropic
0.468
0.3890.7860.6180.393$4.46
4Claude Fable 5.1 (max)Anthropic
0.424
0.3320.7950.5350.369$3.09
5GPT-5.6 Sol (high)OpenAI
0.393
0.3190.6890.4970.341$1.88
6Muse Spark 1.3 (medium)ParetoMeta
0.358
0.2460.8070.6330.221$0.07
7GLM-5.3 Flash (max)open weightsZ.ai
0.320
0.1880.8480.4650.247$0.09
8Qwen 3.8 27B (thinking on)open weightsAlibaba
0.226
0.1180.6600.2970.191$0.26
9DeepSeek V4.1 Flash (high)open weightsParetoDeepSeek
0.210
0.0660.7870.1890.221$0.05
10Qwen 3.8 27B (thinking off)open weightsAlibaba
0.200
0.0830.6670.2660.166$0.28
11Inkling (high)Thinking Machines
0.198
0.0950.6090.1690.213$0.75
12Grok 4.6 (high)xAI
0.184
0.0440.7440.1970.178$0.26
13Gemma 4 31B (thinking on)open weightsParetoGoogle
0.126
0.0030.6150.1270.125$0.03
14Gemma 4 31B (thinking off)open weightsParetoGoogle
0.115
0.0010.5720.1120.116$0.02
+
+

Takeaways

+

What the evaluation shows across the 14 agents. Every chart below is drawn from the paper's numbers; hover for values.

+
Takeaway 1

Construction and editing probe different capabilities: scene-level spatial reasoning and precise control of scene state.

Despite nearly identical overall scores, Astra is stronger at construction and Gemini at editing.

Text-to-Scene (construction)Image-to-Scene (editing)
← Text-to-SceneImage-to-Scene →0.7240.515GPT-6 Astra0.6570.581Gemini 3.8 Flash0.7880.424Claude Fable 5.10.7180.468Claude Opus 50.7070.393GPT-5.6 Sol0.6460.358Muse Spark 1.30.5090.320GLM-5.3 Flash0.5570.200Qwen 3.8 27B · off0.5670.184Grok 4.60.5160.226Qwen 3.8 27B · on0.4250.198Inkling0.4710.126Gemma 4 31B · on0.4550.115Gemma 4 31B · off0.2430.210DeepSeek V4.1 Flash
Text-to-Scene (left) and Image-to-Scene (right) score per agent, in overall order (paper Table 2). Astra and Gemini tie at 0.619 overall.
Takeaway 2

Spatial Composition remains the weakest requirement family for all 14 agents.

Errors persist even when the required objects are present: generating the right objects does not ensure that their relationships satisfy the specification.

Requirement-family score per agent
Identity & EnvironmentContent & QuantityAttributes & MaterialsSpatial Composition
0.20.40.60.81.0GPT-6 AstraGemini 3.8 FlashClaude Fable 5.1Claude Opus 5GPT-5.6 SolMuse Spark 1.3GLM-5.3 FlashQwen 3.8 27B · offGrok 4.6Qwen 3.8 27B · onInklingGemma 4 31B · onGemma 4 31B · offDeepSeek V4.1 FlashMean of 140.360.67
Each row is one agent's four family scores; the orange dot, Spatial Composition, is the lowest in every row (paper Figure 14b).
The objects are there; the relation is not
The judge's view of GPT-5.6 Sol's Egyptian Temple roof
✓ slab roof 0.95✓ stone coping 0.82✗ slab roof enclosed by the coping 0.78
The judge's own frame of GPT-5.6 Sol on Egyptian Temple. It finds the slab roof and the stone coping, but not the roof enclosed by the coping: the coping is separate low wall segments with visible gaps (ringed).
Mismatch rate when the required objects are present
Spatial relation32.0%Distribution19.1%Composition4.3%
Across the benchmark (paper Figure 6a).
Takeaway 3

Editing ability varies substantially across repair types.

Astra performs best on Transform repairs, while Gemini is stronger on Lifecycle and Layout operations.

GPT-6 AstraGemini 3.8 FlashClaude Fable 5.1Claude Opus 5Restore0.3240.3760.2440.388Lifecycle0.5770.7150.5540.593Layout0.5960.7030.4640.512Symmetry0.4660.5440.5550.359Transform0.6130.5320.3530.406
Mean Image-to-Scene score by repair type for four agents; the outlined cell is the best of the four in its row (paper Figure 6b).
Takeaway 4

Recovering the target does not guarantee precise editing.

Across all 14 configurations, 35.8% of complete recoveries still contain unintended changes to the surrounding scene.

1,050Image-to-Scene evaluations (14 agents × 75 cases)
212recovered every target completely
76 · 35.8%of those still changed the surrounding scene
unintended editsclean complete recoveryno complete recovery
One dot per Image-to-Scene evaluation (paper Figure 6c).
+

Turn the scenes yourself

+

For each selected case, every agent's saved scene is exported from Unreal Engine and opens in the browser in 3D, beside the evaluator's scores for it.

+ +
Open the case explorer
+
+

Cite This Work

+
@article{code4scene2026,
+  title   = {Code4Scene: Benchmarking Coding Agents for Constructing and Editing 3D Scenes},
+  author  = {Ye, Xiaokang and Mantri, Siddhant Hitesh and Chen, Zimeng and Zhang, Edward and Zheng, Zhaoxu and Li, Yuanheng and Chen, Yizhao and Huang, Tianyang and Qin, Lianhui},
+  year    = {2026}
+}
+
+ GitHub Repository +SimWorld Main Site
+
+ + + \ No newline at end of file