diff --git "a/index.html" "b/index.html" --- "a/index.html" +++ "b/index.html" @@ -2,7 +2,7 @@ -Code4Scene: Benchmarking Coding Agents for Constructing and Editing 3D Scenes +Code4Scene: Benchmarking Coding Agents for Constructing 3D Scenes -

Constructing and Editing
3D Scenes with Code

Benchmarking coding agents that turn text and reference images into engine-native 3D scenes. Measuring spatial reasoning, task fulfillment and precise control of scene state.

Task Format

Nordic Harbour · Text-to-Scene

An open-ended scene description, an asset catalog and the Unreal Engine editor.

Agent inputText + assets
Task instruction

Build me a compact, sunlit Copenhagen-style canal district with narrow cobblestone streets and rectangular waterways enclosing dense urban blocks.

Townhouses, canals, bridges and a waterfront promenade — with detailed requirements for materials, layout and street furniture.

Given: an empty level and the content pack’s asset catalog.

EvaluationEngine-native scene
CASE SCORE0.877
Detailed Alignment
0.757
Overview Alignment
0.909
Physical Safety
0.900

Task fulfillment, artifact integrity and static physical validity.

0.2 × Detailed + 0.6 × Overview + 0.2 × Physical

+ Full task instructionTEXT-TO-SCENE

Build me a compact, sunlit Copenhagen-style canal district with narrow cobblestone streets and rectangular waterways enclosing dense urban blocks. Arrange roughly thirty attached four- to six-storey townhouses in continuous street and canal-front rows, using slender façades, repetitive white-framed windows, arched doors and carriage entrances, modest cornices, and steep tiled roofs packed with dormers and brick chimneys. Include occasional exposed party walls between staggered building rows. Finish the façades in weathered muted red, mustard yellow, dusty blue, warm grey, beige, and exposed brick, with red, orange, and charcoal roof tiles. Run a principal waterfront street along a stone quay, crossed by smaller streets and linked across the canals by short, gently arched stone-and-metal bridges, including one opening onto a broad paved promenade with rounded landings. Form the canal edges from stepped pale stone, dark timber retaining walls, and iron railings around reflective rippling water. Add regularly spaced leafy trees, ornate black lamps, stone bollards, benches, European road signs, zebra crossings, puddles, grime, and sparse weeds.

320 full benchmark target · 160 construction + 160 editing, including private cases201 public cases · 129 construction + 72 editing (22 indoor, 50 outdoor)
+ +

Constructing
3D Scenes with Code

Benchmarking coding agents that turn open-ended scene descriptions into engine-native 3D scenes in Unreal Engine. Measuring spatial reasoning, task fulfillment and physical validity.

Task Format

Nordic Harbour · Text-to-Scene

An open-ended scene description, an asset catalog and the Unreal Engine editor.

Agent inputText + assets
Task instruction

Build me a compact, sunlit Copenhagen-style canal district with narrow cobblestone streets and rectangular waterways enclosing dense urban blocks.

Townhouses, canals, bridges and a waterfront promenade — with detailed requirements for materials, layout and street furniture.

Given: an empty level and the content pack’s asset catalog.

EvaluationEngine-native scene
CASE SCORE0.877
Detailed Alignment
0.757
Overview Alignment
0.909
Physical Safety
0.900

Task fulfillment, artifact integrity and static physical validity.

0.2 × Detailed + 0.6 × Overview + 0.2 × Physical

+ Full task instructionTEXT-TO-SCENE

Build me a compact, sunlit Copenhagen-style canal district with narrow cobblestone streets and rectangular waterways enclosing dense urban blocks. Arrange roughly thirty attached four- to six-storey townhouses in continuous street and canal-front rows, using slender façades, repetitive white-framed windows, arched doors and carriage entrances, modest cornices, and steep tiled roofs packed with dormers and brick chimneys. Include occasional exposed party walls between staggered building rows. Finish the façades in weathered muted red, mustard yellow, dusty blue, warm grey, beige, and exposed brick, with red, orange, and charcoal roof tiles. Run a principal waterfront street along a stone quay, crossed by smaller streets and linked across the canals by short, gently arched stone-and-metal bridges, including one opening onto a broad paved promenade with rounded landings. Form the canal edges from stepped pale stone, dark timber retaining walls, and iron railings around reflective rippling water. Add regularly spaced leafy trees, ornate black lamps, stone bollards, benches, European road signs, zebra crossings, puddles, grime, and sparse weeds.

160 construction cases in the full benchmark target, including private cases129 public construction cases

Leaderboard

-

14 coding-agent configurations · paper results on the original 95 public cases (20 construction + 75 editing). These results predate the 201-case public release.

-

Code4Scene

Paper evaluation · score out of 100

Swipe to view all 14 configurations →

Sorted by score. Bar labels round scores ×100; the table preserves three decimal places. One color per model provider.

-

Code4Scene Pareto Frontier

Quality against the cost of one case · original 95-case public evaluation

Swipe to explore all configurations →

Solid: the frontier — nothing cheaper scores higher. One point per configuration, at its evaluated reasoning effort. Cost is the mean USD per case; overall averages the two task means.

-
14 configurations
Ranking

Overall = ½ Text-to-Scene + ½ Image-to-Scene over the 95 public cases; its cost is weighted the same way.

#Agent configurationOverallText-to-SceneImage-to-SceneCost / case
1GPT-6 Astra (max)ParetoOpenAI
0.619
0.7240.515$12.90
2Gemini 3.8 Flash (high)ParetoGoogle
0.619
0.6570.581$1.92
3Claude Fable 5.1 (max)Anthropic
0.606
0.7880.424$10.62
4Claude Opus 5 (max)Anthropic
0.593
0.7180.468$19.48
5GPT-5.6 Sol (high)OpenAI
0.550
0.7070.393$2.97
6Muse Spark 1.3 (medium)ParetoMeta
0.502
0.6460.358$0.09
7GLM-5.3 Flash (max)open weightsZ.ai
0.415
0.5090.320$0.40
8Qwen 3.8 27B (thinking off)open weightsAlibaba
0.378
0.5570.200$0.38
9Grok 4.6 (high)xAI
0.376
0.5670.184$2.29
10Qwen 3.8 27B (thinking on)open weightsAlibaba
0.371
0.5160.226$0.30
11Inkling (high)Thinking Machines
0.312
0.4250.198$0.48
12Gemma 4 31B (thinking on)open weightsParetoGoogle
0.298
0.4710.126$0.03
13Gemma 4 31B (thinking off)open weightsParetoGoogle
0.285
0.4550.115$0.02
14DeepSeek V4.1 Flash (high)open weightsDeepSeek
0.227
0.2430.210$0.12

20 public cases. Case score = 0.2 × Detailed + 0.6 × Overview + 0.2 × Physical. Detailed and Overview are means of the per-case scores.

#Agent configurationScoreDetailedOverviewPhysicalCost / case
1Claude Fable 5.1 (max)ParetoAnthropic
0.788
0.7250.7720.898$18.16
2GPT-6 Astra (max)OpenAI
0.724
0.7070.7330.712$22.47
3Claude Opus 5 (max)Anthropic
0.718
0.6310.7160.811$34.50
4GPT-5.6 Sol (high)ParetoOpenAI
0.707
0.6320.7100.772$4.07
5Gemini 3.8 Flash (high)ParetoGoogle
0.657
0.6370.6640.655$2.49
6Muse Spark 1.3 (medium)ParetoMeta
0.646
0.5980.6610.651$0.10
7Grok 4.6 (high)xAI
0.567
0.5220.5760.585$4.33
8Qwen 3.8 27B (thinking off)open weightsAlibaba
0.557
0.5070.5300.688$0.49
9Qwen 3.8 27B (thinking on)open weightsAlibaba
0.516
0.5050.4600.694$0.33
10GLM-5.3 Flash (max)open weightsZ.ai
0.509
0.4030.5370.533$0.71
11Gemma 4 31B (thinking on)open weightsParetoGoogle
0.471
0.4120.4160.694$0.03
12Gemma 4 31B (thinking off)open weightsParetoGoogle
0.455
0.4380.3700.725$0.02
13Inkling (high)Thinking Machines
0.425
0.3930.3370.722$0.20
14DeepSeek V4.1 Flash (high)open weightsDeepSeek
0.243
0.2170.1610.513$0.19

75 public cases (25 indoor, 50 outdoor). Case score = 0.8 × Repair F1 + 0.2 × Physical. Repair F1 counts target actors restored to the withheld ground truth.

#Agent configurationScoreRepair F1PhysicalIndoorOutdoorCost / case
1Gemini 3.8 Flash (high)ParetoGoogle
0.581
0.5270.7960.7170.513$1.35
2GPT-6 Astra (max)OpenAI
0.515
0.4450.7960.7850.380$3.32
3Claude Opus 5 (max)Anthropic
0.468
0.3890.7860.6180.393$4.46
4Claude Fable 5.1 (max)Anthropic
0.424
0.3320.7950.5350.369$3.09
5GPT-5.6 Sol (high)OpenAI
0.393
0.3190.6890.4970.341$1.88
6Muse Spark 1.3 (medium)ParetoMeta
0.358
0.2460.8070.6330.221$0.07
7GLM-5.3 Flash (max)open weightsZ.ai
0.320
0.1880.8480.4650.247$0.09
8Qwen 3.8 27B (thinking on)open weightsAlibaba
0.226
0.1180.6600.2970.191$0.26
9DeepSeek V4.1 Flash (high)open weightsParetoDeepSeek
0.210
0.0660.7870.1890.221$0.05
10Qwen 3.8 27B (thinking off)open weightsAlibaba
0.200
0.0830.6670.2660.166$0.28
11Inkling (high)Thinking Machines
0.198
0.0950.6090.1690.213$0.75
12Grok 4.6 (high)xAI
0.184
0.0440.7440.1970.178$0.26
13Gemma 4 31B (thinking on)open weightsParetoGoogle
0.126
0.0030.6150.1270.125$0.03
14Gemma 4 31B (thinking off)open weightsParetoGoogle
0.115
0.0010.5720.1120.116$0.02
+

14 coding-agent configurations · paper results on the original 20 public construction cases.

+

Code4Scene

Paper evaluation · score out of 100

Swipe to view all 14 configurations →

Sorted by score. Bar labels round scores ×100; the table preserves three decimal places. One color per model provider.

+

Code4Scene Pareto Frontier

Quality against the cost of one case · original 20-case public evaluation

Swipe to explore all configurations →

Solid: the frontier — nothing cheaper scores higher. One point per configuration, at its evaluated reasoning effort. Cost is the mean USD per case.

+
14 configurations
Ranking

20 public cases. Case score = 0.2 × Detailed + 0.6 × Overview + 0.2 × Physical. Detailed and Overview are means of the per-case scores.

#Agent configurationScoreDetailedOverviewPhysicalCost / case
1Claude Fable 5.1 (max)ParetoAnthropic
0.788
0.7250.7720.898$18.16
2GPT-6 Astra (max)OpenAI
0.724
0.7070.7330.712$22.47
3Claude Opus 5 (max)Anthropic
0.718
0.6310.7160.811$34.50
4GPT-5.6 Sol (high)ParetoOpenAI
0.707
0.6320.7100.772$4.07
5Gemini 3.8 Flash (high)ParetoGoogle
0.657
0.6370.6640.655$2.49
6Muse Spark 1.3 (medium)ParetoMeta
0.646
0.5980.6610.651$0.10
7Grok 4.6 (high)xAI
0.567
0.5220.5760.585$4.33
8Qwen 3.8 27B (thinking off)open weightsAlibaba
0.557
0.5070.5300.688$0.49
9Qwen 3.8 27B (thinking on)open weightsAlibaba
0.516
0.5050.4600.694$0.33
10GLM-5.3 Flash (max)open weightsZ.ai
0.509
0.4030.5370.533$0.71
11Gemma 4 31B (thinking on)open weightsParetoGoogle
0.471
0.4120.4160.694$0.03
12Gemma 4 31B (thinking off)open weightsParetoGoogle
0.455
0.4380.3700.725$0.02
13Inkling (high)Thinking Machines
0.425
0.3930.3370.722$0.20
14DeepSeek V4.1 Flash (high)open weightsDeepSeek
0.243
0.2170.1610.513$0.19

Click a column to sort. Pareto: on the score-against-cost frontier.

Visualizing agent outputs

-

Drag any scene to inspect the saved 3D output. Compare repairs with their corrupted inputs and ground truth, or open a case to explore every agent.

-
Loading 3D scene…Drag to orbit · Scroll to zoom
Text-to-Scene

Nordic Harbour

GPT-6 Astra (max) · score 0.877

Compare all 14 agents ↗
Loading 3D scene…Drag to orbit · Scroll to zoom
Text-to-Scene

Medieval Big Farm Town

Claude Fable 5.1 (max) · score 0.830

Compare all 14 agents ↗
Loading 3D scene…Drag to orbit · Scroll to zoom
Text-to-Scene

Egyptian Temple

Claude Opus 5 (max) · score 0.777

Compare all 14 agents ↗
Loading 3D scene…Drag to orbit · Scroll to zoom
Text-to-Scene

Bazaar

Claude Fable 5.1 (max) · score 0.930

Compare all 14 agents ↗
Loading 3D scene…Drag to orbit · Scroll to zoom
Image-to-Scene

Old Industrial Pallet Bay

GPT-6 Astra (max) · score 1.000

Compare all 14 agents ↗
Loading 3D scene…Drag to orbit · Scroll to zoom
Image-to-Scene

New York Mailbox Pair

Claude Fable 5.1 (max) · score 1.000

Compare all 14 agents ↗
- +

Drag any scene to inspect the saved 3D output, or open a case to explore every agent.

+
Loading 3D scene…Drag to orbit · Scroll to zoom
Text-to-Scene

Nordic Harbour

GPT-6 Astra (max) · score 0.877

Compare all 14 agents ↗
Loading 3D scene…Drag to orbit · Scroll to zoom
Text-to-Scene

Medieval Big Farm Town

Claude Fable 5.1 (max) · score 0.830

Compare all 14 agents ↗
Loading 3D scene…Drag to orbit · Scroll to zoom
Text-to-Scene

Egyptian Temple

Claude Opus 5 (max) · score 0.777

Compare all 14 agents ↗
Loading 3D scene…Drag to orbit · Scroll to zoom
Text-to-Scene

Bazaar

Claude Fable 5.1 (max) · score 0.930

Compare all 14 agents ↗
+
-

Takeaways

+

Takeaway

What the evaluation shows across the 14 agents. Every chart below is drawn from the paper's numbers; hover for values.

-
Takeaway 1

Construction and editing probe different capabilities: scene-level spatial reasoning and precise control of scene state.

Despite nearly identical overall scores, Astra is stronger at construction and Gemini at editing.

Text-to-Scene (construction)Image-to-Scene (editing)
← Text-to-SceneImage-to-Scene →0.7240.515GPT-6 Astra0.6570.581Gemini 3.8 Flash0.7880.424Claude Fable 5.10.7180.468Claude Opus 50.7070.393GPT-5.6 Sol0.6460.358Muse Spark 1.30.5090.320GLM-5.3 Flash0.5570.200Qwen 3.8 27B · off0.5670.184Grok 4.60.5160.226Qwen 3.8 27B · on0.4250.198Inkling0.4710.126Gemma 4 31B · on0.4550.115Gemma 4 31B · off0.2430.210DeepSeek V4.1 Flash
Text-to-Scene (left) and Image-to-Scene (right) score per agent, in overall order (paper Table 2). Astra and Gemini both round to 0.619 overall.
Takeaway 2

Spatial Composition remains the weakest requirement family for all 14 agents.

Errors persist even when the required objects are present: generating the right objects does not ensure that their relationships satisfy the specification.

Requirement-family score per agent
Identity & EnvironmentContent & QuantityAttributes & MaterialsSpatial Composition
0.20.40.60.81.0GPT-6 AstraGemini 3.8 FlashClaude Fable 5.1Claude Opus 5GPT-5.6 SolMuse Spark 1.3GLM-5.3 FlashQwen 3.8 27B · offGrok 4.6Qwen 3.8 27B · onInklingGemma 4 31B · onGemma 4 31B · offDeepSeek V4.1 FlashMean of 140.360.67
Each row is one agent's four family scores; the orange dot, Spatial Composition, is the lowest in every row (paper Figure 14b).
The objects are there; the relation is not
Loading 3D scene…Drag to orbit · Scroll to zoom
✓ slab roof 0.95✓ stone coping 0.82✗ slab roof enclosed by the coping 0.78
GPT-5.6 Sol’s saved Egyptian Temple scene. The judge finds the roof and stone coping, but flags their spatial relationship. View the judge’s evidence ↗
Mismatch rate when the required objects are present
Spatial relation32.0%Distribution19.1%Composition4.3%
Across the benchmark (paper Figure 6a).
Takeaway 3

Editing ability varies substantially across repair types.

Astra performs best on Transform repairs, while Gemini is stronger on Lifecycle and Layout operations.

GPT-6 AstraGemini 3.8 FlashClaude Fable 5.1Claude Opus 5Restore0.3240.3760.2440.388Lifecycle0.5770.7150.5540.593Layout0.5960.7030.4640.512Symmetry0.4660.5440.5550.359Transform0.6130.5320.3530.406
Mean Image-to-Scene score by repair type for four agents; the outlined cell is the best of the four in its row (paper Figure 6b).
Takeaway 4

Recovering the target does not guarantee precise editing.

Across all 14 configurations, 35.8% of complete recoveries still contain unintended changes to the surrounding scene.

1,050Image-to-Scene evaluations (14 agents × 75 cases)
212recovered every target completely
76 · 35.8%of those still changed the surrounding scene
unintended editsclean complete recoveryno complete recovery
One dot per Image-to-Scene evaluation (paper Figure 6c).
+
Takeaway

Spatial Composition remains the weakest requirement family for all 14 agents.

Errors persist even when the required objects are present: generating the right objects does not ensure that their relationships satisfy the specification.

Requirement-family score per agent
Identity & EnvironmentContent & QuantityAttributes & MaterialsSpatial Composition
0.20.40.60.81.0GPT-6 AstraGemini 3.8 FlashClaude Fable 5.1Claude Opus 5GPT-5.6 SolMuse Spark 1.3GLM-5.3 FlashQwen 3.8 27B · offGrok 4.6Qwen 3.8 27B · onInklingGemma 4 31B · onGemma 4 31B · offDeepSeek V4.1 FlashMean of 140.360.67
Each row is one agent's four family scores; the orange dot, Spatial Composition, is the lowest in every row (paper Figure 14b).
The objects are there; the relation is not
Loading 3D scene…Drag to orbit · Scroll to zoom
✓ slab roof 0.95✓ stone coping 0.82✗ slab roof enclosed by the coping 0.78
GPT-5.6 Sol’s saved Egyptian Temple scene. The judge finds the roof and stone coping, but flags their spatial relationship. View the judge’s evidence ↗
Mismatch rate when the required objects are present
Spatial relation32.0%Distribution19.1%Composition4.3%
Across the benchmark (paper Figure 6a).

Cite This Work

@misc{ye2026code4scene,
   title         = {Code4Scene: Benchmarking Coding Agents for Constructing and Editing 3D Scenes},
@@ -50,12 +48,12 @@ svg.ch .dim{opacity:.42}svg.ch g.row:hover{opacity:1}svg.ch .hot{fill:var(--c4s-
   url           = {https://arxiv.org/abs/2609.36777},
 }
- arXiv - GitHub Repository + arXiv + GitHub Repository SimWorld Main Site
- \ No newline at end of file + \ No newline at end of file