diff --git "a/index.html" "b/index.html" --- "a/index.html" +++ "b/index.html" @@ -38,11 +38,7 @@ svg.ch .dim{opacity:.42}svg.ch g.row:hover{opacity:1}svg.ch .hot{fill:var(--c4s-

Takeaways

What the evaluation shows across the 14 agents. Every chart below is drawn from the paper's numbers; hover for values.

-
Takeaway 1

Construction and editing probe different capabilities: scene-level spatial reasoning and precise control of scene state.

Despite nearly identical overall scores, Astra is stronger at construction and Gemini at editing.

Text-to-Scene (construction)Image-to-Scene (editing)
← Text-to-SceneImage-to-Scene →0.7240.515GPT-6 Astra0.6570.581Gemini 3.8 Flash0.7880.424Claude Fable 5.10.7180.468Claude Opus 50.7070.393GPT-5.6 Sol0.6460.358Muse Spark 1.30.5090.320GLM-5.3 Flash0.5570.200Qwen 3.8 27B · off0.5670.184Grok 4.60.5160.226Qwen 3.8 27B · on0.4250.198Inkling0.4710.126Gemma 4 31B · on0.4550.115Gemma 4 31B · off0.2430.210DeepSeek V4.1 Flash
Text-to-Scene (left) and Image-to-Scene (right) score per agent, in overall order (paper Table 2). Astra and Gemini both round to 0.619 overall.
Takeaway 2

Spatial Composition remains the weakest requirement family for all 14 agents.

Errors persist even when the required objects are present: generating the right objects does not ensure that their relationships satisfy the specification.

Requirement-family score per agent
Identity & EnvironmentContent & QuantityAttributes & MaterialsSpatial Composition
0.20.40.60.81.0GPT-6 AstraGemini 3.8 FlashClaude Fable 5.1Claude Opus 5GPT-5.6 SolMuse Spark 1.3GLM-5.3 FlashQwen 3.8 27B · offGrok 4.6Qwen 3.8 27B · onInklingGemma 4 31B · onGemma 4 31B · offDeepSeek V4.1 FlashMean of 140.360.67
Each row is one agent's four family scores; the orange dot, Spatial Composition, is the lowest in every row (paper Figure 14b).
The objects are there; the relation is not
Loading 3D scene…Drag to orbit · Scroll to zoom
✓ slab roof 0.95✓ stone coping 0.82✗ slab roof enclosed by the coping 0.78
GPT-5.6 Sol’s saved Egyptian Temple scene. The judge finds the roof and stone coping, but flags their spatial relationship. View the judge’s evidence ↗
Mismatch rate when the required objects are present
Spatial relation32.0%Distribution19.1%Composition4.3%
Across the benchmark (paper Figure 6a).
Takeaway 3

Editing ability varies substantially across repair types.

Astra performs best on Transform repairs, while Gemini is stronger on Lifecycle and Layout operations.

GPT-6 AstraGemini 3.8 FlashClaude Fable 5.1Claude Opus 5Restore0.3240.3760.2440.388Lifecycle0.5770.7150.5540.593Layout0.5960.7030.4640.512Symmetry0.4660.5440.5550.359Transform0.6130.5320.3530.406
Mean Image-to-Scene score by repair type for four agents; the outlined cell is the best of the four in its row (paper Figure 6b).
Takeaway 4

Recovering the target does not guarantee precise editing.

Across all 14 configurations, 35.8% of complete recoveries still contain unintended changes to the surrounding scene.

1,050Image-to-Scene evaluations (14 agents × 75 cases)
212recovered every target completely
76 · 35.8%of those still changed the surrounding scene
unintended editsclean complete recoveryno complete recovery
One dot per Image-to-Scene evaluation (paper Figure 6c).
-

How Code4Scene scores a scene

-

The leaderboard, one scene through the evaluator, two scenes compared, and what the benchmark finds, in 90 seconds.

-
-
+
Takeaway 1

Construction and editing probe different capabilities: scene-level spatial reasoning and precise control of scene state.

Despite nearly identical overall scores, Astra is stronger at construction and Gemini at editing.

Text-to-Scene (construction)Image-to-Scene (editing)
← Text-to-SceneImage-to-Scene →0.7240.515GPT-6 Astra0.6570.581Gemini 3.8 Flash0.7880.424Claude Fable 5.10.7180.468Claude Opus 50.7070.393GPT-5.6 Sol0.6460.358Muse Spark 1.30.5090.320GLM-5.3 Flash0.5570.200Qwen 3.8 27B · off0.5670.184Grok 4.60.5160.226Qwen 3.8 27B · on0.4250.198Inkling0.4710.126Gemma 4 31B · on0.4550.115Gemma 4 31B · off0.2430.210DeepSeek V4.1 Flash
Text-to-Scene (left) and Image-to-Scene (right) score per agent, in overall order (paper Table 2). Astra and Gemini both round to 0.619 overall.
Takeaway 2

Spatial Composition remains the weakest requirement family for all 14 agents.

Errors persist even when the required objects are present: generating the right objects does not ensure that their relationships satisfy the specification.

Requirement-family score per agent
Identity & EnvironmentContent & QuantityAttributes & MaterialsSpatial Composition
0.20.40.60.81.0GPT-6 AstraGemini 3.8 FlashClaude Fable 5.1Claude Opus 5GPT-5.6 SolMuse Spark 1.3GLM-5.3 FlashQwen 3.8 27B · offGrok 4.6Qwen 3.8 27B · onInklingGemma 4 31B · onGemma 4 31B · offDeepSeek V4.1 FlashMean of 140.360.67
Each row is one agent's four family scores; the orange dot, Spatial Composition, is the lowest in every row (paper Figure 14b).
The objects are there; the relation is not
Loading 3D scene…Drag to orbit · Scroll to zoom
✓ slab roof 0.95✓ stone coping 0.82✗ slab roof enclosed by the coping 0.78
GPT-5.6 Sol’s saved Egyptian Temple scene. The judge finds the roof and stone coping, but flags their spatial relationship. View the judge’s evidence ↗
Mismatch rate when the required objects are present
Spatial relation32.0%Distribution19.1%Composition4.3%
Across the benchmark (paper Figure 6a).
Takeaway 3

Editing ability varies substantially across repair types.

Astra performs best on Transform repairs, while Gemini is stronger on Lifecycle and Layout operations.

GPT-6 AstraGemini 3.8 FlashClaude Fable 5.1Claude Opus 5Restore0.3240.3760.2440.388Lifecycle0.5770.7150.5540.593Layout0.5960.7030.4640.512Symmetry0.4660.5440.5550.359Transform0.6130.5320.3530.406
Mean Image-to-Scene score by repair type for four agents; the outlined cell is the best of the four in its row (paper Figure 6b).
Takeaway 4

Recovering the target does not guarantee precise editing.

Across all 14 configurations, 35.8% of complete recoveries still contain unintended changes to the surrounding scene.

1,050Image-to-Scene evaluations (14 agents × 75 cases)
212recovered every target completely
76 · 35.8%of those still changed the surrounding scene
unintended editsclean complete recoveryno complete recovery
One dot per Image-to-Scene evaluation (paper Figure 6c).

Cite This Work

@misc{ye2026code4scene,
   title         = {Code4Scene: Benchmarking Coding Agents for Constructing and Editing 3D Scenes},