Text-to-Scene (left) and Image-to-Scene (right) score per agent, in overall order (paper Table 2). Astra and Gemini both round to 0.619 overall.
Takeaway 2
Spatial Composition remains the weakest requirement family for all 14 agents.
Errors persist even when the required objects are present: generating the right objects does not ensure that their relationships satisfy the specification.
Each row is one agent's four family scores; the orange dot, Spatial Composition, is the lowest in every row (paper Figure 14b).
The objects are there; the relation is not
Loading 3D scene…Drag to orbit · Scroll to zoom
✓ slab roof 0.95✓ stone coping 0.82✗ slab roof enclosed by the coping 0.78
GPT-5.6 Sol’s saved Egyptian Temple scene. The judge finds the roof and stone coping, but flags their spatial relationship. View the judge’s evidence ↗
Mismatch rate when the required objects are present
Across the benchmark (paper Figure 6a).
Takeaway 3
Editing ability varies substantially across repair types.
Astra performs best on Transform repairs, while Gemini is stronger on Lifecycle and Layout operations.
Mean Image-to-Scene score by repair type for four agents; the outlined cell is the best of the four in its row (paper Figure 6b).
Takeaway 4
Recovering the target does not guarantee precise editing.
Across all 14 configurations, 35.8% of complete recoveries still contain unintended changes to the surrounding scene.
Text-to-Scene (left) and Image-to-Scene (right) score per agent, in overall order (paper Table 2). Astra and Gemini both round to 0.619 overall.
Takeaway 2
Spatial Composition remains the weakest requirement family for all 14 agents.
Errors persist even when the required objects are present: generating the right objects does not ensure that their relationships satisfy the specification.
Each row is one agent's four family scores; the orange dot, Spatial Composition, is the lowest in every row (paper Figure 14b).
The objects are there; the relation is not
Loading 3D scene…Drag to orbit · Scroll to zoom
✓ slab roof 0.95✓ stone coping 0.82✗ slab roof enclosed by the coping 0.78
GPT-5.6 Sol’s saved Egyptian Temple scene. The judge finds the roof and stone coping, but flags their spatial relationship. View the judge’s evidence ↗
Mismatch rate when the required objects are present
Across the benchmark (paper Figure 6a).
Takeaway 3
Editing ability varies substantially across repair types.
Astra performs best on Transform repairs, while Gemini is stronger on Lifecycle and Layout operations.
Mean Image-to-Scene score by repair type for four agents; the outlined cell is the best of the four in its row (paper Figure 6b).
Takeaway 4
Recovering the target does not guarantee precise editing.
Across all 14 configurations, 35.8% of complete recoveries still contain unintended changes to the surrounding scene.