diff --git a/.gitattributes b/.gitattributes
index f0df735fd3eab65a211ca6657220109c73c44cbf..c41838348fc08cb89f48356aa98d8522733c3573 100644
--- a/.gitattributes
+++ b/.gitattributes
@@ -95,3 +95,4 @@ media/t2s/nordic-harbour/qwen-on.jpg filter=lfs diff=lfs merge=lfs -text
media/t2s/nordic-harbour/sol.jpg filter=lfs diff=lfs merge=lfs -text
media/teaser-poster.jpg filter=lfs diff=lfs merge=lfs -text
media/teaser.mp4 filter=lfs diff=lfs merge=lfs -text
+media/teaser-v31.mp4 filter=lfs diff=lfs merge=lfs -text
diff --git a/README.md b/README.md
index be804997e5b4f2aa7fea31e01ab15f35e6a28933..50582a7137114074f68fcd8b0c5e7e6d949621ba 100644
--- a/README.md
+++ b/README.md
@@ -6,13 +6,14 @@ colorTo: gray
sdk: static
app_file: index.html
pinned: false
-short_description: Benchmarking coding agents that build and edit 3D scenes
+short_description: Benchmarking coding agents that build 3D scenes
---
# Code4Scene
-Project page for *Code4Scene: Benchmarking Coding Agents for Constructing and Editing 3D Scenes*: the leaderboard on the
-95-case public set, the paper's takeaways, and selected cases whose saved Unreal Engine scenes can be viewed in 3D.
+Project page for Code4Scene's construction (Text-to-Scene) setting: the leaderboard on the 20-case public set, the paper's
+Spatial Composition takeaway, and selected cases whose saved Unreal Engine scenes can be viewed in 3D. Since 2026-10-01 the
+page shows Text-to-Scene only; `scripts/t2s_only.py` records that edit.
Paper: [arXiv:2609.36777](https://arxiv.org/abs/2609.36777) · Code: [SimWorld-AI/Code4Scene](https://github.com/SimWorld-AI/Code4Scene)
@@ -20,23 +21,23 @@ Part of the [SimWorld](https://simworld.org) project.
## Website design
-Layout, navigation, wireframe mark, typography and card geometry follow the user-provided [BuildingBench](https://enactra.ai/buildingbench/) design. The bar chart rendering and company palette are adapted from its [Figure Kit](http://ds-serv12.ucsd.edu:8875/figure-kit/#bar), specifically `bar-static/build.py`. All scores, configurations, authors, media, models and evaluator content remain Code4Scene's. The original 95-public-case paper evaluation is distinguished from the current 201 public cases and the 320-case full benchmark target, including private cases.
+Layout, navigation, wireframe mark, typography and card geometry follow the user-provided [BuildingBench](https://enactra.ai/buildingbench/) design. The bar chart rendering and company palette are adapted from its [Figure Kit](http://ds-serv12.ucsd.edu:8875/figure-kit/#bar), specifically `bar-static/build.py`. All scores, configurations, authors, media, models and evaluator content remain Code4Scene's. The paper's 20-case public Text-to-Scene evaluation is distinguished from the current 129 public construction cases and the 160-case full construction target, including private cases.
The reference toolkit's Inter, IBM Plex Mono and Source Serif 4 font files are bundled with their original SIL Open Font License notices in `assets/fonts/`. Provider marks are from the toolkit's Lobehub icons, used to identify the corresponding model providers.
Deployment remains this Hugging Face static Space: `index.html` and `cases.html`, with the existing media/model paths and headers. The pre-restyle version is recoverable at commit `9ff787d0f9d72905d92f780da56bd2229782b981`.
-The Pareto renderer (`assets/pareto.js`) ports the reference toolkit's reversed logarithmic cost axis, connected frontier, provider palette and greedy collision-aware label placement. `data/pareto.json` combines unrounded scores from the existing `data/site.js` with exact mean costs from the original published tables; all three task views recompute non-dominance from those values. The table displays three decimals but sorts at source precision.
+The Pareto renderer (`assets/pareto.js`) ports the reference toolkit's reversed logarithmic cost axis, connected frontier, provider palette and greedy collision-aware label placement. `data/pareto.json` combines unrounded scores from the existing `data/site.js` with exact mean costs from the original published tables; the Text-to-Scene view recomputes non-dominance from those values. The table displays three decimals but sorts at source precision.
-Homepage scenes (`assets/scene-cards.js`) adapt BuildingBench's shared WebGL renderer and per-view canvas engine. Existing public packed models are decoded with the same adapter as `cases.html`; geometry and the full case explorer remain unchanged. Visible scenes load automatically; repeated scenes share decoded assets. Drag or arrow keys orbit, wheel or +/- zoom, and the pause control stops rotation. Both Image-to-Scene cards switch between the saved repair, corrupted input and ground truth using the reference camera.
+Homepage scenes (`assets/scene-cards.js`) adapt BuildingBench's shared WebGL renderer and per-view canvas engine. Existing public packed models are decoded with the same adapter as `cases.html`; geometry and the full case explorer remain unchanged. Visible scenes load automatically; repeated scenes share decoded assets. Drag or arrow keys orbit, wheel or +/- zoom, and the pause control stops rotation.
## Interaction and loading reliability
-Task Format switches the entire construction/editing example, including instructions, inputs, saved output and evaluator. Bar marks, labels and Pareto points share pointer, keyboard and tap tooltips with source-precision scores and costs in `data/leaderboard-details.json`. The ranking supports filtering and source-precision sorting.
+Task Format shows the construction example: instructions, inputs, saved output and evaluator. Bar marks, labels and Pareto points share pointer, keyboard and tap tooltips with source-precision scores and costs in `data/leaderboard-details.json`. The ranking supports filtering and source-precision sorting.
Homepage model work uses a shared two-job download/decode queue. Hidden/offscreen views release their request; stale results cannot replace the selected state. Card textures are limited to 512px and render resolution is bounded, while original model files and the full case explorer stay intact. Idle decoded assets are evicted; the 128 MiB cache target is not a hard limit on currently visible geometry. Downloads show progress, abort after 20 seconds without data or 90 seconds total, retry once, then expose an explicit Retry button. Context loss pauses drawing and restoration rebuilds the environment; failures expose recovery instead of an indefinite loading overlay.
-Queue tests: `node --test tests/scene-assets.test.mjs`. Browser QA includes public-model cold loads, quick scrolling, task/repair switching, desktop/mobile resize, forced HTTP failures, stalled downloads, retry and WebGL context loss.
+Queue tests: `node --test tests/scene-assets.test.mjs`. Browser QA includes public-model cold loads, quick scrolling, desktop/mobile resize, forced HTTP failures, stalled downloads, retry and WebGL context loss.
## Geometry-first scene cards
diff --git a/assets/buildingbench.css b/assets/buildingbench.css
index b3c018df21be8b7e4b3b041b55419430c7ee98d7..d66d949dfa3029df984a09ac647ed5981b809005 100644
--- a/assets/buildingbench.css
+++ b/assets/buildingbench.css
@@ -14,7 +14,7 @@ html{background:var(--bg);scroll-padding-top:24px}body{color:var(--ink);backgrou
section,section.alt,section.dark{padding:34px 0;background:transparent;color:var(--text-p)}.section-header,.section-header.centered{text-align:left;margin-bottom:22px;max-width:none}.section-label{display:none}.section-title,section.dark .section-title{font-size:clamp(30px,4.2vw,46px);line-height:1.04;letter-spacing:-1.9px;font-weight:700;color:#171816;margin:6px 0 20px}.section-desc,section.dark .section-desc{font-size:14px;line-height:1.65;max-width:850px;margin:0;color:var(--text-p)}.card{background:white;border:1px solid var(--border);border-radius:20px;box-shadow:none}.card:hover{box-shadow:none}.btn{font-size:12px;padding:10px 14px;border-radius:6px;box-shadow:none;font-weight:600}.btn:hover{transform:none;box-shadow:none}.btn-primary{background:var(--green-dark);color:white;border:1px solid var(--green-dark);box-shadow:none}.btn-primary:hover{background:var(--green);color:white;box-shadow:none}.btn-outline,.btn-dark,section.dark .btn-outline{background:#fff;border:1px solid var(--border);color:var(--green-dark)}
.task-section{padding-top:0}.task-shell{padding:20px;border:1px solid #d7dbd5;border-radius:28px;background:white}.task-shell>h3{font-size:22px;letter-spacing:-.6px;line-height:1.35;color:var(--text-h);margin:0 0 5px}.task-shell>h3 span{font-weight:500}.task-subtitle{font-size:14px;margin:0 0 20px;color:#60655e}.task-columns{display:grid;grid-template-columns:1fr 1.55fr 1fr;gap:14px}.task-panel{min-width:0;background:#fbfcf9;border:1px solid #d7dbd5;border-radius:18px;padding:14px;display:flex;flex-direction:column}.panel-heading{display:flex;justify-content:space-between;gap:10px;align-items:baseline;margin-bottom:14px;font-size:13px;color:#171816}.panel-heading span{font-size:11px;color:#626960;text-align:right}.prompt-preview{border-radius:12px;background:#f0f3ed;padding:16px;flex:1}.prompt-preview p{font-size:14px;line-height:1.6;margin:9px 0}.prompt-preview .prompt-detail{font-size:12px;color:#747b71}.output-panel img{width:100%;height:253px;object-fit:cover;border:1px solid #e4e4e2;border-radius:12px}.task-foot{font-size:11px;line-height:1.6;color:#626960;margin:12px 0 0}.score-block{padding:20px 18px;border-radius:14px;background:#171816;color:#fff;flex:1;display:flex;flex-direction:column;justify-content:center}.score-block>span{font-size:11px;color:#bbbfba;letter-spacing:.06em}.score-block>strong{font-size:52px;line-height:1.2;letter-spacing:-2px}.score-block dl{margin-top:20px;display:grid;gap:12px}.score-block dl>div{display:flex;justify-content:space-between;gap:8px;align-items:center;font-size:11px}.score-block dd{font-size:15px;font-weight:500;font-variant-numeric:tabular-nums}.task-full{margin-top:18px;border:1px solid var(--border);border-radius:12px}.task-full summary{padding:13px 16px;display:flex;justify-content:space-between;gap:12px;font-size:12px;font-weight:500;cursor:pointer}.task-full summary span+span{font-size:10px;color:var(--text-muted)}.task-full>p{padding:0 18px 18px;font-size:13px;line-height:1.8}.task-settings .grid2{margin:18px 0 0!important;gap:14px}.task-settings .setting-card{padding:16px;background:#fafbf8;border-radius:14px}.setting-card h3{font-size:18px;margin-bottom:7px}.setting-card p{font-size:12px;line-height:1.65}.setting-card .kicker{font-size:10px;color:var(--muted)}.formula{font-size:11px;white-space:normal}.availability{border-bottom:1px solid var(--border);padding:18px 4px;display:flex;gap:10px 24px;flex-wrap:wrap;font-size:11px;line-height:1.7;color:var(--text-muted)}.availability b{color:var(--text-h);font-size:14px}
.lb-tabs{justify-content:flex-start;margin-bottom:16px}.seg{background:#f0f3ef;border:1px solid var(--border);border-radius:7px;padding:3px;gap:3px}.seg button,.seg-lg button{font-size:12px;padding:8px 14px;border-radius:4px;color:#778178}.seg button.on{background:white;color:var(--ink);box-shadow:0 1px 4px #152a1915}.bar-card{padding:20px 18px 12px;margin-bottom:18px;overflow:hidden}.bar-heading{display:flex;justify-content:space-between;align-items:baseline;padding:0 12px;gap:20px}.bar-heading h3{font:400 31px 'Source Serif 4',Georgia,serif;color:#1a1a1a}.bar-heading span{font-size:12px;color:#8d8d8d}.bar-scroll{overflow:auto;scrollbar-width:thin}.kit-bars{display:block;width:100%;min-width:720px;height:auto}.bar-card .note{margin:0 12px;font-size:11px;color:#738078}.chart-card{padding:20px 24px;margin-bottom:18px}.sub-title{font-size:18px;letter-spacing:-.3px;color:var(--text-h);font-weight:700}.sub-desc{font-size:12px;color:var(--text-muted);line-height:1.6}.ch-legend{font-size:11px;gap:6px 16px;margin-top:14px}.note,.board-note{font-size:11px;color:var(--text-muted)}.board{font-size:13px}.board th{font-size:10px;font-weight:500;color:#7d877d;background:#f8faf6;letter-spacing:.03em}.board td{padding:12px 10px}.board tbody tr:hover{background:#f5f8f2}.board .tag{font-size:9px;background:#edf3ef;color:#4b715d}.board .cellbar i{opacity:.13}.board .cellbar b{color:#284e3f}.board-wrap{overflow:auto}
-.case-grid{grid-template-columns:repeat(3,minmax(0,1fr));gap:16px}.case-grid .case-card{border:1px solid #d7ded5;border-radius:16px;box-shadow:none;background:white;overflow:hidden}.case-grid .case-card:hover{transform:none;border-color:var(--green);box-shadow:none}.case-card .th{aspect-ratio:16/10;background-color:#f1f1ee}.case-card .th i{font-size:10px;color:#36544b;background:#ffffffed;border:1px solid #e1e4de;border-radius:5px}.case-card .meta{padding:15px 16px}.case-card .meta h4{font-size:15px;color:var(--text-h)}.case-card .meta p{font-size:11px;color:var(--text-muted);margin-top:6px}.tk{padding:24px;border:1px solid var(--border);border-radius:22px;background:white;margin:18px 0}.tk .n{font-size:10px;text-transform:uppercase;letter-spacing:.08em;color:var(--muted);background:none;padding:0;border:0}.tk h3{font-size:24px;line-height:1.25;letter-spacing:-.7px;max-width:1000px}.tk p.s{font-size:14px;color:#62695e}.tk figure{background:#fbfcf9;border-color:var(--line);border-radius:14px}.tk figcaption{font-size:11px;color:var(--text-muted)}
+.case-grid{grid-template-columns:repeat(2,minmax(0,1fr));gap:16px}.case-grid .case-card{border:1px solid #d7ded5;border-radius:16px;box-shadow:none;background:white;overflow:hidden}.case-grid .case-card:hover{transform:none;border-color:var(--green);box-shadow:none}.case-card .th{aspect-ratio:16/10;background-color:#f1f1ee}.case-card .th i{font-size:10px;color:#36544b;background:#ffffffed;border:1px solid #e1e4de;border-radius:5px}.case-card .meta{padding:15px 16px}.case-card .meta h4{font-size:15px;color:var(--text-h)}.case-card .meta p{font-size:11px;color:var(--text-muted);margin-top:6px}.tk{padding:24px;border:1px solid var(--border);border-radius:22px;background:white;margin:18px 0}.tk .n{font-size:10px;text-transform:uppercase;letter-spacing:.08em;color:var(--muted);background:none;padding:0;border:0}.tk h3{font-size:24px;line-height:1.25;letter-spacing:-.7px;max-width:1000px}.tk p.s{font-size:14px;color:#62695e}.tk figure{background:#fbfcf9;border-color:var(--line);border-radius:14px}.tk figcaption{font-size:11px;color:var(--text-muted)}
.teaser-frame{border:1px solid var(--border);border-radius:20px;box-shadow:none;background:#171816;max-width:none}.teaser-frame video{width:100%;max-height:680px}.author-block{margin:22px 0}.hero-authors{font-size:13px;color:var(--ink);margin-bottom:8px}.hero-affiliations{font-size:11px;color:var(--muted);margin-bottom:0}.paper-abstract{border:1px solid var(--border);border-radius:16px;background:white}.paper-abstract summary{padding:18px 20px;display:flex;justify-content:space-between;gap:20px;cursor:pointer;font-size:15px;font-weight:600;color:var(--text-h)}.paper-abstract summary span{font-size:11px;font-weight:400;color:var(--muted)}.paper-abstract .abstract-card{border:0;box-shadow:none;padding:0 20px 20px;max-width:none;font-size:13px;line-height:1.8}.method-flow .flow{gap:12px;margin-top:24px}.flow .step{background:#fff;border:1px solid var(--border);border-radius:12px;padding:15px}.flow .step:after{display:none}.flow .step .n{font-size:10px;color:var(--green)}.flow .step b{font-size:13px}.flow .step span{font-size:11px}.cite-box{background:white;border:1px solid var(--border);border-radius:16px;max-width:none}.cite-box pre{color:#4d5a50;font-size:12px}.cite-copy-btn{background:#f4f7f0;color:var(--ink);border:1px solid var(--border);border-radius:5px}footer{max-width:1360px;margin:auto;background:transparent;border-top:1px solid var(--border);padding:24px 0;color:var(--muted);text-align:left;font-size:11px}footer a{color:var(--green)}footer p:last-child{color:var(--muted)!important}
.case-page>section{max-width:1440px;padding-left:40px;padding-right:40px;margin:auto}.case-page .card{border-radius:22px}.case-page .viewer{background:#f1f1ee;border:1px solid #e4e4e2;border-radius:14px}.case-page .casebar button{background:#fff;font-family:Inter,sans-serif;border-radius:10px}.case-page .casebar button.on{background:#edf3f5;border-color:var(--green);color:var(--green-dark)}.case-page .casebar button.on small{color:#647a82}.case-page .panel{border-radius:14px}.case-page .section-title{font-size:38px}.case-page .agents{grid-template-columns:repeat(7,minmax(0,1fr))}
@media(min-width:1500px){.intro{padding-top:39px;padding-bottom:35px}}
diff --git a/assets/leaderboard-bars.js b/assets/leaderboard-bars.js
index 2e2916c625e84a9553188a8ef8429c5396638254..d42536c70d3e6e124cc2430eb0638aaafce26fc9 100644
--- a/assets/leaderboard-bars.js
+++ b/assets/leaderboard-bars.js
@@ -3,7 +3,7 @@
(async function(){
const {bindChartTooltip}=await import("./chart-interactions.js?v=20261001c");
await Promise.all([document.fonts.load('400 13px Inter'),document.fonts.load('700 13px Inter'),document.fonts.load('500 15.5px "IBM Plex Mono"')]);
-const ALL = {"overall": [{"name": "GPT-6 Astra (max)", "org": "OpenAI", "score": 0.6193352866666667, "colour": "#10A37F", "ink": "#ffffff", "label": "62", "new": false}, {"name": "Gemini 3.8 Flash (high)", "org": "Google", "score": 0.6189750266666667, "colour": "#7b1fa2", "ink": "#ffffff", "label": "62", "new": false}, {"name": "Claude Fable 5.1 (max)", "org": "Anthropic", "score": 0.60604898, "colour": "#D97757", "ink": "#ffffff", "label": "61", "new": false}, {"name": "Claude Opus 5 (max)", "org": "Anthropic", "score": 0.5931299933333334, "colour": "#D97757", "ink": "#ffffff", "label": "59", "new": false}, {"name": "GPT-5.6 Sol (high)", "org": "OpenAI", "score": 0.5499994666666667, "colour": "#10A37F", "ink": "#ffffff", "label": "55", "new": false}, {"name": "Muse Spark 1.3 (medium)", "org": "Meta", "score": 0.5022053666666666, "colour": "#42a5f5", "ink": "#111111", "label": "50", "new": false}, {"name": "GLM-5.3 Flash (max)", "org": "Z.ai", "score": 0.4147088666666666, "colour": "#96650b", "ink": "#ffffff", "label": "41", "new": false}, {"name": "Qwen 3.8 27B (thinking off)", "org": "Alibaba", "score": 0.37818169333333335, "colour": "#fe7016", "ink": "#111111", "label": "38", "new": false}, {"name": "Grok 4.6 (high)", "org": "xAI", "score": 0.37567506, "colour": "#111111", "ink": "#ffffff", "label": "38", "new": false}, {"name": "Qwen 3.8 27B (thinking on)", "org": "Alibaba", "score": 0.37094399333333333, "colour": "#fe7016", "ink": "#111111", "label": "37", "new": false}, {"name": "Inkling (high)", "org": "Thinking Machines", "score": 0.31150237333333336, "colour": "#686868", "ink": "#ffffff", "label": "31", "new": false}, {"name": "Gemma 4 31B (thinking on)", "org": "Google", "score": 0.29812257999999997, "colour": "#7b1fa2", "ink": "#ffffff", "label": "30", "new": false}, {"name": "Gemma 4 31B (thinking off)", "org": "Google", "score": 0.28498897333333334, "colour": "#7b1fa2", "ink": "#ffffff", "label": "28", "new": false}, {"name": "DeepSeek V4.1 Flash (high)", "org": "DeepSeek", "score": 0.22658958, "colour": "#3f51e0", "ink": "#ffffff", "label": "23", "new": false}], "t2s": [{"name": "Claude Fable 5.1 (max)", "org": "Anthropic", "score": 0.78775, "colour": "#D97757", "ink": "#ffffff", "label": "79", "new": false}, {"name": "GPT-6 Astra (max)", "org": "OpenAI", "score": 0.72382, "colour": "#10A37F", "ink": "#ffffff", "label": "72", "new": false}, {"name": "Claude Opus 5 (max)", "org": "Anthropic", "score": 0.7181599999999999, "colour": "#D97757", "ink": "#ffffff", "label": "72", "new": false}, {"name": "GPT-5.6 Sol (high)", "org": "OpenAI", "score": 0.70704, "colour": "#10A37F", "ink": "#ffffff", "label": "71", "new": false}, {"name": "Gemini 3.8 Flash (high)", "org": "Google", "score": 0.65689, "colour": "#7b1fa2", "ink": "#ffffff", "label": "66", "new": false}, {"name": "Muse Spark 1.3 (medium)", "org": "Meta", "score": 0.646195, "colour": "#42a5f5", "ink": "#111111", "label": "65", "new": false}, {"name": "Grok 4.6 (high)", "org": "xAI", "score": 0.56721, "colour": "#111111", "ink": "#ffffff", "label": "57", "new": false}, {"name": "Qwen 3.8 27B (thinking off)", "org": "Alibaba", "score": 0.55666, "colour": "#fe7016", "ink": "#111111", "label": "56", "new": false}, {"name": "Qwen 3.8 27B (thinking on)", "org": "Alibaba", "score": 0.51571, "colour": "#fe7016", "ink": "#111111", "label": "52", "new": false}, {"name": "GLM-5.3 Flash (max)", "org": "Z.ai", "score": 0.5094299999999999, "colour": "#96650b", "ink": "#ffffff", "label": "51", "new": false}, {"name": "Gemma 4 31B (thinking on)", "org": "Google", "score": 0.470705, "colour": "#7b1fa2", "ink": "#ffffff", "label": "47", "new": false}, {"name": "Gemma 4 31B (thinking off)", "org": "Google", "score": 0.45494, "colour": "#7b1fa2", "ink": "#ffffff", "label": "45", "new": false}, {"name": "Inkling (high)", "org": "Thinking Machines", "score": 0.42498500000000006, "colour": "#686868", "ink": "#ffffff", "label": "42", "new": false}, {"name": "DeepSeek V4.1 Flash (high)", "org": "DeepSeek", "score": 0.24269, "colour": "#3f51e0", "ink": "#ffffff", "label": "24", "new": false}], "i2s": [{"name": "Gemini 3.8 Flash (high)", "org": "Google", "score": 0.5810600533333334, "colour": "#7b1fa2", "ink": "#ffffff", "label": "58", "new": false}, {"name": "GPT-6 Astra (max)", "org": "OpenAI", "score": 0.5148505733333334, "colour": "#10A37F", "ink": "#ffffff", "label": "51", "new": false}, {"name": "Claude Opus 5 (max)", "org": "Anthropic", "score": 0.46809998666666675, "colour": "#D97757", "ink": "#ffffff", "label": "47", "new": false}, {"name": "Claude Fable 5.1 (max)", "org": "Anthropic", "score": 0.4243479600000001, "colour": "#D97757", "ink": "#ffffff", "label": "42", "new": false}, {"name": "GPT-5.6 Sol (high)", "org": "OpenAI", "score": 0.3929589333333333, "colour": "#10A37F", "ink": "#ffffff", "label": "39", "new": false}, {"name": "Muse Spark 1.3 (medium)", "org": "Meta", "score": 0.35821573333333334, "colour": "#42a5f5", "ink": "#111111", "label": "36", "new": false}, {"name": "GLM-5.3 Flash (max)", "org": "Z.ai", "score": 0.3199877333333333, "colour": "#96650b", "ink": "#ffffff", "label": "32", "new": false}, {"name": "Qwen 3.8 27B (thinking on)", "org": "Alibaba", "score": 0.2261779866666667, "colour": "#fe7016", "ink": "#111111", "label": "23", "new": false}, {"name": "DeepSeek V4.1 Flash (high)", "org": "DeepSeek", "score": 0.21048916, "colour": "#3f51e0", "ink": "#ffffff", "label": "21", "new": false}, {"name": "Qwen 3.8 27B (thinking off)", "org": "Alibaba", "score": 0.19970338666666668, "colour": "#fe7016", "ink": "#111111", "label": "20", "new": false}, {"name": "Inkling (high)", "org": "Thinking Machines", "score": 0.19801974666666666, "colour": "#686868", "ink": "#ffffff", "label": "20", "new": false}, {"name": "Grok 4.6 (high)", "org": "xAI", "score": 0.18414012, "colour": "#111111", "ink": "#ffffff", "label": "18", "new": false}, {"name": "Gemma 4 31B (thinking on)", "org": "Google", "score": 0.12554015999999998, "colour": "#7b1fa2", "ink": "#ffffff", "label": "13", "new": false}, {"name": "Gemma 4 31B (thinking off)", "org": "Google", "score": 0.11503794666666667, "colour": "#7b1fa2", "ink": "#ffffff", "label": "12", "new": false}]};
+const ALL = {"t2s": [{"name": "Claude Fable 5.1 (max)", "org": "Anthropic", "score": 0.78775, "colour": "#D97757", "ink": "#ffffff", "label": "79", "new": false}, {"name": "GPT-6 Astra (max)", "org": "OpenAI", "score": 0.72382, "colour": "#10A37F", "ink": "#ffffff", "label": "72", "new": false}, {"name": "Claude Opus 5 (max)", "org": "Anthropic", "score": 0.7181599999999999, "colour": "#D97757", "ink": "#ffffff", "label": "72", "new": false}, {"name": "GPT-5.6 Sol (high)", "org": "OpenAI", "score": 0.70704, "colour": "#10A37F", "ink": "#ffffff", "label": "71", "new": false}, {"name": "Gemini 3.8 Flash (high)", "org": "Google", "score": 0.65689, "colour": "#7b1fa2", "ink": "#ffffff", "label": "66", "new": false}, {"name": "Muse Spark 1.3 (medium)", "org": "Meta", "score": 0.646195, "colour": "#42a5f5", "ink": "#111111", "label": "65", "new": false}, {"name": "Grok 4.6 (high)", "org": "xAI", "score": 0.56721, "colour": "#111111", "ink": "#ffffff", "label": "57", "new": false}, {"name": "Qwen 3.8 27B (thinking off)", "org": "Alibaba", "score": 0.55666, "colour": "#fe7016", "ink": "#111111", "label": "56", "new": false}, {"name": "Qwen 3.8 27B (thinking on)", "org": "Alibaba", "score": 0.51571, "colour": "#fe7016", "ink": "#111111", "label": "52", "new": false}, {"name": "GLM-5.3 Flash (max)", "org": "Z.ai", "score": 0.5094299999999999, "colour": "#96650b", "ink": "#ffffff", "label": "51", "new": false}, {"name": "Gemma 4 31B (thinking on)", "org": "Google", "score": 0.470705, "colour": "#7b1fa2", "ink": "#ffffff", "label": "47", "new": false}, {"name": "Gemma 4 31B (thinking off)", "org": "Google", "score": 0.45494, "colour": "#7b1fa2", "ink": "#ffffff", "label": "45", "new": false}, {"name": "Inkling (high)", "org": "Thinking Machines", "score": 0.42498500000000006, "colour": "#686868", "ink": "#ffffff", "label": "42", "new": false}, {"name": "DeepSeek V4.1 Flash (high)", "org": "DeepSeek", "score": 0.24269, "colour": "#3f51e0", "ink": "#ffffff", "label": "24", "new": false}]};
const LOGOS = {"Anthropic": {"icon": "anthropic", "fill": "#141413", "viewBox": "0 0 24 24", "inner": " ", "rule": "evenodd"}, "OpenAI": {"icon": "openai", "fill": "#141413", "viewBox": "0 0 24 24", "inner": " ", "rule": "evenodd"}, "Meta": {"icon": "meta-color", "viewBox": "0 0 24 24", "inner": " ", "rule": null}, "Google": {"icon": "google-color", "viewBox": "0 0 24 24", "inner": " ", "rule": null}, "DeepSeek": {"icon": "deepseek-color", "viewBox": "0 0 24 24", "inner": " ", "rule": null}, "xAI": {"icon": "xai", "tile": "#111111", "shape": "square", "viewBox": "0 0 24 24", "inner": " ", "rule": "evenodd"}, "Z.ai": {"icon": "zai", "tile": "#111111", "shape": "square", "viewBox": "0 0 24 24", "inner": " ", "rule": "evenodd"}, "Moonshot": {"icon": "kimi", "tile": "#111111", "shape": "square", "viewBox": "0 0 24 24", "inner": " ", "rule": "evenodd"}, "StepFun": {"icon": "stepfun", "tile": "#01d9cf", "shape": "circle", "viewBox": "0 0 24 24", "inner": " ", "rule": "evenodd"}, "Undisclosed": {"icon": "unionalpha-color", "viewBox": "0 0 24 24", "inner": " ", "rule": null}, "Cognition": {"icon": "devin-color", "viewBox": "0 0 24 24", "inner": " ", "rule": null}, "Fireworks": {"icon": "fireworks-color", "viewBox": "0 0 24 24", "inner": " ", "rule": null}, "Stealth": {"wordmark": ["STEALTH"]}, "Thinking Machines": {"wordmark": ["THINKING", "MACHINES"]}, "Alibaba": {"wordmark": ["QWEN"]}};
for(const [key,DATA] of Object.entries(ALL)){
const svg=document.getElementById('bars-'+key), W=1200, H=420;
diff --git a/cases.html b/cases.html
index d8a13b1eb627c534741bc3a21657fc3fac913db0..d0d921c05d5631b0d339861f09b13adfc8561d19 100644
--- a/cases.html
+++ b/cases.html
@@ -82,12 +82,12 @@ table.cs td,table.cs th{white-space:nowrap}.panel code{font-size:.82em;backgroun
.in-open{border:1px solid var(--primary);background:#fff;color:var(--primary-dark);border-radius:8px;padding:.45rem .9rem;font:600 .85rem Inter,sans-serif;cursor:pointer}
.in-open:hover{background:var(--primary);color:#fff}
@media(max-width:900px){.in-i2s{grid-template-columns:1fr}}
-
Skip to benchmark explorer
+Skip to benchmark explorer
+evaluator scored that scene. 4 cases, 29 scenes in 3D. Pick a case, then an agent.
@@ -99,7 +99,7 @@ evaluator scored that scene. 6 cases, 41 scenes in 3D. Pick a case, then an agen
-
Skip to benchmark explorer
-
Constructing and Editing3D Scenes with Code Benchmarking coding agents that turn text and reference images into engine-native 3D scenes. Measuring spatial reasoning, task fulfillment and precise control of scene state.
Your browser does not support embedded video. Watch the Code4Scene video .Code4Scene in 90 seconds Construction, editing and evaluation
+
Skip to benchmark explorer
+
Constructing3D Scenes with Code Benchmarking coding agents that turn open-ended scene descriptions into engine-native 3D scenes in Unreal Engine. Measuring spatial reasoning, task fulfillment and physical validity.
-
Overall Text-to-Scene Image-to-Scene
Code4Scene Paper evaluation · score out of 100 Swipe to view all 14 configurations →
Sorted by score. Bar labels round scores ×100; the table preserves three decimal places. One color per model provider.
-
Quality against the cost of one case · original 95-case public evaluation
Swipe to explore all configurations →
Solid: the frontier — nothing cheaper scores higher. One point per configuration, at its evaluated reasoning effort. Cost is the mean USD per case; overall averages the two task means.
-
Search configurations 14 configurations
Ranking
Overall = ½ Text-to-Scene + ½ Image-to-Scene over the 95 public cases; its cost is weighted the same way.
# Agent configuration Overall Text-to-Scene Image-to-Scene Cost / case 1 OpenAI GPT-6 Astra (max)Pareto OpenAI 0.619
0.724 0.515 $12.90 2 Gemini Gemini 3.8 Flash (high)Pareto Google 0.619
0.657 0.581 $1.92 3 Claude Claude Fable 5.1 (max)Anthropic 0.606
0.788 0.424 $10.62 4 Claude Claude Opus 5 (max)Anthropic 0.593
0.718 0.468 $19.48 5 OpenAI GPT-5.6 Sol (high)OpenAI 0.550
0.707 0.393 $2.97 6 Meta Muse Spark 1.3 (medium)Pareto Meta 0.502
0.646 0.358 $0.09 7 Z.ai GLM-5.3 Flash (max)open weights Z.ai 0.415
0.509 0.320 $0.40 8 Qwen Qwen 3.8 27B (thinking off)open weights Alibaba 0.378
0.557 0.200 $0.38 9 Grok Grok 4.6 (high)xAI 0.376
0.567 0.184 $2.29 10 Qwen Qwen 3.8 27B (thinking on)open weights Alibaba 0.371
0.516 0.226 $0.30 11 Thinking Machines
- Inkling (high)Thinking Machines 0.312
0.425 0.198 $0.48 12 Gemini Gemma 4 31B (thinking on)open weights Pareto Google 0.298
0.471 0.126 $0.03 13 Gemini Gemma 4 31B (thinking off)open weights Pareto Google 0.285
0.455 0.115 $0.02 14 DeepSeek DeepSeek V4.1 Flash (high)open weights DeepSeek 0.227
0.243 0.210 $0.12
20 public cases. Case score = 0.2 × Detailed + 0.6 × Overview + 0.2 × Physical. Detailed and Overview are means of the per-case scores.
# Agent configuration Score Detailed Overview Physical Cost / case 1 Claude Claude Fable 5.1 (max)Pareto Anthropic 0.788
0.725 0.772 0.898 $18.16 2 OpenAI GPT-6 Astra (max)OpenAI 0.724
0.707 0.733 0.712 $22.47 3 Claude Claude Opus 5 (max)Anthropic 0.718
0.631 0.716 0.811 $34.50 4 OpenAI GPT-5.6 Sol (high)Pareto OpenAI 0.707
0.632 0.710 0.772 $4.07 5 Gemini Gemini 3.8 Flash (high)Pareto Google 0.657
0.637 0.664 0.655 $2.49 6 Meta Muse Spark 1.3 (medium)Pareto Meta 0.646
0.598 0.661 0.651 $0.10 7 Grok Grok 4.6 (high)xAI 0.567
0.522 0.576 0.585 $4.33 8 Qwen Qwen 3.8 27B (thinking off)open weights Alibaba 0.557
0.507 0.530 0.688 $0.49 9 Qwen Qwen 3.8 27B (thinking on)open weights Alibaba 0.516
0.505 0.460 0.694 $0.33 10 Z.ai GLM-5.3 Flash (max)open weights Z.ai 0.509
0.403 0.537 0.533 $0.71 11 Gemini Gemma 4 31B (thinking on)open weights Pareto Google 0.471
0.412 0.416 0.694 $0.03 12 Gemini Gemma 4 31B (thinking off)open weights Pareto Google 0.455
0.438 0.370 0.725 $0.02 13 Thinking Machines
- Inkling (high)Thinking Machines 0.425
0.393 0.337 0.722 $0.20 14 DeepSeek DeepSeek V4.1 Flash (high)open weights DeepSeek 0.243
0.217 0.161 0.513 $0.19
75 public cases (25 indoor, 50 outdoor). Case score = 0.8 × Repair F1 + 0.2 × Physical. Repair F1 counts target actors restored to the withheld ground truth.
# Agent configuration Score Repair F1 Physical Indoor Outdoor Cost / case 1 Gemini Gemini 3.8 Flash (high)Pareto Google 0.581
0.527 0.796 0.717 0.513 $1.35 2 OpenAI GPT-6 Astra (max)OpenAI 0.515
0.445 0.796 0.785 0.380 $3.32 3 Claude Claude Opus 5 (max)Anthropic 0.468
0.389 0.786 0.618 0.393 $4.46 4 Claude Claude Fable 5.1 (max)Anthropic 0.424
0.332 0.795 0.535 0.369 $3.09 5 OpenAI GPT-5.6 Sol (high)OpenAI 0.393
0.319 0.689 0.497 0.341 $1.88 6 Meta Muse Spark 1.3 (medium)Pareto Meta 0.358
0.246 0.807 0.633 0.221 $0.07 7 Z.ai GLM-5.3 Flash (max)open weights Z.ai 0.320
0.188 0.848 0.465 0.247 $0.09 8 Qwen Qwen 3.8 27B (thinking on)open weights Alibaba 0.226
0.118 0.660 0.297 0.191 $0.26 9 DeepSeek DeepSeek V4.1 Flash (high)open weights Pareto DeepSeek 0.210
0.066 0.787 0.189 0.221 $0.05 10 Qwen Qwen 3.8 27B (thinking off)open weights Alibaba 0.200
0.083 0.667 0.266 0.166 $0.28 11 Thinking Machines
- Inkling (high)Thinking Machines 0.198
0.095 0.609 0.169 0.213 $0.75 12 Grok Grok 4.6 (high)xAI 0.184
0.044 0.744 0.197 0.178 $0.26 13 Gemini Gemma 4 31B (thinking on)open weights Pareto Google 0.126
0.003 0.615 0.127 0.125 $0.03 14 Gemini Gemma 4 31B (thinking off)open weights Pareto Google 0.115
0.001 0.572 0.112 0.116 $0.02
+
14 coding-agent configurations · paper results on the original 20 public construction cases.
+
Code4Scene Paper evaluation · score out of 100 Swipe to view all 14 configurations →
Sorted by score. Bar labels round scores ×100; the table preserves three decimal places. One color per model provider.
+
Quality against the cost of one case · original 20-case public evaluation
Swipe to explore all configurations →
Solid: the frontier — nothing cheaper scores higher. One point per configuration, at its evaluated reasoning effort. Cost is the mean USD per case.
+
Search configurations 14 configurations
Ranking
20 public cases. Case score = 0.2 × Detailed + 0.6 × Overview + 0.2 × Physical. Detailed and Overview are means of the per-case scores.
# Agent configuration Score Detailed Overview Physical Cost / case 1 Claude Claude Fable 5.1 (max)Pareto Anthropic 0.788
0.725 0.772 0.898 $18.16 2 OpenAI GPT-6 Astra (max)OpenAI 0.724
0.707 0.733 0.712 $22.47 3 Claude Claude Opus 5 (max)Anthropic 0.718
0.631 0.716 0.811 $34.50 4 OpenAI GPT-5.6 Sol (high)Pareto OpenAI 0.707
0.632 0.710 0.772 $4.07 5 Gemini Gemini 3.8 Flash (high)Pareto Google 0.657
0.637 0.664 0.655 $2.49 6 Meta Muse Spark 1.3 (medium)Pareto Meta 0.646
0.598 0.661 0.651 $0.10 7 Grok Grok 4.6 (high)xAI 0.567
0.522 0.576 0.585 $4.33 8 Qwen Qwen 3.8 27B (thinking off)open weights Alibaba 0.557
0.507 0.530 0.688 $0.49 9 Qwen Qwen 3.8 27B (thinking on)open weights Alibaba 0.516
0.505 0.460 0.694 $0.33 10 Z.ai GLM-5.3 Flash (max)open weights Z.ai 0.509
0.403 0.537 0.533 $0.71 11 Gemini Gemma 4 31B (thinking on)open weights Pareto Google 0.471
0.412 0.416 0.694 $0.03 12 Gemini Gemma 4 31B (thinking off)open weights Pareto Google 0.455
0.438 0.370 0.725 $0.02 13 Thinking Machines
+ Inkling (high)Thinking Machines 0.425
0.393 0.337 0.722 $0.20 14 DeepSeek DeepSeek V4.1 Flash (high)open weights DeepSeek 0.243
0.217 0.161 0.513 $0.19
Click a column to sort. Pareto: on the score-against-cost frontier.
-
Loading 3D scene… Drag to orbit · Scroll to zoom Ⅱ
Loading 3D scene… Drag to orbit · Scroll to zoom Ⅱ
Loading 3D scene… Drag to orbit · Scroll to zoom Ⅱ
Loading 3D scene… Drag to orbit · Scroll to zoom Ⅱ
Loading 3D scene… Drag to orbit · Scroll to zoom Ⅱ
Repair Corrupted input Ground truth
Loading 3D scene… Drag to orbit · Scroll to zoom Ⅱ
Repair Corrupted input Ground truth
-
+
Drag any scene to inspect the saved 3D output, or open a case to explore every agent.
+Loading 3D scene… Drag to orbit · Scroll to zoom Ⅱ
Loading 3D scene… Drag to orbit · Scroll to zoom Ⅱ
Loading 3D scene… Drag to orbit · Scroll to zoom Ⅱ
Loading 3D scene… Drag to orbit · Scroll to zoom Ⅱ
+
-
+
Takeaway Spatial Composition remains the weakest requirement family for all 14 agents. Errors persist even when the required objects are present: generating the right objects does not ensure that their relationships satisfy the specification.
Requirement-family score per agent
Identity & Environment Content & Quantity Attributes & Materials Spatial Composition
0.2 0.4 0.6 0.8 1.0 GPT-6 Astra Gemini 3.8 Flash Claude Fable 5.1 Claude Opus 5 GPT-5.6 Sol Muse Spark 1.3 GLM-5.3 Flash Qwen 3.8 27B · off Grok 4.6 Qwen 3.8 27B · on Inkling Gemma 4 31B · on Gemma 4 31B · off DeepSeek V4.1 Flash Mean of 14 0.36 0.67 Each row is one agent's four family scores; the orange dot, Spatial Composition, is the lowest in every row (paper Figure 14b). The objects are there; the relation is not
Loading 3D scene… Drag to orbit · Scroll to zoom Ⅱ
✓ slab roof 0.95 ✓ stone coping 0.82 ✗ slab roof enclosed by the coping 0.78
GPT-5.6 Sol’s saved Egyptian Temple scene. The judge finds the roof and stone coping, but flags their spatial relationship. View the judge’s evidence ↗ Mismatch rate when the required objects are present
Spatial relation 32.0% Distribution 19.1% Composition 4.3% Across the benchmark (paper Figure 6a).
Copy @misc{ye2026code4scene,
title = {Code4Scene: Benchmarking Coding Agents for Constructing and Editing 3D Scenes},
@@ -50,12 +48,12 @@ svg.ch .dim{opacity:.42}svg.ch g.row:hover{opacity:1}svg.ch .hot{fill:var(--c4s-
url = {https://arxiv.org/abs/2609.36777},
}
-