LMBuild: Evaluating LLM Agents for Generating Buildable and Functional Structures
Abstract
LLM-based agents are increasingly capable of generating complex 3D structures, with the potential to reshape how objects are designed and realized in the physical world. Yet, producing elegant geometry is fundamentally different from producing objects that can be built and perform their intended functions. Existing evaluations largely focus on geometric quality while overlooking physical realizability. We introduce LMBuild, a benchmark for evaluating LLM agents on generating buildable and functional structures. LMBuild represents generated objects as assembled structures comprising part decompositions, joints, materials, and sequences. To support reproducible evaluation, we provide a unified framework consisting of: (1) an interactive environment in which agents can use tools to retrieve, create, and place components to construct objects; (2) a curated benchmark that repurposes established CAD datasets and augments them with knowledge from Wikipedia; and (3) a evaluation framework covering structural soundness, functional affordance, design quality, and physical realization. Evaluations across 30 systems reveal several intriguing findings: (a) Soundness and alignment are no longer the primary bottlenecks for frontier closed-source models, while functional affordance and physical operability remain substantially more challenging; (b) stronger models more effectively create new components, whereas weaker models tend to rely on retrieval; and (c) providing functional specifications substantially improves part completeness, kinematics, and physical operability. These results show that generating real-world structures requires deeper reasoning about functional affordances, mechanics, and designing and creating novel components. We expect LMBuild to provide a foundation for measuring progress and incentivizing research toward agents that generate buildable and functional structures.
Community
HF Daily Papers submission: LMBuild (arXiv 2610.04292)
Links
- GitHub: https://github.com/Lumos-Jiateng/LMBuild
- Project page: https://lumos-jiateng.github.io/LMBuild/
- Dataset: https://huggingface.co/datasets/Lumos-Jiateng/LMBuild
Comment
๐๏ธ LMBuild: can LLM agents build objects that actually work, and not just look right?
LLM agents can already produce good-looking 3D geometry. But a wheelchair whose wheels don't turn, or a desk that can't be assembled, isn't a real object. LMBuild evaluates agents on generating buildable and functional structures: part decompositions, joints, materials and assembly sequences.
๐ง Interactive environment: agents use tools to retrieve catalog parts, create new parts, place and adjust them, then declare joints, materials and an assembly sequence.
๐ฆ Benchmark: we repurpose established CAD datasets (LEGO/BrickComposer, Fusion 360, Artiverse, PartNeXt, open-source product CAD) and ground them in Wikipedia knowledge. LMBuild-Core has 200 objects.
๐ 12 hierarchical metrics in four groups: Soundness (connectivity, collision, stability), Affordance (functional geometry, parts, kinematics), Design (decomposition, aesthetics, alignment; the VLM judges are human-calibrated) and Realization (assembly sequence, material, simulation-based operability).
๐ 30 systems evaluated: frontier closed APIs, open-source (M)LLMs, and domain-specific generators (BrickGPT, PartCrafter, PartPacker, Cube3D, PhysX-Anything โฆ).
Key findings:
1๏ธโฃ Soundness and visual alignment are no longer the main bottleneck for frontier models. Functional affordance and physical operability are: even the best systems score only about 20 on simulation-based operability.
2๏ธโฃ Stronger models create; weaker models retrieve. Allowing part creation helps frontier models but can hurt weaker ones.
3๏ธโฃ Stating functional requirements explicitly substantially improves part completeness, kinematics and operability.
Code, data and evaluation are all open source. Feedback is welcome! ๐
Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper