Conventional single-image 3D scene reconstruction typically produces static meshes or radiance fields (NeRF/3DGS), which lack modularity, CAD interpretability, and programmatic editability. Researchers introduce LEGO-Anything, an Image-to-Code framework where a coding agent autonomously writes and executes Blender Python scripts, renders multiview checks, and iteratively refines the 3D scene program. To systematically evaluate programmatic 3D recovery, the authors release LEGO-Bench across 208 images from 104 realistic indoor and outdoor environments. Diagnosing three recurrent failure modes in agentic generation—weak layout initialization, regressive edits, and uncalibrated self-evaluation—they propose LEGO-Plugin, a training-free harness delivering up to 62.7% relative improvements across evaluated foundation models while supporting zero-shot downstream 3D spatial queries.
Key Takeaways
- ✓Image-to-Code 3D Reconstruction Paradigm: Transcends uneditable neural radiance fields by representing 3D scenes as explicit, human-readable Blender Python programs that can be inspected, queried, and modified.
- ✓LEGO-Plugin Training-Free Harness: Systematically mitigates regressive edits and unreliable agent self-evaluations, delivering up to a 62.7% relative performance gain across state-of-the-art LLMs without fine-tuning.
- ✓Deterministic 3D Vision Querying: The resulting scene programs directly enable deterministic queries for object detection, instance segmentation, and depth readouts, bridging computer vision and CAD simulation.
Finished reading? Explore benchmark rankings & pricing
Real-world SWE-bench scores & $20/mo vs API cost break-even calculator

Discussion & Comments
0Sign in to join the discussion
Connect with AI developers to exchange benchmark insights.