Conventional single-image 3D scene reconstruction typically produces static meshes or radiance fields (NeRF/3DGS), which lack modularity, CAD interpretability, and programmatic editability. Researchers introduce LEGO-Anything, an Image-to-Code framework where a coding agent autonomously writes and executes Blender Python scripts, renders multiview checks, and iteratively refines the 3D scene program. To systematically evaluate programmatic 3D recovery, the authors release LEGO-Bench across 208 images from 104 realistic indoor and outdoor environments. Diagnosing three recurrent failure modes in agentic generation—weak layout initialization, regressive edits, and uncalibrated self-evaluation—they propose LEGO-Plugin, a training-free harness delivering up to 62.7% relative improvements across evaluated foundation models while supporting zero-shot downstream 3D spatial queries.

Key Takeaways

  • ✓Image-to-Code 3D Reconstruction Paradigm: Transcends uneditable neural radiance fields by representing 3D scenes as explicit, human-readable Blender Python programs that can be inspected, queried, and modified.
  • ✓LEGO-Plugin Training-Free Harness: Systematically mitigates regressive edits and unreliable agent self-evaluations, delivering up to a 62.7% relative performance gain across state-of-the-art LLMs without fine-tuning.
  • ✓Deterministic 3D Vision Querying: The resulting scene programs directly enable deterministic queries for object detection, instance segmentation, and depth readouts, bridging computer vision and CAD simulation.
🧭

Finished reading? Explore benchmark rankings & pricing

Real-world SWE-bench scores & $20/mo vs API cost break-even calculator

🔬

In-Depth Technical Analysis

Core Background & Industry Pain Points Reconstructing a 3D scene from a single input image is fundamental to robotics, simulation, and computer graphics. While Neural Radiance Fields (NeRF) and 3D Gaussian Splatting (3DGS) produce photorealistic renderings, they hit fundamental boundaries in engineering production: 1. Uneditable Black-Box Representations: Tens of millions of unconstrained Gaussian splats or point clouds cannot be natively edited, layered, or parametrically adjusted in standard DCC tools like Blender or Maya; 2. Absence of Semantic Hierarchy & Physical Structure: Implicit fields lack semantic abstractions (e.g., distinguishing a coffee cup sitting atop a desk), failing to support collision detection, mechanical queries, or interactive scene manipulation. ### Architectural Highlights & Underlying Mechanics The authors introduce LEGO-Anything, reframing 3D scene reconstruction as an iterative code generation loop executed by an autonomous coding agent: 1. Image-to-Code Closed Loop: Using Blender Python as the underlying domain-specific language, the coding agent parses the image, drafts an initial scene generation script, invokes the Blender kernel, renders multiview camera passes, and iteratively debugs the scene program; 2. Simulator-Grounded LEGO-Bench: Spans 208 images across 104 realistic indoor and outdoor simulation environments, rigorously scoring artifact code validity, visible-surface geometry fidelity, and rendered visual appearance; 3. Training-Free LEGO-Plugin Harness: Diagnostic traces revealed three systemic failures: poor spatial initialization, regressive edits (degrading previously correct objects), and uncalibrated self-critiques. LEGO-Plugin imposes structured layout constraints and multi-angle differential anchors without requiring model retraining. ### Benchmark & Experimental Validation - LEGO-Bench Evaluation: Across state-of-the-art vision-language models, baseline GPT-6-astra establishes the frontier, achieving 53.4% indoor and 39.6% outdoor scores; - 62.7% Relative Performance Leap: Integrating the plug-and-play LEGO-Plugin improves performance across all six evaluated foundational models, delivering relative gains of up to 62.7% in overall scene score; - Zero-Shot 3D Vision Queries: Within the LEGO-World testbed, executing reconstructed Blender programs yields deterministic bounding boxes, instance masks, and metric depth readouts, demonstrating the utility of programmatic 3D scene representation. ### Engineering Takeaways & Practical Guide - Paper & Benchmark Access: The research and experimental framework are cataloged under arXiv:2609.36380; - Impact on 3D Asset Creation Pipelines: Game developers and architectural visualization teams can deploy this Image-to-Code loop to transform 2D concept art directly into structured .blend scenes with editable node trees, lighting rigs, and textures; - Embodied Simulation Integration: Because Blender scripts define clean collision meshes and physical mass properties, the generated scene programs easily convert into URDF or Isaac Sim formats for robotic manipulation policy testing.