Developer 3s Key Decision Metrics
Conventional vision-language-action (VLA) and world-action models map raw observations directly to robot control commands, resulting in brittle policies vulnerable to slight layout variations because environment state, constraints, and error recovery remain implicitly buried inside latent vectors. Researchers introduce Physical Coding (arXiv:2609.35432), bridging digital software agent methodologies with embodied robotics: representing physical environments as 'Code as World' and orchestrating execution via 'Code as Policy.' The team builds HexaAnything, an autonomous agent orchestrating perception, planning, and control tools while utilizing in-the-loop environmental feedback. Evaluated across RoboCasa365, PhyBench, and real-world AgileX dual-arm robots executing laboratory physics experiments, HexaAnything systematically outperforms baseline VLAs and internalizes physical execution experience into self-evolving model weights.
Key Takeaways
- ✓Introduces Physical Coding, treating world state as structured code and robotic execution as algorithmic policies
- ✓Builds HexaAnything, outperforming leading VLA baselines like XR-1 across Composite-Unseen splits on RoboCasa365
- ✓Successfully deploys to real-world AgileX dual-arm robots, autonomously executing complex laboratory physics workflows
Turn your technical choice into a development budget
Compare 29+ dev plans & simulate token costs vs $20/mo subscriptions
Project Links & Resources
Direct AccessIn-Depth Technical Analysis
核心背景与行业痛点
Vision-Language-Action (VLA) and World-Action Models (WAM) govern modern robotics by directly mapping camera frames and text goals to joint velocities or end-effector trajectories. However, implicitly encoding world dynamics, task progress, and exception handling inside low-level continuous action vectors renders policies brittle: slight alterations in object poses or camera viewpoints trigger catastrophic execution crashes. Crucially, raw action streams prevent high-level symbolic reflection, debugging, or reusable procedural recall.
架构亮点与底层机制
Researchers introduce Physical Coding (arXiv:2609.35432), porting digital software agent principles into embodied physical workflows:
- Code as World: Represents spatial geometries, object relations, and dynamic progress metrics explicitly as structured, queryable programmatic data structures.
- Code as Policy: Formulates planning, execution loops, safety verifications, and exception rollbacks as executable Python routines that orchestrate specialized tools alongside low-level VLA control primitives.
- HexaAnything Autonomous Orchestrator: Integrates multimodal perception, LLM reasoning, and robotic execution. In-the-loop sensory feedback enables real-time code rewriting, while verified successes crystallize into reusable procedural memory.
- Hierarchical Self-Evolution: Synthesizes execution traces into curated post-training datasets to update model weights (HexaModel), optimize tool harnesses, and iteratively refine physical task protocols.
权威 Benchmark 与实测跑分对比
Benchmarked across simulated environments, physical reasoning suites, and physical dual-arm robotic systems:
- Outperforms XR-1 VLA on RoboCasa365: Establishes clear superiority across Composite-Unseen scenarios and overall success metrics on RoboCasa365.
- Weight Internalization Superiority: The harness-trained HexaModel systematically outperforms foundation baselines across every task split.
- Autonomous Physics Experimentation on Dual-Arm AgileX Robots: Deployed on physical hardware, HexaAnything autonomously executes complex physics experiments on PhyBench and diverse tabletop manipulation workflows with lower latency than published baselines.
开发者实战落地与开箱指南
The HexaAnything codebase, simulation harnesses, and robotic control bridges are publicly released. Robotics teams can deploy Physical Coding as an interpretable symbolic middleware over low-level motor primitives, transforming robotic manipulators into self-correcting autonomous coding agents capable of continuous physical self-evolution.
Benchmark side-by-side against alternatives, or calculate monthly token cost vs subscription break-even.

Discussion & Comments
0Sign in to join the discussion
Connect with AI developers to exchange benchmark insights.