Vision-language model (VLM) agents control robots via visual feedback and motion primitives, but repetitive model queries and redundant visual frames inflict massive token overhead. Peking University researchers introduce PyRUA-Lean, an interactive code-execution framework coupling feedback-driven primitive composition with selective observation. Rather than emitting atomic tool calls, the agent composes robot primitives and learned vision-language-action (VLA) policies into executable Python cells that execute localized conditional checks and retries, requesting sensory images only upon necessary replanning. Across 700 tasks on LIBERO-PRO, RoboTwin 2.0, and RoboCasa365, PyRUA-Lean boosts success from 63.1% to 71.7% under identical budgets, slashing LLM calls by 49% and input tokens by 65%.

Key Takeaways

  • ✓Replaces atomic tool calls with interactive Python cells combining kinematic primitives and local retry loops
  • ✓Elevates overall task completion from 63.1% to 71.7% across 700 benchmark instances on LIBERO-PRO and RoboCasa365
  • ✓Cuts cloud model calls by 49% and slashes visual input token consumption by 65% while smoothing physical motion
🧭

Finished reading? Explore benchmark rankings & pricing

Real-world SWE-bench scores & $20/mo vs API cost break-even calculator

🔬

In-Depth Technical Analysis

核心背景与行业痛点

Vision-Language-Action (VLA) robotics typically relies on iterative step-by-step tool calling: at each action tick, high-resolution multi-view images are piped into an LLM/VLM, which returns a single motion primitive. This creates twin pathologies in robotics: massive token bloat from streaming static visual frames, and crippling latency loops where minor grip misalignments trigger multi-second round-trip network delays instead of local retries.

架构亮点与底层机制

Peking University researchers introduce PyRUA-Lean to replace atomic API calls with interactive code-execution units:

  1. Interactive Code Cells: The planner emits Python code blocks embedding localized control loops, coupling deterministic kinematic primitives with learned VLA sub-policies.
  2. Edge-Side Conditional Checks: Embedded try-except blocks and sensor conditionals (e.g., gripper torque thresholds) execute local retries instantly at the edge without querying the cloud model.
  3. Selective Observation Gateways: The code explicitly triggers request_observation() only upon catastrophic branch failures or major phase transitions, suppressing redundant continuous frame streaming.

权威 Benchmark 与实测跑分对比

Benchmarked across 700 simulation instances on LIBERO-PRO, RoboTwin 2.0, and RoboCasa365 using GPT-6 Astra:

  1. 14% Relative Success Rate Surge: Boosts aggregate task completion from 63.1% to 71.7% under matched budget constraints.
  2. 65% Input Token Reductions: On mutually solved tasks, PyRUA-Lean cuts model invocation counts by 49% and slashes input token expenditure by 65%.
  3. Fluid Motion Continuity: Eliminating latency bottlenecks produces smooth, uninterrupted physical execution trajectories.

开发者实战落地与开箱指南

PyRUA-Lean is open-sourced on GitHub with ready-to-deploy ROS/ROS2 adapters. Robotics developers can register hardware driver primitives into Python environments, allowing foundation models to generate structured control code that maximizes autonomy while minimizing token expenses.