While LLM agents excel in multi-step decision-making, they transfer poorly to unseen environments. Conventional world-model approaches predict future observations, incurring high training overhead and compounding planning errors. EVOKE demonstrates that digital world knowledge is already internalized during pretraining, and failure to transfer stems from single-goal post-training that encourages superficial contextual shortcuts. By holding environment states fixed while ranking candidate actions across diverse counterfactual goals, EVOKE compels policies to elicit internalized world dynamics, boosting generalization across unseen environments.

Key Takeaways

  • ✓Replaces heavy explicit world models by eliciting internalized causal dynamics via fixed-state goal diversity
  • ✓Improves task success by 14.3 percentage points on unseen environments while cutting redundant exploration steps by 26%
  • ✓Achieves superior generalization using only 40% of the training trajectory data compared to standard single-goal recipes
🧭

Finished reading? Explore benchmark rankings & pricing

Real-world SWE-bench scores & $20/mo vs API cost break-even calculator

🔬

In-Depth Technical Analysis

核心背景与行业痛点 LLM agents struggle when transferring to unfamiliar environments. Conventional world-model approaches predict future visual/state observations, suffering from high training burdens and compounding roll-out errors. Because digital world knowledge is already internalized within base LLMs during pretraining, the core hurdle lies in post-training: single-goal supervision per state drives agents toward superficial contextual habits rather than causal reasoning. ### 架构亮点与底层机制 EVOKE supplies targeted pressure to elicit latent world knowledge through goal diversity at fixed states: 1. Fixed-State Counterfactual Objectives: Holds the environmental state and conversation history invariant while introducing alternative goals. 2. Action Preference Re-ranking: Compels the agent to re-rank identical candidate actions under varying goals. Superficial context-matching fails under this regime, forcing the policy to simulate underlying state-action transitions. 3. Direct Decision Supervision: Bypasses explicit observation prediction networks, activating latent world representations directly through preference alignment objectives. ### 权威 Benchmark 与实测跑分对比 Evaluated across multiple multi-step agent decision benchmarks across three model backbones: 1. 14.3% Boost on Unseen Environments: EVOKE lifts task success on completely unseen environments by 14.3 percentage points over standard single-goal post-training while reducing redundant exploratory steps by 26%. 2. Exceptional Sample Efficiency: Matches and exceeds standard full-dataset baselines using only 40% of the trajectory training data. 3. Shortcut Suppression: Probing experiments confirm enhanced state-goal decoupling in representation space, eliminating vulnerability to misleading prompt artifacts. ### 开发者实战落地与开箱指南 EVOKE is open-sourced on GitHub with modular loss implementations and data augmentation scripts. Practitioners building enterprise RPA agents or cross-platform assistants can integrate fixed-state counterfactual goal swapping into existing DPO pipelines to unlock robust out-of-distribution transfer without adding inference latency.