Developer 3s Key Decision Metrics
As language-model agents gain write access to operating systems, cloud APIs, and databases, an insidious attack vector emerges: actions that appear innocuous in isolation can cause catastrophic damage when executed after earlier actions alter system permissions, configuration files, or database records. Researchers introduce SEAD, a unified state-based control framework for analyzing attacks and defenses in tool-using agents. The attacker module, DART, decomposes malicious goals into locally benign steps guided by execution feedback, improving attack success rates by 18.8 to 35.9 percentage points. To counter this, the defender module, SAGE, dynamically dispatches safe read-only queries to verify system state before authorizing actions. SAGE preserves 95.79% of benign workflows while intercepting 92.73% of harmful trajectories, reducing live executable attack success from 48.0% down to 4.0%. Code and benchmarks are open-sourced on GitHub.
Key Takeaways
- ✓State-Based Security Formulation: Formalizes agent safety as partially observed state control, recognizing that previous actions alter system states to make subsequent benign-looking actions lethal.
- ✓DART Adversarial Trajectory Search: Decomposes malicious objectives into locally plausible steps, leveraging real execution feedback to boost attack success by 18.8–35.9 percentage points across four major models.
- ✓SAGE Proactive Read-Only Defense: Investigates critical system states via safe queries prior to action approval, preserving 95.79% of benign tasks while crushing live attack success from 48.0% to 4.0%.
Finished reading? Explore benchmark rankings & pricing
Real-world SWE-bench scores & $20/mo vs API cost break-even calculator

Discussion & Comments
0Sign in to join the discussion
Connect with AI developers to exchange benchmark insights.