Before invoking tools, agentic LLMs navigate a K-way action space: executing calls, clarifying ambiguities, answering directly, or declining. While internal activation steering attempts to align pre-execution tool decisions, aggregate metrics obscure where manipulated latent states land and the severe collateral damage they inflict. Researchers from Aberdeen and Oxford introduce SAKIKO, an auditing framework formalizing representation repair via directional error discovery, router-conditioned intervention, destination-resolved verification, and prospectively frozen statistical licensing. Crucially, destination auditing demonstrates that behavioral movement does not equal repair: an intervention scoring a net +55 gain corrupted more than half of baseline-correct decisions, establishing the necessity of outcome-resolved adjudication.

Key Takeaways

  • ✓Pioneers SAKIKO, the first mechanistic auditing framework for internal interventions in tool-using foundation models
  • ✓Demonstrates that an intervention achieving a net +55 gain corrupted over 50% of baseline-correct agent decisions
  • ✓Introduces destination-resolved adjudication and frozen statistical licensing to prevent harmful steering in production
🧭

Finished reading? Explore benchmark rankings & pricing

Real-world SWE-bench scores & $20/mo vs API cost break-even calculator

🔬

In-Depth Technical Analysis

核心背景与行业痛点

Tool invocation requires agents to navigate complex multi-way action boundaries. To rectify erroneous tool-calling decisions without retraining, researchers increasingly leverage activation steering—injecting directional vectors into residual streams. However, conventional benchmarks rely on aggregate net gain metrics, blinding practitioners to severe latent shifts and collateral damage where manipulated states land on unintended decision spaces.

架构亮点与底层机制

Researchers from Aberdeen and Oxford introduce SAKIKO, an interpretable auditing framework formalizing internal representation repairs across four rigorous pillars:

  1. Directional Error Discovery: Isolates localized geometric steering directions responsible for distinct misclassification routes.
  2. Router-Conditioned Interventions: Restricts activation steering specifically to tokens routed through targeted decision channels rather than applying uniform tensor offsets.
  3. Destination-Resolved Adjudication: Traces the final destination of steered latent vectors, explicitly categorizing outcomes into genuine repairs, over-steered distortions, and corrupted baseline-correct choices.
  4. Frozen Statistical Licensing: Employs strict hypothesis-testing gates to disqualify fragile point estimates.

权威 Benchmark 与实测跑分对比

Audited across seven foundation models on When2Call and MetaTool benchmarks:

  1. Exposing Hidden Collateral Damage: An intervention boasting a +55 net accuracy gain was proven to silently corrupt more than 50% of the baseline-correct decisions it touched.
  2. Zero Random Competitors: None of 59 budget-matched random vectors reproduced the calibrated targeted direction gains across sealed evaluations.
  3. Statistical Disqualification: Promising superficial point estimates on Qwen3-4B and Gemma-2-9B were formally rejected due to finite-sample variance.

开发者实战落地与开箱指南

SAKIKO is open-sourced on GitHub. Engineering teams deploying agentic tool routers or safety alignment steering vectors can integrate SAKIKO's diagnostic probes to audit model interventions, guaranteeing that alignment patches do not compromise baseline production reliability.