On 2026-10-03, Microsoft Copilot Studio and Hugging Face released ThinkingBox plus ThinkingBox-Bench (507 tasks × 20 trials): isolated MCP sessions graded by executable terminal-state/side-effect checks, wired into OpenEnv. Vendor board: Claude Opus 5.5 67.16% pass@1; Kimi-K3 93.89% pass@20 but only 13.41% all-20 — one success ≠ reliability. MIT framework microsoft/thinkingbox; CDLA data microsoft/thinkingbox-data; paper arXiv:2608.19741.

Key Takeaways

  • ✓Sources: HF blog 2026-10-03, OpenEnv docs, microsoft/thinkingbox (MIT), thinkingbox-data (CDLA), arXiv:2608.19741
  • ✓507 stateful business workflows × 20 isolated MCP trials each
  • ✓Grade terminal DB/side effects; ~67% of failing traces still terminate cleanly after a mutating tool
  • ✓Opus 5.5 67.16% pass@1; Kimi-K3 broadest pass@20 (93.89%) but 13.41% all-20; GPT-5.4 ~$6.80/dependable task
  • ✓Run: OpenEnv + [email protected] → Typesense + MCP proxy → thinkingbox-eval
🧭

Turn your technical choice into a development budget

Compare 29+ dev plans & simulate token costs vs $20/mo subscriptions

🔬

In-Depth Technical Analysis

Background

Agents that look finished often leave the wrong DB state. On 2026-10-03, Microsoft Copilot Studio and Hugging Face open-sourced ThinkingBox and ThinkingBox-Bench: 507 stateful business workflows (retail, auto insurance, travel, neobank IT, consulting IT/HR), each run 20 times from a clean backend, graded on terminal state and side effects—not a polite closing sentence. Paper: arXiv:2608.19741 (One Success Isn't Reliability).

Mechanism

Isolated MCP tool sessions; simulated users release private facts only when asked. A side-effect extractor plus deterministic executable judges accept any trajectory that hits the required end state and reject wrong/missing/extra effects (477/507 state-only; 30 add narrow response rubrics). The same loop yields sparse RL rewards. OpenEnv exposes reset/step against pinned thinkingbox-bench-v1.0. Licenses: MIT framework, CDLA-Permissive-2.0 data, BSD-3 OpenEnv env.

Benchmarks (vendor tables)

pass@1: Claude Opus 5.5 67.16%, Opus 5 66.50%, GPT-5.4 65.36%, Kimi-K3 57.37%. Breadth vs consistency: Kimi-K3 93.89% pass@20 but only 13.41% all-20; Opus 5 hits 47.53% all-20. ~79.9% of failure signatures are tool-usage/recovery; among clean-terminating mutating failures, 77.61% still have wrong field values. Estimated cost per dependable task: GPT-5.4 ~$6.80, GPT-6 Astra ~$7.45, Opus 5.5 ~$7.80 (OpenRouter list rates).

Playbook

Clone OpenEnv + [email protected]; install tb; start Typesense + MCP Session Proxy; set thinkingbox.yaml; run OpenEnv server + thinkingbox-eval. Keep system errors in a sidecar. For production agents: check terminal state before commit, classify retryable tool errors, shrink the tool surface, require humans on irreversible writes.

Action HubReady to adopt this in production?

Benchmark side-by-side against alternatives, or calculate monthly token cost vs subscription break-even.