At its Oct 7 Windows event Microsoft shipped an on-device MAI-Code-1.1-Flash (137B total / 6.8B active MoE, 256K context, ~3.3 bits per weight) on RTX Spark PCs, scoring 70.80% SWE-Bench Verified and 66.29% Terminal-Bench 2.1 locally. GitHub HydraFusion will route Copilot tasks between local and cloud models in experimental preview later in October, with no inference charge for local calls; MXC agent containment went GA the same day.

Key Takeaways

  • ✓On-device: 70.80% SWE-Bench Verified (cloud 72.6%) and 66.29% Terminal-Bench 2.1 (cloud 62.9%); GPT-OSS-120B scored 32.0% / 23.6% in the same test (Microsoft Command Line)
  • ✓Footprint: ~3.3 bits/weight mixed precision plus DFlash2 speculative decoding; 75.5GB peak memory at 256K; 923.5 / 769.8 tok/s prompt processing at 64K / 128K
  • ✓Routing: HydraFusion local+cloud routing reaches the Copilot app, Copilot CLI and VS Code in experimental preview later in October, with no inference charge for local calls (Windows Blog)
  • ✓Available now: Copilot CLI 1.0.94-0+ discovers local Ollama models via /model (tool calling + streaming required) (GitHub Changelog)
  • ✓Security: MXC agent containment is GA on Windows 11, already supported by Codex, Copilot, OpenClaw and LM Studio, with Claude Code and Hermes Agent among those to follow
Microsoft takes MAI-Code-1.1-Flash local: 137B MoE scores 70.8% SWE-Bench Verified on-device, GitHub Copilot to route between local and cloud
🖼️Official Media
Click to view high-res
🧭

Turn your technical choice into a development budget

Compare 40 dev plans & simulate token costs vs $20/mo subscriptions

🔬

In-Depth Technical Analysis

At its Oct 7 Windows event Microsoft introduced 'hybrid intelligence': run agent work locally when it makes sense and reach the cloud for frontier capability, with no inference charge for local calls. MAI-Code-1.1-Flash, a 137B-total / 6.8B-active MoE coding model, now ships in an on-device build using ~3.3 bits-per-weight mixed-precision quantization (nearly 80% smaller) and DFlash2 sliding-window speculative decoding on a Windows ARM64 llama.cpp CUDA runtime, keeping a 256K context. In Microsoft's Oct 5 tests on Surface Laptop Ultra it scored 70.80% on SWE-Bench Verified (cloud 72.6%) and 66.29% on Terminal-Bench 2.1 (cloud 62.9%), versus 32.0% / 23.6% for GPT-OSS-120B, with 75.5GB peak memory at 256K and 923.5 / 769.8 tok/s prompt processing at 64K / 128K; these are vendor numbers without independent replication yet. GitHub's HydraFusion orchestrator will route Copilot tasks between local and cloud models in experimental preview later in October across the Copilot app, Copilot CLI and VS Code. Available today: Copilot CLI 1.0.94-0 discovers local Ollama models via /model (tool calling and streaming required; offline mode stays explicit via COPILOT_OFFLINE=true). Microsoft Execution Containers (MXC) went GA on Windows 11 for policy-enforced agent sandboxing, already supported by Codex, GitHub Copilot, OpenClaw, Replit, LM Studio and others, and llama.cpp support was added to Windows ML.

Action HubReady to adopt this in production?

Benchmark side-by-side against alternatives, or calculate monthly token cost vs subscription break-even.