Strata, an open-source C++ inference engine, runs the 125B-parameter (6B active) Qwen3.8-Flash-Next on a 12 GB gaming GPU plus 32–64 GB RAM, exposing OpenAI, Anthropic and Responses-compatible APIs on localhost so Claude Code and Codex CLI can use it. The repo passed 10k GitHub stars and hit the Hacker News front page (507 points).

Key Takeaways

  • ✓Scale: Qwen3.8-Flash-Next is 125B total / 6B active (512 experts, 10 routed + 1 shared), 262K native context
  • ✓Measured: RTX 5070 12 GB + 64 GB RAM, Q2_0 decodes ~94 tok/s and prefills a 32K prompt at ~2,650 tok/s; RX 9070 XT ~60 tok/s
  • ✓Requirements: 12 GB VRAM and 32 GB RAM minimum, ~70 GB download; the Coder pack drops half the experts and keeps ~91% of SWE-bench Verified per its authors
  • ✓Integration: localhost OpenAI, Anthropic /v1/messages and Responses endpoints plus an MCP server
  • ✓Traction: ~10.7k stars, 946 forks, MIT license, v0.1.39 released Oct 4
🧭

Heavy Claude Code use: compare subscription limits and API bills

Compare 29+ dev plans & simulate token costs vs $20/mo subscriptions

🔬

In-Depth Technical Analysis

Strata is an MIT-licensed C++ inference engine that runs Alibaba's Qwen3.8-Flash-Next (125B total, ~6B active per token, 512 experts) on ordinary gaming PCs with a 12 GB+ NVIDIA or AMD GPU and 32–64 GB RAM. It keeps hot experts in VRAM and the rest in RAM, adaptively swapping them as the conversation evolves; at long context it streams the KV cache from RAM so more experts fit on the GPU, and offers 4-bit/K8V4 KV options, a speculative draft head and an SSD-mapped low-RAM mode. The project's own measurements on an RTX 5070 + 64 GB RAM show ~94 tok/s decode and ~2,650 tok/s prefill at 32K for the Q2_0 pack (IQ3_S ~53 tok/s, Coder ~55 tok/s); an RX 9070 XT reaches ~60 tok/s. A 3090 figure of 100–140 tok/s is the author's estimate, not a measurement, and the Coder pack's claimed 91% retention of SWE-bench Verified is self-reported. It serves OpenAI, Anthropic /v1/messages and Responses-compatible endpoints on localhost:8080, so Claude Code and Codex CLI can point at it directly, and ships an MCP server. The repo has ~10.7k stars and reached the Hacker News front page with 507 points on Oct 4.

Action HubReady to adopt this in production?

Benchmark side-by-side against alternatives, or calculate monthly token cost vs subscription break-even.