On Oct 9 Prime Intellect shipped a Rust rewrite of Prime Agent, its open-source coding agent harness. A root orchestrator ran 2,209 sub-agents across 10,000+ Prime Sandboxes and ~228.7B GLM-5.3 tokens over two weeks, with differential TUI, harness and protocol parity tests gating every merge. Vendor-measured cold time-to-type dropped from 737.8ms to 55.8ms and post-startup memory from 607MB to 106MB; Windows (beta) and Homebrew installs were added.
Key Takeaways
- ✓Scale: 2,209 agents, 16,758 agent-to-agent messages, 228.70B tokens (192.99B/1,981 agents for the port, 35.70B/228 for perf hillclimbing) across 10,000+ Prime Sandboxes
- ✓Vendor-measured: cold time-to-type 737.8ms→55.8ms (13.22x), first paint 722.8ms→23.6ms, post-startup RSS 607.4MB→106.0MB, install 172.1MB→59.6MB
- ✓Each task ran a Planner→Implementer→adversarial Reviewer (different model)→Verifier state machine; failures loop back to the implementer
- ✓A 3-day target-free hillclimb loop logged 144+ experiment/audit records and merged 69+ performance changes
- ✓Codebase now 9 crates, largest file ~2,500 lines (was ~15,000); native Windows (beta) and Homebrew installs, still open source

Key Decision Metrics at a Glance
Heavy Claude Code use: compare subscription limits and API bills
Compare 40 dev plans & simulate token costs vs $20/mo subscriptions
Project Links & Resources
Direct AccessIn-Depth Technical Analysis
Prime Intellect rewrote Prime Agent, its open-source long-running coding agent harness, from TypeScript to Rust, with the agent doing most of the work. A root orchestrator that wrote no product code split the port into topologically ordered tasks; each passed through a planner, an implementer in its own worktree, an adversarial reviewer on a different model, and a verifier running parity checks in a fresh Prime Sandbox. Humans mainly built objective verification: terminal-frame TUI diffs, transcript and model-request diffs, daemon protocol checks and a per-component feature audit. In total 2,209 agents exchanged 16,758 messages and used 228.7B GLM-5.3 tokens over about two weeks, followed by a three-day hillclimbing loop that merged 69+ performance changes. Vendor-measured results: cold time-to-type 737.8ms to 55.8ms, first paint 722.8ms to 23.6ms, post-startup memory 607.4MB to 106.0MB, install size 172.1MB to 59.6MB; Prime Intellect cautions that cross-harness comparisons lack a common benchmark. The codebase is now nine crates with per-session worker isolation, native Windows (beta) and Homebrew installs.
Benchmark side-by-side against alternatives, or calculate monthly token cost vs subscription break-even.
Discussion & Comments
0Sign in to join the discussion
Connect with AI developers to exchange benchmark insights.