Developer 3s Key Decision Metrics
On 2026-10-02 ggml-org announced llama.cpp server support for Decision Models via /v1/systemone (PR #29818, merged 2026-10-02T09:56Z). Send a state plus typed questions; get option probabilities in one forward pass. Day-0 models: Julia-1, Laya, Kev-4B, lev, OpenJev (vision). Cloudflare Clef is planned next.
Key Takeaways
- ✓Official post 2026-10-02 + merged PR #29818 (2026-10-02T09:56Z, +2139/−15)
- ✓Endpoint POST /v1/systemone with choice/score/noul; probabilities + confidence; output_tokens=0
- ✓Day-0 GGUFs: Julia-1 / Laya / Kev-4B / lev / OpenJev; median ~3–43 ms/question on RTX PRO 6000 (blog table)
- ✓Distinct from Ollama SystemOne & Cloudflare Clef hosting; Clef listed as next in llama.cpp
- ✓Start: llama serve -hf ggml-org/Kev-4B-GGUF then curl /v1/systemone
Finished reading? Explore benchmark rankings & pricing
Real-world SWE-bench scores & $20/mo vs API cost break-even calculator
Project Links & Resources
Direct AccessIn-Depth Technical Analysis
Core Background & Industry Pain Points
Agent pipelines spend many steps on routing, gating, and step verification. Chat models burn tokens and need brittle parsing. Decision models score the options you supply in one forward pass. System One (Jev) standardized the API; Ollama 0.35 and Cloudflare Clef covered container/edge hosting. On 2026-10-02, ggml-org announced native llama.cpp server support; PR #29818 merged at 2026-10-02T09:56Z.
Architecture & Mechanisms
New endpoint POST /v1/systemone: state (text/JSON/chat, optional images) + typed questions (choice / score / noul). Responses include probabilities and confidence with output_tokens: 0. Decision heads ride existing embedding graphs via GGUF metadata; server logic lives in server_decision_context. Router mode picks a model per request. Blog: Cloudflare Clef is next (complementary to Workers AI Clef, not a rehash).
Benchmarks & Measured Numbers
No invented SWE-bench scores. Official RTX PRO 6000 medians: Julia-1 3 ms, Laya 5 ms, Kev-4B 12 ms, lev 36 ms, OpenJev 43 ms per question. PR parity tables show small probability drift vs reference. Specs: Julia-1 144M, Laya 421M, Kev-4B/lev on Qwen3.5-4B-class bases, OpenJev 27B with vision (CC BY-NC 4.0). GGUFs under ggml-org/*-GGUF.
Developer Quickstart
Upgrade llama.cpp → llama serve -hf ggml-org/Kev-4B-GGUF → curl /v1/systemone. Describe option criteria; calibrate confidence cutoffs; quantize after merging the decision head. See server README for the full reference.

Discussion & Comments
0Sign in to join the discussion
Connect with AI developers to exchange benchmark insights.