On Oct 11 the vLLM Semantic Router (vllm-sr) team released the Decision 3.0 family (d3, d3-flash, d3-mini, d3-nano, d3-lite, d3-edge) on Hugging Face under Apache-2.0. Built on Qwen3.5/3.8, the models take text, images and video and return probabilities for choice, yes/no and score questions in one call without generating text. Vendor-reported: the 27B d3 scores 64.7 on Jev Decision Index 0.3.1, ahead of Perplexity Decider v1.1 (62.8) and Decision 2.0 (55.9), with a 55 ms median text latency on one AMD MI325X.

Key Takeaways

  • ✓Six sizes: d3-edge 0.59B, d3-lite 0.85B, d3-nano 2.2B, d3-mini 4.5B, d3-flash 8.4B, d3 26.1B, all Apache-2.0 (HF collection)
  • ✓Vendor-reported Jev Decision Index 0.3.1: d3 64.7 vs Perplexity Decider v1.1 62.8, Jev 60.1, Decision 2.0 55.9
  • ✓Vision Index 0.3.1: d3 71.6 vs Perplexity Decider v1.1 70.6 and JEV-27B-VL 69.6
  • ✓Median latency on one AMD MI325X, batch 1: d3 55 ms text / 273 ms image / 553 ms 10-second video; d3-edge 6.7 ms text
  • ✓No text generation: every option gets a probability, and several Choice / Yes-No / Score questions share one call
vLLM Semantic Router open-sources Decision 3.0: six multimodal decision models from 0.6B to 27B, #1 on Jev Decision Index in its own evaluation
🖼️Official Media
Click to view high-res
🧭

Not enough VRAM? Compare cloud API and self-hosting costs

Compare 40 dev plans & simulate token costs vs $20/mo subscriptions

🔬

In-Depth Technical Analysis

Decision 3.0 from the vLLM Semantic Router team is a family of six Apache-2.0 'System 1' decision models (0.59B to 26.09B parameters) fine-tuned from Qwen3.5 and Qwen3.8-27B. Instead of generating text, they take text or JSON plus multiple images and videos (sampled at 2 fps, up to 32 frames) and return a probability for every option of Choice, Yes/No and Score questions, several questions per call. According to the team's own model card, the 27B d3 scores 64.7 on Jev Decision Index 0.3.1 versus 62.8 for Perplexity Decider v1.1, 60.1 for Jev and 55.9 for Decision 2.0, and 71.6 on the vision board versus 69.6 for JEV-27B-VL. Median latency on a single AMD MI325X at batch 1 is 55 ms for text, 273 ms with an image and 553 ms with a 10-second video; the 0.59B d3-edge answers text in 6.7 ms but scores only 28.82 on the public suite. Load it with transformers 5.17 and trust_remote_code, then call model.system_one(). Small sizes suit gateway routing and tool-call screening, with the 27B model or a general LLM for harder cases. No independent benchmark replication yet.

Action HubReady to adopt this in production?

Benchmark side-by-side against alternatives, or calculate monthly token cost vs subscription break-even.