Liquid AI released two open-weight multimodal decision models: d1-3B (text+vision) and the experimental d1-omni-600M (text+image or text+audio). Instead of generating tokens they answer in a single forward pass. d1-3B scores 48.57 on Decision Index v0.2.1 (public split), first among sub-10B models and on par with the 12x larger Decider 35B-A3B, and answers in 8 ms on an RTX 4090 and 50 ms on a Jetson Orin Nano, with day-one llama.cpp support.

Key Takeaways

  • ✓d1-3B: 48.57 on Decision Index v0.2.1 public split, #1 under 10B, on par with Decider 35B-A3B; d1-omni-600M: 15.95
  • ✓Text benchmark mean: d1-3B 82.9 vs Decider 4B 81.1; d1-omni-600M 78.4 vs Decider 2B 77.1 (Civil Comments 95.8)
  • ✓Single-question latency: 8 ms RTX 4090, 9 ms MI325X, 16 ms Jetson AGX Thor, 30 ms Apple M5 Pro, 50 ms Jetson Orin Nano
  • ✓Packed throughput: 475 states/s on RTX 4090, 1,106/s on MI325X; three questions cost ~1.3x one
  • ✓Weights on Hugging Face (incl. GGUF and w8a8), LFM Open License v1.0, day-one llama.cpp and NVFP4 support
Liquid AI open-sources Open d1: d1-3B and d1-omni-600M decision models answer in one forward pass, 8 ms on RTX 4090
🖼️Official Media
Click to view high-res
🧭

Turn your technical choice into a development budget

Compare 40 dev plans & simulate token costs vs $20/mo subscriptions

🔬

In-Depth Technical Analysis

Liquid AI open-sourced two models from its d1 decision family on Oct 7, 2026. Unlike generative LFMs, d1 models emit no tokens and return a calibrated answer (yes/no, choice, score) in one forward pass. d1-3B is built on the decoder-only LFM2.5-VL-3B (text+image), using weight averaging with LFM2.5-2.6B and merged seed/data-mix fine-tunes; d1-omni-600M builds on the bidirectional LFM2.5-Encoder-350M with a FastConformer audio encoder and the LFM2.5-VL-450M vision encoder. d1-3B scores 48.57 on Decision Index v0.2.1 (public split), first under 10B and on par with Decider 35B-A3B; its seven-benchmark text mean is 82.9 vs 81.1 for Decider 4B, while d1-omni-600M hits 78.4 vs 77.1 for Decider 2B. Single-question latency is 8 ms on RTX 4090, 9 ms on MI325X, 16 ms on Jetson AGX Thor and 50 ms on Jetson Orin Nano. Weights (including GGUF and w8a8) are on Hugging Face under the LFM Open License v1.0 with day-one llama.cpp and NVFP4 support. Target uses: agent guardrails, reranking, visual inspection and voice-command routing.

Action HubReady to adopt this in production?

Benchmark side-by-side against alternatives, or calculate monthly token cost vs subscription break-even.