Large language model serving exhibits bifurcated hardware constraints: prefill is compute-bound, benefiting from high-throughput low-precision tensor operations (e.g., NVFP4), while decode is memory-bandwidth-bound, requiring ultra-compact weight storage (e.g., 2-3 bit weight-only) to minimize token-generation memory traffic. Uniform quantization forces suboptimal compromises across both phases. Researchers from IST Austria introduce Disaggregated Quantization (DQ, arXiv:2609.26333), decoupling computation formats, quantization precision, and physical memory placement across prefill and decode. Paired with Offloaded Disaggregated Prefill (ODP) to stream NVFP4 prefill weights asynchronously from NVMe SSDs, DQ delivers a 1.78x time-to-first-token (TTFT) speedup on Qwen3.8-27B at 8K prompt length within llama.cpp, while supporting disaggregated serving in vLLM on architectures up to 2.8T parameters.

Key Takeaways

  • ✓Pioneers Disaggregated Quantization (DQ) to specialize precision formats independently for compute-bound prefill and bandwidth-bound decode
  • ✓Introduces Offloaded Disaggregated Prefill (ODP), streaming NVFP4 weights from SSD to slash 8K prompt TTFT by 1.78x in llama.cpp
  • ✓Boosts MMLU-Pro by 32.5 points and MMMU-Pro by 35.3 points on compact 2-3 bit decode checkpoints, scaling to 2.8T architectures
Disaggregated Quantization: Specializing LLM Prefill and Decode via DQ and Offloaded Prefill for 1.78x TTFT Acceleration
🖼️Official Media
Click to view high-res
🧭

Not enough VRAM? Compare cloud API and self-hosting costs

Compare 29+ dev plans & simulate token costs vs $20/mo subscriptions

🔬

In-Depth Technical Analysis

核心背景与行业痛点

Serving modern large language models forces an irreconcilable conflict between prefill and decode phases. Prefill is heavily compute-bound, demanding maximal tensor FLOPS to ingest extended multi-turn prompts, while decode is strictly memory-bandwidth-bound, where token latency is governed entirely by memory bus transfer speeds. Existing quantization schemes impose uniform precision across both phases, forcing painful compromises: aggressive low-bit quantization (e.g., 2-3 bit) accelerates decode but wrecks prefill mathematical precision, while higher precision bloats memory traffic during rollout.

架构亮点与底层机制

Researchers from IST Austria introduce Disaggregated Quantization (DQ, arXiv:2609.26333), decoupling precision formats across execution stages:

  1. Execution-Stage Disaggregation: Specializes computation formats independently: decode adopts compact 2-3 bit weight-only representations without activation quantization to eliminate memory bus bottlenecks, while prefill employs compute-native low-precision arithmetic (e.g., NVFP4) to unlock peak hardware FLOPS.
  2. Offloaded Disaggregated Prefill (ODP): Resolves single-GPU VRAM limits by streaming specialized NVFP4 prefill weights directly from NVMe SSDs over PCIe, amortizing weight transfer seamlessly across long-context prompt computation.
  3. Frontier Architecture Validation: Validates unified format disaggregation across models from Qwen3 and Gemma 3 to massive 2.8-trillion-parameter MoE foundation architectures.
  4. Dual Runtime Integration: Seamlessly integrates with disaggregated vLLM distributed setups and local llama.cpp environments.

权威 Benchmark 与实测跑分对比

Benchmarked across MMLU-Pro, MMMU-Pro, and end-to-end production serving traces:

  1. 1.78x Speedup in Time-to-First-Token (TTFT): Across 8K prompt lengths in llama.cpp on Qwen3.8-27B, ODP delivers a 1.78x TTFT acceleration over standard weight-only baselines.
  2. Catapults Low-Bit Accuracy by 35 Points: Coupling an NVFP4 prefiller with compact 2-3 bit GGUF decoders restores precision dramatically, boosting MMLU-Pro by 32.5 points and MMMU-Pro by 35.3 points.
  3. Zero VRAM Bloat: Achieves dramatic latency reductions without increasing runtime active GPU memory footprints.

开发者实战落地与开箱指南

Disaggregated Quantization represents an essential architectural milestone for AI infrastructure engineers and edge AI builders. System teams can implement DQ within vLLM to disaggregate prefill and decode nodes with tailored arithmetic, while local developers running models via llama.cpp can leverage NVMe streaming to run large 27B-70B models at unprecedented long-context speeds.

Action HubReady to adopt this in production?

Benchmark side-by-side against alternatives, or calculate monthly token cost vs subscription break-even.