Addressing the fundamental tension where prefill is compute-bound while decode is memory-bandwidth-bound, researchers from ISTA and Neural Magic introduced Disaggregated Quantization (DQ). By assigning compute-native NVFP4 weights to prefill and compact 2-3 bit weights to decode, paired with Offloaded Disaggregated Prefill (ODP) that streams prefill weights from SSD, DQ cuts Time-To-First-Token (TTFT) by 1.78x at 8K context and lifts 1-bit decode accuracy by 32.5 points on MMLU-Pro.
- ✓Stage-Specialized Formats: Eliminates uniform end-to-end quantization by applying compute-native NVFP4 for compute-bound prefill and compact 2-3 bit representations for bandwidth-bound decode token generation.
- ✓Zero-Footprint SSD Streaming: Offloaded Disaggregated Prefill (ODP) dynamically streams specialized prefill weights from high-speed SSDs during prompt ingestion, completely amortizing I/O latency across sequence lengths.
- ✓1.78x TTFT Speedup and 2.8T Scale: Delivers a 1.78x TTFT speedup on Qwen 27B at 8K prompt length in llama.cpp, raises 1-bit decode accuracy by 32.5 points on MMLU-Pro, and validates up to 2.8T parameter scales under vLLM disaggregation.
🧭Finished reading? Explore benchmark rankings & pricing
Real-world SWE-bench scores & $20/mo vs API cost break-even calculator
🔗
Project Links & Resources
Direct AccessDirect access to official project resources and documentation🔬
In-Depth Technical Analysis
核心背景与行业痛点 / Background & Pain Points LLM inference inherently exhibits opposing physical bottlenecks across its two execution phases: prefill is compute-bound over thousands of prompt tokens, demanding maximum Tensor Core arithmetic throughput, while autoregressive decode is memory-bandwidth-bound, constrained by memory traffic per token. Modern quantization schemes enforce uniform formats across both phases, inevitably sacrificing either compute efficiency during prefill or memory bandwidth during decode. ### 架构亮点与底层机制 / Architectural Highlights Researchers from ISTA and Neural Magic formulated Disaggregated Quantization (DQ): 1. Phase-Specialized Arithmetic and Weights: Deploys compute-native NVFP4 weights to maximize arithmetic throughput during prefill, paired with 2-3 bit ultra-compact weights to minimize memory bandwidth consumption during decode; 2. Activation Decoupling on Decode: Demonstrates that eliminating activation quantization strictly during decode restores task accuracy without increasing serving latency; 3. Offloaded Disaggregated Prefill (ODP): For memory-constrained single-GPU setups, ODP streams specialized prefill weights asynchronously from NVMe SSDs, completely overlapping I/O transfer latency with prompt execution and freeing memory immediately for decode caches. ### 权威 Benchmark 与实测跑分对比 / Benchmark & Evaluation Extensive evaluations across Qwen 3 and Gemma 3 on llama.cpp and vLLM reveal breakthrough performance: - Time-to-First-Token (TTFT): Slashes prefill latency by 1.78x over standard weight-only baselines at 8K prompt length on Qwen3.8-27B; - Accuracy Recovery in Low Bits: Adding an NVFP4 prefiller to a 1-bit GGUF decoder boosts MMLU-Pro by +32.5 points and MMMU-Pro by +35.3 points without modifying the decode checkpoint; - Trillion-Parameter Scale: Empirically validates shared-weight format disaggregation up to 2.8 trillion parameters under disaggregated serving configurations. ### 开发者实战落地与开箱指南 / Developer Practical Guide - Technical Reference: Detailed proofs and profiles are available in arXiv preprint 2609.26333; - vLLM Integration: Ideal for vLLM disaggregated serving clusters by provisioning NVFP4 on Prefill workers and 2-3 bit quantization on Decode workers; - Edge Deployment: ODP streaming enables running 27B-70B models with long-context windows on single consumer GPUs with NVMe storage.
Discussion & Comments
0Sign in to join the discussion
Connect with AI developers to exchange benchmark insights.