Despite modern LLMs touting million-token context windows, multi-hop reasoning degrades steeply as inputs grow—a pervasive phenomenon termed 'context rot.' This collapse stems from architectural entanglement, where the immense search burden of contextual grounding exhausts the representational capacity required for complex reasoning. Adobe Research and Penn State introduce DISCO (Grounding-Reasoning Disaggregation via Distributed Long Context Scaling). Inspired by MapReduce and Apache Spark, DISCO partitions long documents across parallel Worker LLMs for localized grounding, coordinated by a central Driver LLM trained via GRPO to formulate extraction plans and synthesize evidence. On the 1M-token RULER-QA benchmark, DISCO maintains 78.4% accuracy where monolithic models collapse, boosts LongBench v2 scores by up to 9.8 points, matches Gemini-3-Pro-Preview performance, and slashes serving costs by over 80%.

Key Takeaways

  • ✓Eliminating Context Rot at Scale: Systematically decouples contextual grounding from reasoning, eliminating attention dispersion and KV cache noise that exhaust model capacity over 1M+ tokens.
  • ✓MapReduce Distributed Agent Topology: A central Driver LLM (trained with GRPO) dynamically generates atomic extraction plans, directing a fleet of parallel Worker LLMs to harvest localized evidence before synthesizing final answers.
  • ✓78.4% Accuracy on 1M Tokens & 80% Cost Reduction: Maintains 78.4% accuracy on 1M RULER-QA where monolithic models collapse, gains up to 9.8 points on LongBench v2, and matches frontier models at over 80% lower inference cost.
🧭

Finished reading? Explore benchmark rankings & pricing

Real-world SWE-bench scores & $20/mo vs API cost break-even calculator

🔬

In-Depth Technical Analysis

Core Background & Industry Pain Points While contemporary frontier LLMs boast million-token context windows, their real-world reasoning degrades precipitously as inputs scale—a critical limitation termed "Context Rot": 1. Entanglement of Grounding and Reasoning: Monolithic Transformers force identical self-attention layers to simultaneously execute needle-in-a-haystack localization (grounding) across millions of noisy tokens and multi-step relational deduction (reasoning). The search burden exhausts representational bandwidth, causing multi-hop reasoning to collapse; 2. Prohibitive Attention & KV Cache Costs: Serving monolithic 1M-token contexts demands colossal KV cache memory and produces quadratic prefill latency, restricting enterprise adoption for massive codebase audits and long-document synthesis. ### Architectural Highlights & Underlying Mechanics Adobe Research and Penn State introduced DISCO (DIStributed long COntext scaling via Grounding-Reasoning Disaggregation), applying distributed MapReduce principles to long-context LLMs: 1. Physical Disaggregation of Compute Roles: Massive context is partitioned across a fleet of parallel, lightweight Worker LLMs that exclusively perform localized grounding without reasoning overhead; 2. GRPO-Trained Central Driver LLM: A central coordinator model trained via Group Relative Policy Optimization (GRPO) maps user queries into atomic extraction tasks, dispatches them across workers, and synthesizes answers purely over clean, validated evidence chunks; 3. KV Cache Isolation: Eliminates massive cross-node KV cache distribution and synchronization bubbles; communication is restricted to concise extracted candidate passages. ### Benchmark & Experimental Validation Empirical results demonstrate transformative gains across massive-context evaluations: - 1M-Token RULER-QA Benchmark: Under extreme 1,000,000-token conditions where monolithic baselines collapse, DISCO retains a robust 78.4% accuracy; - LongBench v2 Superiority: On realistic multi-domain long-form datasets, DISCO beats full-context baselines by up to +9.8 points; - Frontier Performance at >80% Cost Reduction: Matches the reasoning fidelity of frontier models such as Gemini-3-Pro-Preview while slashing operational inference expenditures by more than 80%. ### Engineering Takeaways & Practical Guide - Paper & System Design: Full algorithmic specifications are available in arXiv:2609.33485; - Strategic Guidance for Code & Enterprise Agents: Rather than feeding entire multi-megabyte codebases into expensive monolithic 1M context windows, engineers should structure architectures around lightweight MapReduce worker fleets paired with a reasoning coordinator to maximize accuracy and minimize GPU memory footprints; - Framework Integration: The worker-driver RPC topology is readily adaptable to distributed orchestration frameworks like Ray, LangGraph, and OpenHands.