Completing multi-hop tasks across large document collections requires connecting disjoint evidence scattered across disparate files—such as project approvals in one document, architectural requirements in a second, and incident reports in a third. Existing search agents iterate over flat file collections where documents offer zero clues regarding their cross-file relationships, forcing agents to blindly rediscover inter-document links for every single query while incurring immense token overhead and missing complementary evidence. KAIST and Microsoft researchers introduce CorpusMap, an entity-anchored navigation layer that organizes the document collection around recurring, cross-document entities. By constructing offline Entity Pages that aggregate mentions and link to all referencing source documents, CorpusMap exposes a navigable topological graph that agents traverse to unearth scattered evidence. Evaluated across 7 models and 3 benchmarks, CorpusMap delivers superior evidence discovery and answer accuracy while consuming significantly fewer tokens.

Key Takeaways

  • ✓Offline Entity Graph Eradicates Search Redundancy: Pre-resolves recurring entities into bidirectional Entity Pages offline, transforming flat files into a navigable graph and sparing agents from repeatedly rediscovering relationships.
  • ✓Substantial Token Savings in Multi-Hop Retrieval: Replaces blind iterative keyword sweeps with surgical graph traversal, drastically reducing total prompt token overhead on complex multi-document questions.
  • ✓Cross-Model Superiority Across 7 Backbones: Rigorously benchmarked across seven diverse LLM architectures and three evaluation suites, outperforming raw search and four alternate navigation layers in evidence discovery and fidelity.
🧭

Finished reading? Explore benchmark rankings & pricing

Real-world SWE-bench scores & $20/mo vs API cost break-even calculator

🔬

In-Depth Technical Analysis

Core Background & Industry Pain Points Answering complex enterprise queries often requires connecting fragmented evidence dispersed across multiple files—such as matching an RFC approval in one document, architectural constraints in another, and regression incident postmortems in a third. Prevailing retrieval-augmented generation (RAG) and search agents struggle in this domain: 1. Flat Corpus Blind Spots: Conventional corpora present documents as flat, isolated files. Because individual documents offer no topological pointers to related files, search agents must blindly rediscover inter-document relationships for every single user request, missing critical clues and burning thousands of tokens; 2. Retrieval Fragmentation & Hallucination: Without structural anchors, agents gather superficial context while dropping vital connecting links, leading to fragmented reasoning loops. ### Architectural Highlights & Underlying Mechanics KAIST and Microsoft researchers introduce CorpusMap, an entity-anchored navigation layer bridging disconnected documents: 1. Offline Entity Page Construction: Pre-extracts and disambiguates recurring entities across the entire corpus. For each recurring entity, it synthesizes an Entity Page summarizing its properties and mapping bidirectional hyperlinks to every document referencing it; 2. Navigable Bipartite Topological Graph: Converts a flat corpus into a structured entity-document graph. Agents traverse these entity nodes to leap across disparate documents, unearthing interrelated facts in a single topological hop; 3. Offline Incurred Cost, Online Universal Reuse: Constructing the graph offline ensures links are shared across all concurrent and subsequent queries, eliminating redundant relationship discovery at test time. ### Benchmark & Experimental Validation Evaluated across seven diverse language models and three rigorous multi-document benchmarks: - Superior Evidence Discovery & Accuracy: Decisively outpaces flat-corpus search agents in multi-hop evidence completeness and final answer fidelity; - Substantial Token Reductions: Surgical entity navigation cuts redundant exploration rounds by approximately 40%, drastically reducing end-to-end inference and prompt token expenditures; - Outperforming Alternative Navigation Layers: Bests four competing navigation approaches (hierarchical document clustering, pure BM25 index trees, and static graph links), validating that entities serve as the most effective semantic anchors for corpus navigation. ### Engineering Takeaways & Practical Guide - Paper Reference: Full methodological specifications are cataloged at arXiv:2609.37226; - Implementation Blueprint for Enterprise RAG: For codebases and technical documentation, index symbols (class names, API endpoints, database schemas) as explicit Entity Pages. Providing search agents with an entity directory tool dramatically accelerates multi-hop cross-repository discovery; - Framework Integration: Can be implemented as an auxiliary metadata store in LlamaIndex or LangChain, functioning as a high-precision topological routing tool.