Developer 3s Key Decision Metrics
Completing multi-hop tasks across large document collections requires connecting disjoint evidence scattered across disparate files—such as project approvals in one document, architectural requirements in a second, and incident reports in a third. Existing search agents iterate over flat file collections where documents offer zero clues regarding their cross-file relationships, forcing agents to blindly rediscover inter-document links for every single query while incurring immense token overhead and missing complementary evidence. KAIST and Microsoft researchers introduce CorpusMap, an entity-anchored navigation layer that organizes the document collection around recurring, cross-document entities. By constructing offline Entity Pages that aggregate mentions and link to all referencing source documents, CorpusMap exposes a navigable topological graph that agents traverse to unearth scattered evidence. Evaluated across 7 models and 3 benchmarks, CorpusMap delivers superior evidence discovery and answer accuracy while consuming significantly fewer tokens.
Key Takeaways
- ✓Offline Entity Graph Eradicates Search Redundancy: Pre-resolves recurring entities into bidirectional Entity Pages offline, transforming flat files into a navigable graph and sparing agents from repeatedly rediscovering relationships.
- ✓Substantial Token Savings in Multi-Hop Retrieval: Replaces blind iterative keyword sweeps with surgical graph traversal, drastically reducing total prompt token overhead on complex multi-document questions.
- ✓Cross-Model Superiority Across 7 Backbones: Rigorously benchmarked across seven diverse LLM architectures and three evaluation suites, outperforming raw search and four alternate navigation layers in evidence discovery and fidelity.
Finished reading? Explore benchmark rankings & pricing
Real-world SWE-bench scores & $20/mo vs API cost break-even calculator

Discussion & Comments
0Sign in to join the discussion
Connect with AI developers to exchange benchmark insights.