From Local to Global: A Graph RAG Approach to Query-Focused Summarization
- Type
- other
- Venue
- arXiv / Microsoft
Summary
GraphRAG: LLM extracts entities/relations/claims into a knowledge graph, Leiden-partitions it hierarchically, and pregenerates community summaries; query-time map-reduce over those summaries answers global QFS that vector RAG misses. Podcast (~1M tok) and news (~1.7M tok) graphs: 8,564/15,754 nodes. vs vector RAG, comprehensiveness win rates ~72–83% and diversity ~62–82% (GPT-4 judge, 125 questions×5). Root-level summaries use ~2–3% of source tokens. Code https://github.com/microsoft/graphrag.
Keywords
graphrag · rag · knowledge-graph · leiden · qfs · microsoft · community-detection · sensemaking
Topics
RAG, knowledge graphs, query-focused summarization
Research notes
- Primary: arxiv abs (cs.CL; also cs.AI, cs.IR). License CC BY 4.0 on HTML at check. Microsoft Strategic Missions and Technologies / Office of the CTO / Microsoft Research. Correspondence daedge/trinhha/newmancheng/joshbradley/achao/moapurva/steventruitt/dasham/robertness/jolarso @microsoft.com. Code https://github.com/microsoft/graphrag (35,498 stars at check; MIT). HF paper page 7 upvotes; unofficial linked model muthuk1/graphrag-inference-hackathon not copied into hf_* fields. Discord posted PDF. Eval uses public podcast transcripts plus MultiHop-RAG news rather than a new hosted corpus, so no datasets_local row.