GraphRAG

Standard RAG retrieves text chunks similar to a question. That works well for questions whose answer lives in one or a few passages, and poorly for questions that require connecting information across many documents ("how are these three suppliers related to last year's recall?") or summarizing a whole corpus ("what are the main themes in these 5,000 customer interviews?"). GraphRAG addresses both by building a knowledge graph of entities and relationships from the documents, then retrieving over that structure.

Popularized by Microsoft Research's GraphRAG project, and implemented in graph databases like Neo4j and many frameworks, the approach trades higher indexing cost and complexity for better multi-hop reasoning, corpus-level answers, and explainable connections between facts.

TL;DR

Quick Example

Extracting a graph with an LLM, and a local-search style query in Neo4j Cypher:

The retrieved entities, relationship descriptions, and evidence chunks form the context for the LLM's answer.

Core Concepts

Building the Graph

  1. Chunk documents (see chunking).
  2. Extract entities and relationships per chunk with an LLM using structured outputs, optionally with a predefined schema (ontology) of entity and relation types.
  3. Resolve entities: merge duplicates ("IBM", "International Business Machines", "I.B.M.") with normalization, embeddings, and LLM-assisted matching. This step is critical for graph quality.
  4. Store in a graph database or graph structure, keeping links back to source chunks for citations.
  5. Detect communities: cluster the graph (Leiden or Louvain algorithms, see graph algorithms) hierarchically.
  6. Summarize communities: an LLM writes a report per community at each level, describing key entities, relationships, and claims.

Local Search

For questions about specific entities ("What issues affected the X200 battery supplier?"):

  1. Identify entities in the question (extraction or embedding match against entity descriptions).
  2. Traverse to related entities and relationships (1–2 hops), ranked by relevance and relationship weight.
  3. Gather descriptions, community summaries, and linked source chunks.
  4. Generate an answer from that structured context, with citations.

Global Search

For corpus-wide, sensemaking questions ("What are the major risk themes across all audit reports?"):

  1. Select a level of the community hierarchy.
  2. Map: ask the LLM to answer the question from each community summary independently, with a relevance score.
  3. Reduce: combine the highest-rated partial answers into a final response.

Vector RAG handles these questions poorly, because no single chunk contains "the themes of the whole corpus".

Variants and Hybrids

When to Use GraphRAG

Best Practices

Define a Schema for Your Domain

Unconstrained extraction yields inconsistent entity types and relationship names. A small ontology (the entity and relationship types you actually need) improves consistency and makes queries reliable.

Invest in Entity Resolution

Duplicate or fragmented entities break traversal: the same company as five disconnected nodes can't show connections. Combine normalization rules, embedding similarity, and LLM verification, and review high-impact merges.

Keep Provenance

Link every entity and relationship to the chunks it came from. Answers need citations, and extraction errors need to be traceable. See LLM hallucinations.

Evaluate Against Vector RAG

Measure both approaches on your questions, including multi-hop and global ones. GraphRAG's benefits are real for certain question types, but they don't justify its cost everywhere. See RAG evaluation.

Common Mistakes

Underestimating Indexing Cost

Extracting entities and relationships from every chunk, then summarizing every community, can require LLM calls numbering in the hundreds of thousands for large corpora. Estimate token costs up front, use smaller models for extraction, and index incrementally.

Treating Extracted Graphs as Ground Truth

LLM extraction misses relationships, invents weak ones, and misattributes facts. Treat the graph as a noisy index into the source text, and ground final answers in the original chunks.

Using GraphRAG for Simple Lookups

"What's our parental leave policy?" doesn't need graph traversal. Route simple questions to vector or hybrid retrieval, and reserve graph queries for relational and global questions.

FAQ

What's the difference between GraphRAG and regular RAG?

Regular RAG retrieves text chunks by semantic or keyword similarity to the question. GraphRAG first builds a knowledge graph of entities, relationships, and community summaries from the corpus, then retrieves by traversing that structure, or by summarizing communities. That enables multi-hop and corpus-wide questions that chunk similarity handles poorly.

Do I need a graph database for GraphRAG?

Not strictly. Microsoft's GraphRAG reference implementation stores its graph and summaries in files and tables. Graph databases like Neo4j help when you need flexible traversal queries, incremental updates, or to combine the extracted graph with existing structured data.

How expensive is GraphRAG?

Indexing is the main cost: LLM calls for extraction, entity resolution, and community summaries across the whole corpus, often an order of magnitude more than embedding-based indexing. Query costs vary: local search is modest, and global search, which reads many community summaries, can be expensive per question.

Can GraphRAG and vector search be combined?

Yes, and that's the most common production design. Vector or hybrid search finds relevant entities or chunks, graph traversal adds connected context, and a reranker selects the final context. Routing by question type (lookup vs relational vs global) keeps costs down.

Related Topics

References