Skip to content

Which Graph RAG to Choose: GraphRAG v3.1.2 vs LightRAG vs HippoRAG 2 — Design, Cost, and Selection

Aug 25, 2026 1 min
TL;DR Same 'knowledge graph + retrieval' label, three different bets: Microsoft GraphRAG v3.1.2 pays indexing cost for global summarization, LightRAG cuts cost with dual-level retrieval and incremental updates, HippoRAG 2 turns RAG into growing associative memory via PPR — this guide splits the trade-offs by component with four query modes, indexing pipelines, and a selection matrix.
Table of Contents
  1. The short version: three graphs, three products
  2. Microsoft GraphRAG v3.1.2: cook documents into a summarizable graph
  3. LightRAG: keep the graph light, make incremental updates cheap
  4. HippoRAG 2: treat RAG as growing memory
  5. Against vector-only and LongRAG: when graphs are waste
  6. How to choose: three axes
  7. Architecture
  8. Takeaway
  9. References

🌏 中文版

Vector search finds similar passages but misses connections — "who is related to whom" and "what does the whole corpus say" require links between documents. Graph RAG fills that gap by extracting entities and relations into a graph and retrieving over it. The question is how to pay for the graph — this guide puts three mainstream approaches side by side so you can choose by cost, update frequency, and question type.

You will get: the full indexing and four-query model of Microsoft GraphRAG v3.1.2, why LightRAG achieves cheap incremental updates, how HippoRAG 2 turns RAG into continual memory, and a matrix for "when to use a graph, when vectors are enough, and when you need neither."

The short version: three graphs, three products

GraphRAG is the heavy, global-summarization pipeline. LightRAG is the light, incrementally updatable engineering compromise. HippoRAG 2 is the continual-learning memory system. The split visible in papers and docs is where you pay: GraphRAG pays at index time (graph + community summaries) for coverage on global questions; LightRAG and HippoRAG deliberately reduce LLM calls and shift cost toward query time.

If you remember one thing: the decision hinges on update frequency and question shape, not leaderboard numbers — daily-changing vs. quarterly-stable corpora and point lookups vs. corpus-wide summaries lead to opposite choices.

Microsoft GraphRAG v3.1.2: cook documents into a summarizable graph

Microsoft GraphRAG (docs | releases v3.1.2) structures the corpus before querying it. As of v3.1.2 (2025-08-21) the canonical indexing flow is:

TextUnit → Entity / Relationship extraction → Leiden clustering → Community Summary (bottom-up)

Documents are split into TextUnits (overlapping chunks). Each TextUnit is passed to an LLM to extract entities and relations, deduplicated into a graph, then partitioned with the Leiden algorithm into hierarchical communities. Each community's entities and relations are summarized bottom-up with an LLM. The docs put it directly: GraphRAG builds a knowledge graph and then generates community summaries bottom-up — those summaries are the real index for global queries, and they explain why initial indexing is expensive.

Compared with alternatives, GraphRAG keeps both "graph + summaries" layers compared with vector-only and LightRAG (arXiv:2410.05779), which drops the heavyweight community reports to save cost, and HippoRAG 2 (arXiv:2502.14802), which replaces build-time summarization with query-time diffusion. The shared failure mode is extraction quality — wrong entities downstream corrupt everything.

Good fit: relation-dense, global-question workloads such as regulation, medical literature, financial research, and enterprise wikis. Poor fit: frequently changing corpora (daily updates), one-off demos, or purely factoid single-passage QA — the graph tax never pays back there.

Config sample:

# settings.yaml — GraphRAG v3.x indexing (illustrative)
input:
  type: csv
  file_pattern: ".*\\.csv$"
chunks:
  size: 1200
  overlap: 100
  group_by_columns: [id]
extract_graph:
  model_id: gpt-4o-mini
  prompt: "extract_graph.txt"
  max_gleanings: 1
cluster_graph:
  max_cluster_size: 10
  use_lcc: true
summarize_descriptions:
  model_id: gpt-4o-mini
  max_length: 500
community_reports:
  model_id: gpt-4o-mini
  max_length: 2000
  max_input_length: 8000
# Index and query (v3 CLI)
graphrag index --root ./ragtest
graphrag query --root ./ragtest --method global "What common risks appear across these contracts?"
graphrag query --root ./ragtest --method local  "What is the relationship between Company A and Company B?"

v3.1.2 ships four query modes, as defined in the docs:

  • Global: answers corpus-wide summaries from community reports.
  • Local: entity-centric graph expansion for point lookups and relation tracing.
  • DRIFT: dynamic hybrid — starts global, then drills local.
  • Basic: pure vector similarity search as a baseline / cheap fallback.

Limitations: indexing still costs many LLM calls; v3's streaming and storage backends (including the new CosmosTableProvider with namespace partitioning) improve throughput and tenant isolation without eliminating the cost. Pilot indexing on a subset before full builds. v3 also carries breaking changes — read breaking-changes.md before upgrading.

LightRAG: keep the graph light, make incremental updates cheap

LightRAG (paper arXiv:2410.05779, v3 2025-04-28) keeps entity/relation extraction but removes GraphRAG's most expensive layer — community reports — and compensates with dual-level retrieval and incremental updates.

Dual-level means retrieving on two tracks: low-level on entities (precise relations and attributes) and high-level on topics/concepts (coverage), then merging. This gives LightRAG point-lookup precision plus reasonable global coverage without pre-built summaries. Against GraphRAG it may trail on extreme global summarization but wins on cost for mixed workloads; against HippoRAG 2, LightRAG's graph is more explicit and its update path is more direct, while HippoRAG leans toward memory and associative diffusion.

Good fit: knowledge bases that change often (product docs, support KBs, research notes) where budget is tight but you need more than vectors. Poor fit: one-shot massive summarization demanding strict global consistency, or extremely noisy corpora where lightweight extraction degrades quickly.

Sample usage (Python, illustrative):

# pip install lightrag-hku
from lightrag import LightRAG, QueryParam
from lightrag.llm.openai import gpt_4o_mini_complete
from lightrag.utils import EmbeddingFunc
import openai

rag = LightRAG(
    working_dir="./lightrag_cache",
    llm_model_func=gpt_4o_mini_complete,
    embedding_func=EmbeddingFunc(
        embedding_dim=3072,
        max_token_size=8192,
        func=lambda texts: openai.embeddings.create(
            model="text-embedding-3-large", input=texts
        ).data[0].embedding
    ),
)

# Incremental writes — call insert repeatedly, no full rebuild
rag.insert("LightRAG supports incremental updates; deletion removes only the relevant subgraph.")
rag.insert(["Second batch...", "Third batch..."])

# Dual-level retrieval: mix traverses low + high together
result = rag.query("What dependencies does Product A have?", param=QueryParam(mode="mix"))
print(result)

# Deletion (backend-dependent): clears only related entities/edges
# rag.delete_by_doc_id("doc-123")

Limitations: 39k+ stars on GitHub, and post-2026-05 it merged RAGAnything for MinerU/Docling multimodal chunking and additional backends — but the core constraint remains: documents never extracted into entities will not become retrievable knowledge. Fusion parameters for mix mode need tuning on your own data.

HippoRAG 2: treat RAG as growing memory

HippoRAG 2 (paper arXiv:2502.14802, ICML 2025) takes its metaphor from the hippocampus: retrieval is not a one-off lookup but accumulating associative memory. The core is Personalized PageRank (PPR).

Passages and entities form a joint graph. At query time, a few seed nodes launch PPR random walks that diffuse across the graph; passages mapped from high-scoring nodes are returned for generation. The paper reports ~7% gains on associative tasks over strong embedding baselines — the signal matters less than the implication: treating "how memory is organized" as a first-class decision enables non-parametric continual learning where new knowledge arrives as nodes and edges without retraining.

Against GraphRAG's build-time summaries, HippoRAG shifts computation to query time; against LightRAG's dual-level, HippoRAG's duality is "text similarity + graph diffusion"; against vector-only and LongRAG (large chunks + long context), HippoRAG shines on cross-document multi-hop while not necessarily beating single-passage fact extraction where vectors or long context already suffice.

Good fit: multi-document associative QA, research KBs, personal/org memory that grows over time and must retain historical context. Poor fit: single-hop factoid workloads where a managed vector store's latency matters more than association.

Sample usage (illustrative):

from hipporag import HippoRAG

rag = HippoRAG(
    llm_model="gpt-4o-mini",
    embedding_model="text-embedding-3-large",
    graph_type="openie",  # or llm-based extraction
)

# Index passages and entities together
rag.index(docs=[
    {"id": "doc1", "text": "Drug A inhibits protein X, which interacts with Y."},
    {"id": "doc2", "text": "Protein Y is overexpressed in disease B."},
])

# Retrieval: PPR diffusion over the graph, then mapped passages
answer = rag.query("Is Drug A indirectly related to disease B?", method="ppr")
print(answer)

# Continual memory: new knowledge extends the graph incrementally
rag.index([{"id": "doc3", "text": "New study links Y to Z."}])

Limitations: PPR step count and damping are hyperparameters — over-diffusion injects noise. The graph grows over time and needs weighting / forgetting policies for old nodes. Quality still depends on extraction; noisy extraction plus diffusion just spreads the noise further.

Against vector-only and LongRAG: when graphs are waste

The intuition "vectors not enough → add a graph" overestimates the return curve.

Vector-only is simple, cheap, and operable. Many factoid and single-passage QA workloads are already served. Its blind spot is relations — vectors do not know "A cites B" or "A and C belong to the same group." LongRAG (Xiao et al., 2024) takes a different route: large chunks + long-context models preserve boundary-crossing information for small corpora in one shot, at the cost of tokens and latency.

Graph payoff concentrates in two quadrants:

Question typeVector / LongRAG enough?Graph incremental value
Single-passage fact (article number, API param)Yes — vectors fastestNone, wasted cost
Cross-document association (citation chain, org relations)No — chains missedClear win for PPR / graph traversal
Corpus-wide summary (trends, risk rollups)LongRAG viable for small corporaGraphRAG community summaries most robust
Frequently updated KBVectors incremental; LongRAG must re-stuffLightRAG / HippoRAG incrementality wins

Attribution check: do not credit graphs alone for "answering global questions" — LongRAG can answer global questions on small corpora; graphs earn their keep by maintaining global consistency and traceable associations at scale and under churn. On small, static corpora, graphs are often over-engineering.

Actionable step: run a 50-question mixed eval (summary / associative / single-fact) across three baselines — vector vs. LongRAG (large chunks) vs. your chosen graph — and split results by question type before deciding which quadrant justifies paying for a graph.

How to choose: three axes

AxisChoose GraphRAG v3.1.2Choose LightRAGChoose HippoRAG 2
Cost toleranceOK to pay heavy first indexing for global precisionTight budget, minimize LLM callsModerate, query-time compute OK
Update frequencyLow (monthly/quarterly) fineHigh (daily / write-as-you-go) preferredContinuously growing memory preferred
Question shapeCorpus-wide summaries, hierarchical reportsMixed (point lookup + moderate summary)Cross-document multi-hop, memory recall
OperationsMust run pipeline + community reportsLightest — local delete/insertMust tune PPR + memory growth

Quick rules:

  • Ask update frequency first — daily → LightRAG / HippoRAG; infrequent → GraphRAG.
  • Ask question shape second — global rollups → GraphRAG; multi-hop association → HippoRAG; mixed → LightRAG.
  • Ask ops capacity last — if you cannot tune graphs and PPR, start vector-only and prove with eval that relations are the bottleneck before adding a graph.

Architecture

                        Query

        ┌─────────────────┼─────────────────┐
        ▼                 ▼                 ▼
   GraphRAG v3.1.2    LightRAG         HippoRAG 2
   ─────────────      ────────         ──────────
   Doc → TextUnit     Doc → Entity/Rel Doc + Passage → Graph
         │  (extract)        │  (light)        │  (openie/LLM)
         ▼                   ▼                 ▼
   Entity/Rels Graph    Entity Graph       Entity + Passage Graph
         │                   │                 │
    Leiden clustering   no Community Summary  PPR diffusion
         │                   │                 │
   Community Summary    dual-level retrieval  Passage recall
   (bottom-up LLM)      low + high fusion    (diffusion result)
         │                   │                 │
        Global/Local/DRIFT/Basic  mix / hybrid   PPR-ranked
         │                   │                 │
         └─────────────────┼─────────────────┘

                   LLM generation + citations

                  eval / observability / cost

Costs shift horizontally: GraphRAG pays on the left (index time), HippoRAG pays on the right (query time), LightRAG keeps both sides thin.

Takeaway

The choice is not "which scores highest" but "where you are willing to pay and how the corpus grows." Microsoft GraphRAG v3.1.2 fits teams that accept indexing cost for global explainability. LightRAG fits teams whose corpus keeps growing under tight budget and ops constraints. HippoRAG 2 fits teams treating RAG as a memory system that must accumulate cross-document associations over time. If unsure, start vector-only and measure the share of failures caused by missing relations on a mixed eval — only pay for a graph when that share is large; otherwise better chunking or LongRAG is the cheaper win.

References