RAG Chunking Strategies

Retrieval-augmented generation works by finding the passages most relevant to a question and giving them to a language model. Before anything can be retrieved, documents must be split into chunks: pieces small enough to embed precisely and fit into the model's context window, but large enough to carry meaning on their own. Chunking is one of the most influential, and most overlooked, decisions in a RAG pipeline.

Chunks that are too large dilute embeddings and waste context. Chunks that are too small lose the surrounding meaning ("it increased by 12%": what did?). Splits that ignore document structure cut tables in half and separate headings from their content. Good chunking preserves meaning boundaries, carries useful metadata, and is validated against real queries.

TL;DR

Quick Example

Structure-aware chunking of Markdown documentation with section context:

Each chunk stays within one section, carries its heading path for context and citations, and records who may see it.

Core Concepts

Why Chunk at All?

Chunking Methods

Size and Overlap

There's no universal best size. It depends on the embedding model, document type, and question style:

Metadata

Every chunk should carry metadata used for filtering (product, version, language, date range, tenant), access control (only retrieve what the user may see), citations (URL, title, page, section), and freshness (updated timestamps for recency boosts). Metadata filters in the vector database often matter as much as embedding quality.

Parent-Child and Small-to-Big Retrieval

Index small chunks (sentences or short passages) for precise matching, but return their parent (the full section or a window of neighboring chunks) to the model. You get precise retrieval without starving the model of context. Variants include sentence-window retrieval and hierarchical indexes (summaries → sections → passages).

Contextual Retrieval

Chunks often lose meaning out of context ("The company's revenue grew 3% over the previous quarter": which company, which quarter?). Contextual chunk headers prepend the document title and section path. Contextual retrieval goes further: an LLM writes a short, chunk-specific context sentence for each chunk before embedding and keyword indexing. Anthropic's published results showed large reductions in retrieval failures, especially combined with hybrid search and reranking. Prompt caching of the full document keeps the generation cost reasonable.

Special Content Types

Best Practices

Clean Before You Chunk

Strip navigation, boilerplate, cookie banners, repeated headers and footers, and tracking noise. Garbage text produces garbage embeddings and wastes tokens.

Keep Stable, Deterministic Chunk IDs

Derive IDs from document ID plus section or position, so re-ingestion updates or deletes the right vectors instead of duplicating them. Re-chunk only changed documents.

Evaluate Chunking Empirically

Build a set of real questions with known supporting passages, and compare chunking strategies by retrieval recall@k and answer quality. Intuition about "good chunk size" is often wrong for a given corpus. See RAG evaluation.

Match Chunks to the Embedding Model

Embedding models have maximum input lengths, and often perform best well below them. Chunks exceeding the limit get truncated silently.

Common Mistakes

Fixed Character Splits Through Structure

The exception list is separated from its rule and mixed with another section. Split on structure first, then size.

One Chunk per Document

Embedding whole long documents yields vague vectors that match many queries loosely and flood the context with irrelevant text. Split into focused passages.

Losing Access Control in the Index

Chunking a restricted document without copying its permissions into chunk metadata lets retrieval surface confidential passages to anyone. Enforce ACL filters at query time. See AI guardrails.

FAQ

What's the best chunk size for RAG?

It depends on your documents, questions, and embedding model. A common starting point is 300–600 tokens with 10–20% overlap, split on structural boundaries. Then evaluate a few sizes against real queries and pick what maximizes retrieval recall and answer quality.

Does chunk overlap help?

Usually a little. It reduces the chance that an answer spanning a boundary is split across two chunks that each score poorly. Too much overlap inflates the index and returns near-duplicate chunks. Structure-aware splitting reduces the need for overlap.

Is semantic chunking better than recursive chunking?

Not reliably. Semantic chunking can help with long unstructured text, but it costs more to compute, and studies show mixed results against well-tuned recursive or structure-aware splitting. Evaluate on your data before adopting it.

Do long-context models make chunking unnecessary?

Not entirely. For a few documents, you can often skip retrieval and send them whole. For large or changing corpora, retrieval over chunks remains far cheaper and faster, and chunking quality still determines what gets retrieved. See context windows.

Related Topics

References