Elasticsearch Analyzers & Text Analysis
When a document is indexed into a text field, Elasticsearch doesn't store the string as-is for searching. It runs it through an analyzer that breaks it into tokens (terms) and normalizes them: lowercasing, removing accents, stemming "running" to "run", and optionally adding synonyms. At query time the search input goes through the same, or a compatible, analyzer, and matching happens on tokens. Analysis is why a search for "Café" finds "cafe", and why "shoe" matches "Running Shoes".
Analyzers are the biggest lever for search recall (finding relevant documents) and a major factor in precision. They're set in mappings, so changing them requires reindexing, and it pays to design and test them early.
TL;DR
- An analyzer = character filters → tokenizer → token filters.
- The default
standardanalyzer splits on word boundaries and lowercases; language analyzers (english,german…) add stemming and stop words. - Synonyms (
synonym_graphat search time) broaden recall. Maintain them as reloadable synonym sets. - Edge n-grams or
search_as_you_typepower autocomplete; n-grams allow substring matching, at a storage cost. - Normalizers apply lowercasing and accent folding to
keywordfields for case-insensitive exact matching. - Test with the
_analyzeAPI, and keep index-time and search-time analysis compatible.
Quick Example
A custom analyzer setup for product search with autocomplete:
(The synonyms_set and english_possessive_stemmer filter would be defined via the synonyms API and a stemmer filter, respectively.)
Core Concepts
Anatomy of an Analyzer
- Character filters transform the raw string: strip HTML (
html_strip), map characters (mapping:&→and), or apply regex replacements. - Tokenizer splits text into tokens:
standard(Unicode word boundaries),whitespace,keyword(the whole string as one token),pattern,ngram/edge_ngram,path_hierarchy,uax_url_email, or language-specific tokenizers (ICU, kuromoji for Japanese, smartcn for Chinese). - Token filters modify, add, or remove tokens:
lowercase,asciifolding(é → e),stop, stemmers (porter_stem,snowball, languagestemmer),synonym_graph,word_delimiter_graph(splits "WiFi-6E" into parts),shingle(word pairs),unique,length.
The _analyze API shows the tokens at every step with explain: true.
Built-In Analyzers
Stemming and Stop Words
Stemming reduces words to a root ("running", "runs" → "run"), which increases recall but can conflate distinct words ("university"/"universe" with aggressive stemmers). Light stemmers or lemmatization (via plugins) are gentler. Stop words (the, a, of) were traditionally removed to save space. Modern BM25 handles them reasonably well, and removing them can break phrase queries like "to be or not to be", so use them selectively.
Synonyms
- Equivalent:
sneakers, trainers, running shoes. - Explicit mappings:
tv => television. - Use
synonym_graphat search time. It handles multi-word synonyms correctly, and search-time synonyms can be updated without reindexing, via the Synonyms API (synonym sets) and reloadable analyzers. - Expanding synonyms at index time bloats the index and requires reindexing on every change.
Autocomplete and Partial Matching
Use a different, non-n-gram search analyzer at query time, so the query "sho" isn't itself exploded into n-grams.
Normalizers
keyword fields aren't analyzed, but a normalizer (restricted to character-level filters like lowercase and asciifolding) makes exact matching and aggregations case- and accent-insensitive: emails, tags, usernames.
Multilingual Text
Use separate fields per language (title.en, title.de) with the matching language analyzers, detect language at ingest (the inference processor or an external library), and use ICU plugins for proper Unicode handling. CJK languages need dedicated tokenizers, since whitespace tokenization doesn't work.
Best Practices
Start Simple, Then Tune With Real Queries
Begin with standard or a language analyzer, collect real search queries and failures ("zero results" searches, poor top results), and add synonyms, stemming changes, or decompounding based on evidence.
Keep Index and Search Analyzers Compatible
They may differ (n-grams at index time only, synonyms at search time), but both must produce tokens in the same normalized form: the same lowercasing, folding, and stemming.
Make Synonyms Search-Time and Reloadable
Search-time synonym sets can be updated through the API without reindexing, so merchandisers and search teams can iterate quickly.
Test Analyzers in CI
Add _analyze assertions for tricky inputs (product codes, hyphenated names, accents, plurals) to your mapping tests. Analyzer regressions silently degrade search quality.
Common Mistakes
Stemming Identifiers
Running SKUs, model numbers, or codes ("A-100-X") through text analyzers splits and mangles them. Map identifiers as keyword (with a normalizer), or use word_delimiter_graph deliberately.
Index-Time Synonym Expansion
Adding synonyms to the index analyzer means every synonym change requires a full reindex, and it inflates the index. Prefer search-time synonym_graph.
N-grams on Large Fields
Indexing ngram tokens (min 1, max 20) on long descriptions multiplies index size many times over, and slows indexing. Restrict n-grams to short fields, like names and titles.
FAQ
What is an analyzer in Elasticsearch?
A pipeline that converts text into tokens for the inverted index. It consists of optional character filters (clean the text), exactly one tokenizer (split it into tokens), and optional token filters (normalize, remove, or add tokens). The same process is applied to query text so that queries and documents match on normalized terms.
Should synonyms be applied at index time or search time?
Search time, in most cases. Search-time synonym_graph handles multi-word synonyms correctly, keeps the index smaller, and lets you update synonyms without reindexing. Index-time expansion is rarely worth its costs.
How do I implement autocomplete?
Use a search_as_you_type field, or a subfield analyzed with an edge_ngram filter at index time and a standard analyzer at search time, queried with match_bool_prefix or multi_match of type bool_prefix. For very fast suggestions from a fixed list, the completion suggester works well.
How do I make keyword searches case-insensitive?
Add a normalizer with lowercase (and asciifolding if needed) to the keyword field. Both indexed values and term queries on that field are then normalized, making exact matches and aggregations case-insensitive.
Related Topics
- Elasticsearch — The search engine overview
- Elasticsearch Mappings — Where analyzers are configured
- Elasticsearch Query DSL — Queries that use analyzed fields
- NLP — Text processing fundamentals
- Tokenization — Tokenization in LLMs vs search engines
- PostgreSQL Full-Text Search — Text search configurations in Postgres