Elasticsearch
Elasticsearch is a distributed search and analytics engine built on Apache Lucene. You send it JSON documents; it indexes every field so you can run full-text search with relevance ranking, filter and sort across billions of documents, and compute aggregations (counts, histograms, percentiles) in milliseconds. It's the engine behind product search, site search, log analytics in the ELK/Elastic Stack, and security analytics.
Elasticsearch is not a primary database. It's eventually consistent, has no multi-document transactions, and reindexing is a normal operation. The standard pattern is to keep your source of truth in a database like PostgreSQL and stream changes into Elasticsearch for search.
TL;DR
- Elasticsearch stores JSON documents in indexes and builds an inverted index for fast full-text search.
- Mappings define field types; analyzers control how text is tokenized — design both deliberately.
- Relevance uses BM25 scoring; tune with boosts,
multi_match, and function scores. - Aggregations power facets, dashboards, and analytics.
- Data is split into shards with replicas for scale and resilience.
- Keep a separate source of truth and sync with change data capture or an outbox; plan for reindexing.
Quick Example
Create an index with an explicit mapping, add a product, and search it:
The english analyzer stems "running" and "shoes" so they match "run" and "shoe"; name^3 weights title matches three times higher; filters don't affect scoring and are cached.
Core Concepts
The Inverted Index
For each field, Lucene maps every term to the list of documents containing it — like the index at the back of a book. Searching "shoe" looks up one term instead of scanning every document. Numeric, date, and geo fields use BKD trees for fast range queries; keyword fields support exact matching, sorting, and aggregations through doc values.
Mappings and Analyzers
An analyzer = character filters + a tokenizer + token filters (lowercase, stop words, stemming, synonyms). The same text can be indexed several ways using multi-fields (name as text, name.raw as keyword). Dynamic mapping guesses types from the first document — convenient for exploring, risky for production.
Queries and Relevance
- Query context scores documents (
match,multi_match,match_phrase). - Filter context includes or excludes without scoring (
term,range,exists) and is cacheable. boolcombinesmust,should,filter, andmust_not.- BM25 scores by term frequency, rarity across the index, and field length.
- Tune with field boosts,
function_score(recency, popularity), synonyms, and rescoring.
Aggregations
Bucket aggregations (terms, date_histogram, range) group documents; metric aggregations (avg, sum, percentiles, cardinality) compute over them. They power faceted navigation ("Footwear (42)") and Kibana dashboards.
Shards, Replicas, and the Cluster
An index is split into primary shards distributed across nodes; each primary has replica shards on other nodes for failover and read throughput. Shard count is fixed at index creation (changing it needs split/shrink or a reindex), so size shards deliberately — roughly 10–50 GB each is a common guideline. Writes become searchable after a refresh (default one second), which is why Elasticsearch is near real-time.
Keeping Search in Sync
Use index aliases so you can build a new index in the background and switch atomically.
Best Practices
Define Mappings Explicitly
Turn off or restrict dynamic mapping for production indexes to avoid field explosions and wrong types (a zip code mapped as a number, for example).
Filter, Then Score
Put exact criteria in filter clauses. They're faster, cached, and don't distort relevance.
Reindex Behind Aliases
Write to products_v2, backfill, verify, then flip the products alias. Zero downtime, easy rollback.
Use Time-Based Indexes for Logs
For log aggregation and metrics, use data streams with index lifecycle management (hot → warm → cold → delete) to control cost.
Secure the Cluster
Enable authentication, TLS, and role-based access. Exposed, unauthenticated clusters are a recurring source of data leaks.
Measure Relevance
Keep a set of real queries with expected results and track metrics like precision at k or NDCG as you tune.
Common Mistakes
Using It as the System of Record
Losing an index should mean "reindex from the database," not "restore from backup and hope."
Too Many Small Shards
Thousands of tiny shards waste heap and slow the cluster. Consolidate indexes and size shards sensibly.
Deep Pagination With from/size
Paging to result 100,000 is expensive. Use search_after with a point-in-time for deep scrolling.
Mapping Explosions
Indexing arbitrary user-supplied keys as fields creates thousands of mappings. Use flattened or key/value nested structures.
Wildcard Queries With Leading *
*shoe scans the whole term dictionary. Use n-gram analyzers or the wildcard field type if you need substring search.
Comparison
FAQ
What is Elasticsearch used for?
Full-text search (products, documents, sites), log and event analytics, observability dashboards through Kibana, security analytics, and increasingly hybrid keyword-plus-vector search.
Is Elasticsearch a database?
It stores and retrieves data, but it's designed as a search and analytics engine. It lacks multi-document transactions and is near real-time, so most systems keep a primary database as the source of truth.
What's the difference between text and keyword fields?
text fields are analyzed into tokens for full-text search. keyword fields store the exact value for filtering, sorting, and aggregations. Many fields use both via multi-fields.
How many shards should an index have?
Size shards to roughly 10–50 GB and keep the total shard count per node modest. Start small; you can split or reindex later as data grows.
Should I use Elasticsearch or PostgreSQL full-text search?
If search is a secondary feature on modest data, Postgres full-text search avoids running another system. Choose Elasticsearch when you need advanced relevance tuning, facets at scale, or search across very large datasets.
Related Topics
- PostgreSQL Full-Text Search — Search without a separate engine
- Change Data Capture — Keeping the index in sync with the database
- Log Aggregation — Elasticsearch as a log store
- Vector Databases — Semantic search and embeddings
- Databases — Choosing the right store for each job