Elasticsearch

Elasticsearch is a distributed search and analytics engine built on Apache Lucene. You send it JSON documents; it indexes every field so you can run full-text search with relevance ranking, filter and sort across billions of documents, and compute aggregations (counts, histograms, percentiles) in milliseconds. It's the engine behind product search, site search, log analytics in the ELK/Elastic Stack, and security analytics.

Elasticsearch is not a primary database. It's eventually consistent, has no multi-document transactions, and reindexing is a normal operation. The standard pattern is to keep your source of truth in a database like PostgreSQL and stream changes into Elasticsearch for search.

TL;DR

Quick Example

Create an index with an explicit mapping, add a product, and search it:

The english analyzer stems "running" and "shoes" so they match "run" and "shoe"; name^3 weights title matches three times higher; filters don't affect scoring and are cached.

Core Concepts

The Inverted Index

For each field, Lucene maps every term to the list of documents containing it — like the index at the back of a book. Searching "shoe" looks up one term instead of scanning every document. Numeric, date, and geo fields use BKD trees for fast range queries; keyword fields support exact matching, sorting, and aggregations through doc values.

Mappings and Analyzers

An analyzer = character filters + a tokenizer + token filters (lowercase, stop words, stemming, synonyms). The same text can be indexed several ways using multi-fields (name as text, name.raw as keyword). Dynamic mapping guesses types from the first document — convenient for exploring, risky for production.

Queries and Relevance

Aggregations

Bucket aggregations (terms, date_histogram, range) group documents; metric aggregations (avg, sum, percentiles, cardinality) compute over them. They power faceted navigation ("Footwear (42)") and Kibana dashboards.

Shards, Replicas, and the Cluster

An index is split into primary shards distributed across nodes; each primary has replica shards on other nodes for failover and read throughput. Shard count is fixed at index creation (changing it needs split/shrink or a reindex), so size shards deliberately — roughly 10–50 GB each is a common guideline. Writes become searchable after a refresh (default one second), which is why Elasticsearch is near real-time.

Keeping Search in Sync

Use index aliases so you can build a new index in the background and switch atomically.

Best Practices

Define Mappings Explicitly

Turn off or restrict dynamic mapping for production indexes to avoid field explosions and wrong types (a zip code mapped as a number, for example).

Filter, Then Score

Put exact criteria in filter clauses. They're faster, cached, and don't distort relevance.

Reindex Behind Aliases

Write to products_v2, backfill, verify, then flip the products alias. Zero downtime, easy rollback.

Use Time-Based Indexes for Logs

For log aggregation and metrics, use data streams with index lifecycle management (hot → warm → cold → delete) to control cost.

Secure the Cluster

Enable authentication, TLS, and role-based access. Exposed, unauthenticated clusters are a recurring source of data leaks.

Measure Relevance

Keep a set of real queries with expected results and track metrics like precision at k or NDCG as you tune.

Common Mistakes

Using It as the System of Record

Losing an index should mean "reindex from the database," not "restore from backup and hope."

Too Many Small Shards

Thousands of tiny shards waste heap and slow the cluster. Consolidate indexes and size shards sensibly.

Deep Pagination With from/size

Paging to result 100,000 is expensive. Use search_after with a point-in-time for deep scrolling.

Mapping Explosions

Indexing arbitrary user-supplied keys as fields creates thousands of mappings. Use flattened or key/value nested structures.

Wildcard Queries With Leading *

*shoe scans the whole term dictionary. Use n-gram analyzers or the wildcard field type if you need substring search.

Comparison

FAQ

What is Elasticsearch used for?

Full-text search (products, documents, sites), log and event analytics, observability dashboards through Kibana, security analytics, and increasingly hybrid keyword-plus-vector search.

Is Elasticsearch a database?

It stores and retrieves data, but it's designed as a search and analytics engine. It lacks multi-document transactions and is near real-time, so most systems keep a primary database as the source of truth.

What's the difference between text and keyword fields?

text fields are analyzed into tokens for full-text search. keyword fields store the exact value for filtering, sorting, and aggregations. Many fields use both via multi-fields.

How many shards should an index have?

Size shards to roughly 10–50 GB and keep the total shard count per node modest. Start small; you can split or reindex later as data grows.

Should I use Elasticsearch or PostgreSQL full-text search?

If search is a secondary feature on modest data, Postgres full-text search avoids running another system. Choose Elasticsearch when you need advanced relevance tuning, facets at scale, or search across very large datasets.

Related Topics

References