MongoDB Aggregation Pipeline

The aggregation pipeline is MongoDB's tool for anything beyond simple finds: grouping and totals, reshaping documents, joining collections, computing running averages, and building reports. A pipeline is an array of stages. Documents flow through them in order, and each stage filters, transforms, groups, or joins its input and passes the result to the next.

It plays the role SQL's SELECT … WHERE … GROUP BY … JOIN … ORDER BY plays in relational databases, expressed as composable steps instead of one declarative statement. Written well, with filtering first and indexes used, pipelines are fast enough for production API endpoints, not just offline analytics.

TL;DR

Quick Example

Monthly revenue per product category for paid orders in 2026, top five categories:

Core Concepts

Essential Stages

Expressions

Inside stages, expression operators compute values: arithmetic ($add, $multiply), conditionals ($cond, $switch, $ifNull), strings ($concat, $toLower, $regexMatch), dates ($dateTrunc, $dateToString), arrays ($filter, $map, $reduce, $size), and type conversion ($toDecimal, $toObjectId). Field paths are referenced as "$field.subfield", and variables as "$var".

Joins With $lookup

Index the foreignField in the joined collection; otherwise every input document triggers a collection scan. Heavy $lookup use on hot paths is a sign the schema design may need more embedding.

Pagination With $facet

One round trip returns a page and the total count. For deep pagination, prefer range-based (keyset) queries over large $skip values.

Window Functions

Performance

See MongoDB indexes for designing indexes that support pipelines.

Best Practices

Build Pipelines Incrementally

Develop stage by stage, inspecting output after each one (MongoDB Compass's aggregation builder is excellent for this). Complex pipelines are much easier to debug when each stage's output is known.

Materialize Expensive Aggregations

For dashboards that recompute the same heavy aggregation on every request, run it on a schedule or on change and $merge the results into a summary collection. Reads then become simple indexed finds. This is the document-database version of a materialized view.

Offload Analytics From the Primary

Run heavy reporting pipelines against a secondary (readPreference: "secondaryPreferred") or a dedicated analytics node, so they don't compete with transactional traffic. For large-scale analytics, export to a warehouse; see data warehousing.

Common Mistakes

$match After $unwind or $group

Unindexed $lookup

Joining 100,000 orders to a customers collection without an index on the join field means 100,000 collection scans. Always index foreignField.

Huge $push Accumulators

$group with $push: "$ROOT" builds arrays of entire documents and can exceed the 16 MB document limit or the stage memory limit. Push only the fields you need, cap with $topN/$firstN, or restructure the query.

FAQ

Is the aggregation pipeline slower than find?

For the same filter and projection, no: a pipeline starting with $match and $project uses indexes just like find. Pipelines get expensive when they process many documents through grouping, unwinding, or joins, which find can't do at all.

What replaced map-reduce?

The aggregation pipeline. Map-reduce is deprecated; pipelines, including $function and $accumulator for custom JavaScript logic when truly needed, cover its use cases and run much faster.

Can I update documents with an aggregation?

Yes, in two ways: $merge writes pipeline output into a collection (insert, replace, or merge into matching documents), and update commands accept an aggregation pipeline as the update (updateMany({}, [ { $set: { total: { $sum: "$items.price" } } } ])) to compute new values from existing fields.

How do I do full-text or vector search in a pipeline?

On MongoDB Atlas, $search (Lucene-based full-text) and $vectorSearch stages run as the first stage of a pipeline, followed by any normal stages. Self-managed deployments have basic $text search, and recent versions add search and vector search through the community mongot component.

Related Topics

References