GraphQL DataLoader & the N+1 Problem
GraphQL resolves queries field by field. That keeps resolvers simple, since each one only knows how to fetch its own field, but it creates a classic performance trap. Query 50 orders and their customers, and a naive server runs 1 query for the orders plus 50 queries for the customers: the N+1 problem. Nest another level (each customer's address) and it multiplies again.
DataLoader is the standard fix. It collects all the individual load(id) calls made during one tick of execution, sends them to your data source as one batched request, and caches results for the rest of the request. The resolvers stay simple, and the database sees a handful of queries instead of hundreds.
TL;DR
- N+1: one query for a list, then one query per item for a related field. It's the most common GraphQL performance bug.
- DataLoader batches
load(key)calls made in the same tick into a singlebatchFn(keys)call. - The batch function must return results in the same order and length as the keys, with
nullor anErrorfor misses. - DataLoader also caches per request, so the same key loaded twice hits the source once.
- Create new loaders per request (in context), never globally, to avoid leaking data between users.
- Alternatives: lookahead / query planning (join in the parent resolver), ORM prefetching, or GraphQL-to-SQL compilers.
Quick Example
Node.js with the dataloader package:
That's 51 queries reduced to 2, no matter how many orders are returned.
Core Concepts
Why N+1 Happens
The GraphQL executor calls Order.customer once per order object. Each call independently fetches its customer, and since resolvers don't know about each other, nothing combines those fetches. The same pattern appears with REST calls to downstream services, which is often even more expensive than database queries.
How DataLoader Batches
- During execution, resolvers call
loader.load(key), and each call returns a promise. - DataLoader queues keys instead of fetching immediately.
- At the end of the current event-loop tick (after all sibling resolvers have run), it calls your batch function once with all queued keys.
- It resolves each promise with the result at the matching position.
Because GraphQL executes sibling fields in the same tick, all 50 Order.customer resolvers queue their keys before the batch fires.
The Batch Function Contract
- Input: an array of keys (deduplicated by the cache).
- Output: a promise of an array with exactly the same length and order as the keys.
- For missing records return
null; for per-key failures return anErrorinstance (that key'sloadrejects without failing the others).
Getting the ordering wrong is the classic DataLoader bug: databases don't return rows in IN (...) order, so always map results back by key.
Per-Request Caching
DataLoader memoizes load(key) promises for its lifetime. If the same customer appears on ten orders, it's fetched once. That's why loaders must be created per request: a global loader would cache data across users, a security and staleness problem. Use loader.clear(key) or prime(key, value) after mutations to keep the request's cache consistent.
One-to-Many Loaders
For a field like Customer.orders, the batch function takes customer IDs and returns an array of orders per customer:
Pagination per parent (the first 5 orders for each of 50 customers) is harder to batch. Use window functions (ROW_NUMBER() OVER (PARTITION BY customer_id ...)) or a lateral join in the batch query.
DataLoader in Other Ecosystems
Alternatives and Complements
- Lookahead / query planning: inspect the requested selection set (
infoin resolvers) and eagerly join or prefetch in the parent resolver. It's efficient but couples resolvers to query shapes. - ORM prefetching: Django's
select_related/prefetch_related, SQLAlchemy'sselectinload, and Prisma'sinclude, driven by the selection set. See Django ORM. - GraphQL-to-SQL compilers (Hasura, PostGraphile, Join Monster) translate an entire query into one SQL statement.
- Caching layers (Redis) in front of batch functions for hot, rarely-changing entities. See Redis caching patterns.
Best Practices
Use Loaders for Every Relationship Field
Make it a convention: any resolver fetching by ID goes through a loader. Even fields that seem to be resolved only once today will end up inside a list tomorrow.
Batch Calls to Downstream Services Too
Loaders work for any data source. Batch REST or gRPC calls using bulk endpoints (GET /customers?ids=1,2,3). If a service lacks a batch endpoint, adding one is often the highest-value performance fix.
Enforce Authorization Before or Inside Loading
A shared per-request loader may serve cached objects to different fields. Apply authorization in resolvers or pass the viewer into batch functions, so batching never returns data the user can't see. See GraphQL security.
Watch Query Counts in Tests
Add assertions or tracing that count database queries per GraphQL operation (for example "this query must use at most 5 SQL statements"), so N+1 regressions fail CI instead of reaching production.
Common Mistakes
A Global Loader
Returning Results in Database Order
Awaiting Loads Sequentially in One Resolver
for (const id of ids) await loader.load(id) dispatches one batch per iteration. Use loader.loadMany(ids) or Promise.all(ids.map(id => loader.load(id))) so all keys join one batch.
FAQ
What exactly is the N+1 problem in GraphQL?
Resolving a list of N items and then a related field on each item triggers one query for the list plus one per item: N+1 queries total. It grows with each nested level and page size, turning a single GraphQL request into hundreds of database round trips.
Does DataLoader replace a cache like Redis?
No. DataLoader's cache lives only for one request and exists to deduplicate and batch loads within it. A shared cache like Redis stores data across requests. They complement each other: the batch function can check Redis before hitting the database.
Do I need DataLoader with an ORM?
Usually yes, or an equivalent. ORMs don't know about GraphQL's resolver-by-resolver execution, so lazy-loaded relationships still produce N+1 queries. Either use DataLoader around ORM queries or drive the ORM's eager loading from the GraphQL selection set.
Does DataLoader work across nested levels?
Yes. Each level batches independently: all orders' customers in one batch, then all those customers' addresses in the next. The number of queries grows with the query's depth, not the number of objects.
Related Topics
- GraphQL — The query language overview
- GraphQL Schema Design — Relationship fields that need loaders
- GraphQL Pagination — Batching paginated relationships
- Query Optimization — Making the batched queries fast
- Django ORM —
select_relatedandprefetch_related - Object-Relational Mapping — Lazy loading and N+1 in ORMs