Deploying FastAPI in Production
FastAPI applications run on an ASGI server, most commonly Uvicorn. Getting from fastapi dev on a laptop to a reliable production service involves a few decisions: how many worker processes, and whether to scale with processes or with container replicas; how to build a lean Docker image; how to sit behind a reverse proxy or load balancer correctly; and how to handle startup, shutdown, health checks, configuration, and observability.
None of it is complicated, but mistakes like blocking calls in async endpoints, a single worker on a multi-core machine, or ignored proxy headers cause most "FastAPI is slow" or "redirects break in production" reports.
TL;DR
- Run with
fastapi runoruvicorn app.main:app --host 0.0.0.0 --port 8000. Never use--reloadin production. - On VMs, use multiple worker processes (
--workers N, roughly one per CPU core). On Kubernetes, prefer one process per container and scale with replicas. - Don't block the event loop: use
async defonly with async libraries; use plaindeffor blocking code (it runs in a thread pool). - Build small, non-root Docker images (slim base, uv or pip with a lockfile, multi-stage).
- Behind a proxy, enable proxy headers (
--proxy-headers,--forwarded-allow-ips) and setroot_pathif served under a prefix. - Use lifespan for startup and shutdown, add health endpoints, load config with pydantic-settings, and export metrics, traces, and structured logs.
Quick Example
A production Dockerfile using uv:
Application lifecycle and health:
Core Concepts
ASGI Servers and Workers
- Uvicorn is the standard ASGI server (
uvicorn[standard]adds uvloop and httptools for speed).fastapi runwraps it with production defaults. - Workers: each worker is a separate Python process with its own event loop, and Python's GIL means one process uses about one CPU core for Python code.
uvicorn --workers 4runs four processes. Gunicorn with Uvicorn workers is a traditional alternative on VMs. - Containers and Kubernetes: run a single Uvicorn process per container and let the orchestrator scale replicas and restart failures. It gives simpler resource accounting, per-replica health checks, and autoscaling.
Hypercorn and Granian are alternative ASGI servers with HTTP/2 or HTTP/3 support and different performance profiles.
Async vs Sync Endpoints
async defendpoints run on the event loop. Any blocking call (sync DB driver,requests,time.sleep, heavy CPU work) stalls every request on that worker.defendpoints run in a thread pool (40 threads by default via AnyIO), which is safe for blocking libraries, but bounded.- CPU-heavy work (image processing, ML inference, report generation) belongs in background workers or separate services, not in request handlers.
See Python asyncio.
Reverse Proxies and Load Balancers
FastAPI typically sits behind Nginx, a cloud load balancer, or a Kubernetes ingress:
- TLS termination at the proxy.
- Proxy headers: enable
--proxy-headersand set--forwarded-allow-ipsto the proxy's addresses, sorequest.client.host, scheme, and generated URLs (redirects,url_for) reflect the original request. That fixes the classic "redirect to http:// instead of https://" bug. root_path: when served under a path prefix (/api), configureroot_pathso docs and URLs work.- Timeouts and body sizes: align proxy timeouts with slow endpoints, and set request size limits at the proxy.
See Nginx reverse proxy.
Lifespan and Graceful Shutdown
Use the lifespan context manager to create shared resources (DB engines, HTTP clients, ML models) at startup and close them at shutdown. On SIGTERM, Uvicorn stops accepting connections and waits for in-flight requests (--timeout-graceful-shutdown). Match it to the Kubernetes terminationGracePeriodSeconds, and add readiness gating so load balancers stop routing first. See Linux processes & signals.
Configuration
Load settings from environment variables with pydantic-settings, validated at startup (see FastAPI & Pydantic). Keep secrets in the platform's secret store, disable or protect /docs in production, and set debug=False.
Observability
- Structured logs (JSON) with request IDs, via a logging middleware or
structlog, to stdout. - Metrics: request rate, latency histograms, and error counts via
prometheus-fastapi-instrumentatoror OpenTelemetry metrics, scraped by Prometheus. - Traces: OpenTelemetry's FastAPI, SQLAlchemy, and httpx instrumentation, exported via OTLP. See OpenTelemetry instrumentation.
- Error tracking: Sentry or similar for exceptions with context.
Best Practices
Scale Horizontally With Stateless Instances
Keep no per-user state in process memory (sessions, caches that must be consistent). Use Redis or the database, so any replica can serve any request and scaling is just adding replicas.
Separate Liveness and Readiness
Liveness should only verify the process responds. Readiness can check critical dependencies. Never restart pods because the database blipped. See Kubernetes workloads.
Size Connection Pools Across Workers
Each worker process has its own database pool. Ten replicas × four workers × pool size 10 means 400 connections. Budget pools against database limits, or use a pooler like PgBouncer. See PostgreSQL connection pooling.
Load Test Before Launch
Measure throughput and p99 latency with realistic traffic (k6, Locust), check for event-loop blocking, and tune workers, pools, and timeouts based on data. See performance testing.
Common Mistakes
Blocking Libraries Inside async def
Use httpx.AsyncClient, or declare the endpoint with plain def.
Running --reload or Debug Mode in Production
--reload watches files and restarts processes, which wastes CPU and is unsafe in production. It's for development only.
Ignoring Proxy Headers
Without proxy header handling behind a TLS-terminating proxy, FastAPI believes requests are plain HTTP from the proxy's IP. Redirects downgrade to http://, logs show the proxy IP, and IP-based rate limiting breaks.
FAQ
How many Uvicorn workers should I run?
On a VM or bare metal, start with roughly one worker per CPU core for async-heavy apps, and measure. In Kubernetes, run one worker per container and scale replicas instead; set CPU requests so each replica gets about one core.
Should I use Gunicorn with Uvicorn workers?
It used to be the standard recommendation for process management on VMs. Uvicorn now supports --workers with its own supervisor, and container orchestrators handle restarts, so plain Uvicorn (or fastapi run) is usually sufficient. Gunicorn remains a fine choice if you rely on its features.
Is FastAPI fast enough for production?
Yes. It's one of the fastest Python web frameworks, and bottlenecks are almost always I/O (database queries, external APIs) or blocking code in async endpoints. Efficient queries, connection pooling, caching, and correct async usage matter far more than framework overhead.
Can I serve FastAPI serverlessly?
Yes, with adapters like Mangum (AWS Lambda) or on platforms that run containers on demand (Cloud Run, Azure Container Apps). Watch cold starts, connection pooling to databases (use proxies like RDS Proxy), and request timeouts.
Related Topics
- FastAPI — The framework overview
- Docker Multi-Stage Builds — Lean Python images
- Nginx Reverse Proxy — Fronting ASGI apps
- Kubernetes Workloads — Probes and graceful shutdown
- Python asyncio — Avoiding event-loop blocking
- Python Packaging — uv and lockfiles for deployments