Deploying FastAPI in Production

FastAPI applications run on an ASGI server, most commonly Uvicorn. Getting from fastapi dev on a laptop to a reliable production service involves a few decisions: how many worker processes, and whether to scale with processes or with container replicas; how to build a lean Docker image; how to sit behind a reverse proxy or load balancer correctly; and how to handle startup, shutdown, health checks, configuration, and observability.

None of it is complicated, but mistakes like blocking calls in async endpoints, a single worker on a multi-core machine, or ignored proxy headers cause most "FastAPI is slow" or "redirects break in production" reports.

TL;DR

Quick Example

A production Dockerfile using uv:

Application lifecycle and health:

Core Concepts

ASGI Servers and Workers

Hypercorn and Granian are alternative ASGI servers with HTTP/2 or HTTP/3 support and different performance profiles.

Async vs Sync Endpoints

See Python asyncio.

Reverse Proxies and Load Balancers

FastAPI typically sits behind Nginx, a cloud load balancer, or a Kubernetes ingress:

See Nginx reverse proxy.

Lifespan and Graceful Shutdown

Use the lifespan context manager to create shared resources (DB engines, HTTP clients, ML models) at startup and close them at shutdown. On SIGTERM, Uvicorn stops accepting connections and waits for in-flight requests (--timeout-graceful-shutdown). Match it to the Kubernetes terminationGracePeriodSeconds, and add readiness gating so load balancers stop routing first. See Linux processes & signals.

Configuration

Load settings from environment variables with pydantic-settings, validated at startup (see FastAPI & Pydantic). Keep secrets in the platform's secret store, disable or protect /docs in production, and set debug=False.

Observability

Best Practices

Scale Horizontally With Stateless Instances

Keep no per-user state in process memory (sessions, caches that must be consistent). Use Redis or the database, so any replica can serve any request and scaling is just adding replicas.

Separate Liveness and Readiness

Liveness should only verify the process responds. Readiness can check critical dependencies. Never restart pods because the database blipped. See Kubernetes workloads.

Size Connection Pools Across Workers

Each worker process has its own database pool. Ten replicas × four workers × pool size 10 means 400 connections. Budget pools against database limits, or use a pooler like PgBouncer. See PostgreSQL connection pooling.

Load Test Before Launch

Measure throughput and p99 latency with realistic traffic (k6, Locust), check for event-loop blocking, and tune workers, pools, and timeouts based on data. See performance testing.

Common Mistakes

Blocking Libraries Inside async def

Use httpx.AsyncClient, or declare the endpoint with plain def.

Running --reload or Debug Mode in Production

--reload watches files and restarts processes, which wastes CPU and is unsafe in production. It's for development only.

Ignoring Proxy Headers

Without proxy header handling behind a TLS-terminating proxy, FastAPI believes requests are plain HTTP from the proxy's IP. Redirects downgrade to http://, logs show the proxy IP, and IP-based rate limiting breaks.

FAQ

How many Uvicorn workers should I run?

On a VM or bare metal, start with roughly one worker per CPU core for async-heavy apps, and measure. In Kubernetes, run one worker per container and scale replicas instead; set CPU requests so each replica gets about one core.

Should I use Gunicorn with Uvicorn workers?

It used to be the standard recommendation for process management on VMs. Uvicorn now supports --workers with its own supervisor, and container orchestrators handle restarts, so plain Uvicorn (or fastapi run) is usually sufficient. Gunicorn remains a fine choice if you rely on its features.

Is FastAPI fast enough for production?

Yes. It's one of the fastest Python web frameworks, and bottlenecks are almost always I/O (database queries, external APIs) or blocking code in async endpoints. Efficient queries, connection pooling, caching, and correct async usage matter far more than framework overhead.

Can I serve FastAPI serverlessly?

Yes, with adapters like Mangum (AWS Lambda) or on platforms that run containers on demand (Cloud Run, Azure Container Apps). Watch cold starts, connection pooling to databases (use proxies like RDS Proxy), and request timeouts.

Related Topics

References