LLM Context Windows

A model's context window is the maximum number of tokens it can consider at once: system prompt, conversation history, documents, tool definitions and results, and the response it generates. Anything outside the window doesn't exist for the model. It has no memory beyond what you send in the current request.

Windows have grown from about 4,000 tokens in early chat models to hundreds of thousands and even millions today. That changes what's possible (whole codebases, long contracts, hours of transcripts in one prompt), but a bigger window isn't free and isn't automatically better. Cost, latency, and model attention all degrade as you fill it. Deciding what goes into the window is the core of context engineering.

TL;DR

Quick Example

Keeping a chat within budget by caching a large stable prefix and trimming old turns:

In production, summarizing old turns usually beats simply dropping them. See compaction below.

Core Concepts

What Counts Against the Window

Everything the model processes in a request:

If input plus requested max_tokens exceeds the window, the API rejects the request, or the output gets truncated.

Statelessness and "Memory"

The model retains nothing between API calls. Chat apps create the illusion of memory by re-sending the conversation each time. Long-term memory features (user preferences, past sessions) work by storing information externally and injecting the relevant parts into the context. See AI agents.

Cost and Latency

How Well Models Use Long Context

Benchmarks like "needle in a haystack" show modern models can find a single fact anywhere in huge contexts. Real tasks are harder: synthesizing information spread across a long document, following instructions buried deep in context, or reasoning over many similar passages. Quality often degrades as contexts grow, a phenomenon sometimes called context rot, and information in the middle can be underused ("lost in the middle"). Mitigations:

Managing Conversation History

Agentic systems accumulate tool results fast. Aggressively trimming stale tool outputs and compacting history keeps long-running agents inside their window and focused.

Long Context vs Retrieval

Many systems combine them: retrieve generously (dozens of chunks or whole documents) into a large window, instead of squeezing into a tiny one. See agentic RAG for retrieval driven by the model itself.

Best Practices

Budget the Window Explicitly

Allocate tokens per component (instructions, tools, retrieved context, history, output) and enforce the budgets with the target model's token counter. Surprises in production usually come from one component, like a giant tool result, crowding out the rest.

Order for Caching and Attention

Stable content first (system prompt, tool definitions, reference documents) with cache breakpoints, then dynamic content (retrieved chunks, history), then the user's question last.

Prefer Relevant Over Everything

Filter, deduplicate, and rank before inserting. Sending 20 highly relevant chunks usually beats 200 loosely related ones, in quality and in cost.

Test With Realistic Long Inputs

Evaluate your use case at the context sizes you'll actually send, with the answer placed in different positions. Headline window sizes don't tell you how well a model reasons at that length on your task. See LLM evaluation.

Common Mistakes

Unbounded Chat History

Enforce a budget and compact or trim older turns.

Forgetting Output and Reasoning Tokens

A prompt that uses 195k of a 200k window leaves only 5k tokens for the response, and reasoning models may spend many tokens thinking first. Reserve output headroom.

Dumping Raw Tool Output Into Context

A tool returning a 50,000-token JSON blob when the agent needed three fields wastes budget and distracts the model. Make tools return concise, relevant results, with pagination or filters.

FAQ

What happens when I exceed the context window?

Most APIs return an error when input plus max_tokens exceeds the limit. Some chat products silently drop or summarize older messages. If the output hits max_tokens, the response is truncated, which the stop reason shows.

Is a bigger context window always better?

No. It enables tasks that need lots of material at once, but each request costs more, runs slower, and can be less accurate if filled with marginally relevant text. The best systems put the right tokens in the window, not the most.

Does the model remember previous conversations?

Not by itself. Each request is independent. Continuity comes from your application re-sending history or injecting stored memories. Provider features like "projects" or "memory" are implemented the same way, outside the model.

How does prompt caching work with the context window?

Caching doesn't enlarge the window. Cached tokens still count toward it. It makes processing a repeated prefix cheaper and faster by reusing computed state for a time window (minutes by default, longer with extended options).

Related Topics

References