LLM context window explained — tokens, not pages
A context window is the working memory of a request: system prompt, history, retrieved chunks, user message, and (usually) the room left for the answer. It is measured in tokens, not pages, and it is billed when you fill it.
English is roughly 0.75 words per token (or ~1.3 tokens per word). A 128k window is on the order of a short book — if you actually paste one. Most production prompts are far smaller, which is why paying for a 1M window you never fill is a hobby.
Longest context windows right now
Live OpenRouter pricing and Arena quality, cached about an hour. Not a static blog table.
| Model | Provider | Blend $ / 1M | Quality | Context |
|---|---|---|---|---|
| Grok 4.20 Multi-Agent | xAI | 1.88 | — | 2,000,000 |
| Grok 4.20 | xAI | 1.88 | — | 2,000,000 |
| Auto Router (Beta) | OpenRouter | — | — | 2,000,000 |
| Pareto Code Router | OpenRouter | — | — | 2,000,000 |
| Auto Router | OpenRouter | — | — | 2,000,000 |
Input and output share the window
If max context is 128k and you paste 120k, you may only have 8k left to generate (or the API will error). Check max output separately on the model page. “128k context” does not mean “128k novel as the reply.”
Lost in the middle
Attention is uneven. Facts buried in the center of a huge prompt are missed. Prefer retrieval of the right 4k tokens over dumping 100k tokens of maybe-relevant PDFs.
Chat history
Multi-turn products slowly eat the window. Summarize or drop old turns. Otherwise you will hit the cap and also pay to re-send the novel every message.
How we display context
We take OpenRouter’s context_length / top_provider context when present. If both are missing, the row may be dropped from the leaderboard. Treat the number as the provider’s claimed cap.
Choosing a size
Measure p95 prompt tokens in production. Pick a window with ~2× headroom. If p95 is 6k, you do not need 1M. If p95 is 80k because you paste dumps, fix retrieval before you buy context.
FAQ
How many tokens is 1,000 words?
- Roughly 1,300 tokens for English. Other languages differ.
Is a bigger window always better?
- No. It is slower and more expensive if filled, and quality at the cap can drop.
What’s the difference vs RAG?
- RAG chooses which tokens enter the window. A big window without RAG is just a bigger dump.
Where do I see each model’s window?
- Leaderboard context column, model detail pages, and the live table on the longest-context guide.
Related guides
Also try: LLM leaderboard, compare tool, live arena.