What is a context window, and why a bigger one is not always better

A context window is how much an AI model can hold in mind for a single request, and its size is now a headline selling point. Flagship models have reached about a million tokens. But research on 'lost in the middle' and 'context rot' shows a bigger window does not mean the model reliably uses all of it.

By Zain

Published

Visual abstraction of neural networks in AI technology, featuring data flow and algorithms
Photo: Google DeepMind / Pexels

What a context window actually is

A context window is the most text a model can hold in mind at once, for a single request. It is measured in tokens, the small chunks models read and write, and it counts everything in play: your prompt, the hidden system instructions, any documents you paste, the running chat history, and the model's own answer as it is generated. Think of it as the model's desk, not its long-term memory. Everything it can reason about for one reply has to fit on that desk at the same time. When the desk fills up, the oldest material slides off the edge and the model can no longer see it.

That is the whole concept, and it explains a lot of everyday AI behavior: why a long chat starts "forgetting" what you said at the top, why pasting a giant document sometimes works and sometimes does not, and why the same question can cost very different amounts depending on how much you send with it.

Windows got enormous, fast

For years the window was the binding constraint. Early chat models topped out around a few thousand tokens. By 2026 that ceiling has been blown away: most flagship models have converged on windows of about 1 million tokens, roughly 750,000 words, or several long novels' worth of text in a single request. A few go further, with outliers advertising several million tokens. Size has largely stopped being the thing that separates the top models from each other, and the real differentiators are now reasoning quality, speed and price.

That sounds like the problem is solved. Just paste everything in. It is not that simple.

“Performance is often highest when relevant information occurs at the beginning or end of the input context, and significantly degrades when models must access relevant information in the middle of long contexts.”

Liu et al., 'Lost in the Middle: How Language Models Use Long Contexts,' 2023
Detailed close-up image of an open book showing the textured pages and printed text
A million-token window can hold a 750,000-word library, but models attend to the beginning and end of long inputs far more reliably than the middle. Photo: Ruslan Alekso / Pexels

Bigger is not the same as reliable

The uncomfortable finding, replicated many times, is that a model having room for a million tokens does not mean it uses those tokens equally well. The clearest statement of it is a 2023 study by Liu and colleagues, memorably titled "Lost in the Middle." Their result: "performance is often highest when relevant information occurs at the beginning or end of the input context, and significantly degrades when models must access relevant information in the middle of long contexts." Plot the accuracy and it makes a U shape, strong at the two ends and sagging in the center, even in models built for long context.

Newer models have improved on this, but the effect has not gone away. Independent "context rot" testing across many current models found that reliability degrades as the input grows, on tasks as simple as finding a fact or copying text, and even in models carrying million-token windows. Worse, performance drops when you add irrelevant text, well before the window is anywhere near full. Filling the desk with clutter makes the model worse at finding the one paper that matters.

So why not just paste everything in

Two reasons, beyond the accuracy problem. First, cost: you pay per token for everything in the window, on every single call, so a bloated context is a bill you keep paying. Second, speed: more tokens to read means a slower response. Put the three together, higher cost, higher latency, lower reliability, and "stuff the whole knowledge base into the prompt" stops looking clever.

This is exactly why retrieval-augmented generation did not become obsolete the moment windows hit a million tokens, as many predicted. Retrieving the handful of relevant passages and sending only those is usually cheaper, faster and more accurate than sending everything and hoping the model finds the needle. A big window is a bigger haystack, not a better magnet.

How to actually use the window

Treat context as a scarce, curated resource, not a bucket. Put the most important information at the very start or the very end of what you send, where models attend best, and never bury the key instruction in the middle of a wall of text. Trim anything the model does not need for this specific task, because noise measurably lowers accuracy. And when the knowledge is large or changing, retrieve the relevant slice rather than pasting the whole thing. The window tells you what a model can physically ingest; it says nothing about what it will reliably reason over.

Our take

The march to million-token windows is real and genuinely useful, and it is also oversold. Capacity is not attention. A larger window lets you fit more in front of the model; it does not make the model read all of it with equal care, and the research on this is consistent enough that "we have a huge context window" should never be taken as "you can stop thinking about what you send." The teams getting the most out of long context are the ones still doing the unglamorous work of choosing what goes in and where. Bigger windows raised the ceiling. They did not repeal the need to be deliberate under it.

Frequently asked questions

What is a context window in AI?

It is the maximum amount of text, measured in tokens, that a model can consider at once for a single request. It includes your prompt, the system instructions, any documents you paste, the running chat history, and the model's own response. When the window fills, the oldest content falls out of view.

How big are context windows now?

In 2026 most flagship models have converged on windows of about 1 million tokens, roughly 750,000 words, and some advertise several million. Size has largely stopped being the differentiator between top models; reasoning quality, speed and price matter more.

Does a bigger context window mean better answers?

Not necessarily. The 'lost in the middle' effect, documented by Liu and colleagues in 2023, shows models use information at the start and end of a long context far more reliably than the middle. Independent 'context rot' testing found accuracy degrades as inputs grow, even for million-token models.

Should I just paste everything into the context instead of using RAG?

Usually no. You pay per token for everything in the window on every call, so a stuffed context is slower and costlier, and the reliability drop means it can also be less accurate. Retrieving only the relevant chunks, as RAG does, is often better even when a huge window is available.

How do I get the most out of a context window?

Put the most important information at the very beginning or end, not buried in the middle. Trim irrelevant text, since adding noise measurably hurts performance. And treat the window as a scarce, curated resource rather than a bucket to fill.

Sources

What each one is, and whose it is.

  1. PaperIndependent of the vendor
  2. DocumentationIndependent of the vendor
  3. DocumentationIndependent of the vendor
  4. 4

    LLM Context Window Growth Timeline, hidekazu-konishi.com

    DocumentationIndependent of the vendor