Context windows, tokens and what a model can remember

Language models are stateless. Nothing persists between turns; the entire conversation is re-sent on every message, and the context window is the hard limit on how much of it fits. Almost everything people describe as an AI forgetting is this limit being reached — or, more often, being approached, where the failure is quieter and much harder to spot.

Infrastructure · 7 min read

In short

A context window is the entire memory

When you send a message, the application concatenates the system prompt, any project or persona instructions, attached file text, every previous turn and the model's own previous replies, and submits the whole thing as one input. The model reads it, produces a response, and retains nothing. The next message repeats the process from scratch.

The context window is the maximum size of that input, measured in tokens rather than words or characters. English text runs roughly three-quarters of a word per token, so a thousand tokens is about 750 words — but the ratio varies with punctuation, code, formatting and language, and non-English text is frequently far less efficient. Treat the conversion as an estimate, never as arithmetic.

The window is also shared between input and output on most systems. Reserving room for a long response is part of the budget, which is why a nearly-full context produces truncated answers before it produces an error.

What is actually filling it

More than people expect. The hidden system prompt can be substantial. Custom instructions or project context are prepended to every message. An attached PDF becomes text and can consume tens of thousands of tokens on its own. Every previous turn counts — including the model's replies, which are usually longer than your messages. On reasoning models, the internal thinking trace consumes budget too, sometimes a great deal of it.

This is why a long chat gets progressively more expensive and slower for the same-sized question: you are re-sending a growing transcript each time. It is also why a conversation that started sharp can degrade over an afternoon without anything obviously breaking.

Why long context is expensive

Two costs, with different shapes. Self-attention compares every token to every other token, so compute grows with the square of sequence length — double the context and the attention work roughly quadruples. Various optimisations soften this substantially, but the underlying scaling is why long inputs are slow.

The second cost is memory. Serving keeps a key-value cache of the attention state for every token in the sequence, and that cache grows linearly with length while being large in absolute terms. On long conversations it often dominates GPU memory, and it is the reason a provider's concurrency drops sharply when everyone's contexts get long. When a host caps context, this is usually what they are rationing.

Note also that a config file accepting a large context does not mean the model was trained at that length. Techniques that scale positional encodings can extend a model beyond its training length, but quality past the trained range degrades — sometimes gracefully, sometimes not.

Advertised context versus effective context

A model that accepts 131,072 tokens does not reason equally well across all of them. The best-known finding here is that retrieval accuracy depends heavily on position: information at the very beginning and very end of a long input is recalled far better than information in the middle. It is a robust enough effect to have its own name — lost in the middle — and it means burying a crucial constraint halfway through a long document is a genuine mistake.

The other trap is what a passing test proves. Needle-in-a-haystack tests, where a single odd sentence is planted in a long document and the model asked to find it, are much easier than real work. Harder evaluations that require tracking several facts across a long input, or aggregating rather than locating, show performance falling well before the stated limit. A model can honestly advertise a large window and still be substantially worse at using the second half of it.

The practical reading: treat the advertised number as a ceiling on what fits, not a promise about what will be used well.

What happens when you overflow

Something has to be discarded, and the strategy is usually invisible. The most common approach is a rolling window: keep the system prompt, keep the most recent turns, drop the oldest. The conversation continues and the early context is simply gone. The second is rolling summarisation, where older turns are compressed into a synopsis that stays in context — better than deletion, and lossy in ways that favour plot over detail. The third is retrieval, where older material is stored externally and relevant fragments are pulled back in when they seem needed, which works well for documents and poorly for narrative continuity.

The symptoms are recognisable once you know them. A character's eye colour changes. A constraint you set at the start quietly stops being honoured. The model asks for information you already gave it, or contradicts a decision made an hour ago. None of this is the model being unreliable in a general sense — it is the model reasoning correctly over a context that no longer contains the thing you are annoyed it forgot.

Working with the window instead of against it

Front-load what must persist. Stable material — character sheets, style rules, project constraints — belongs in the system prompt or a project instruction block, because those survive turn eviction while ordinary messages do not.

Restate the constraints that matter, periodically and briefly. A one-line reminder near the end of a long conversation is worth more than the original instruction twenty thousand tokens back, both because of position effects and because it may no longer be in context at all.

Start a new conversation for a new task. Dragging an unrelated hour of history along costs money, costs latency and actively degrades attention on the thing you now care about. And match the model to the job: an 8K-context model is a fine fast conversationalist and the wrong tool for a manuscript, while a 131K model is the right tool for a document set and unnecessary for a quick question.

Terms used here

Token
The unit a model reads and writes — typically a word fragment. English averages roughly three-quarters of a word per token; code, punctuation and other languages differ substantially.
Context window
The maximum number of tokens the model can process in a single request, covering the system prompt, the conversation, attachments and usually the response.
KV cache
Stored attention keys and values for every token in the sequence, which avoids recomputing them. Grows linearly with context length and frequently dominates memory during serving.
Lost in the middle
The observed tendency for models to recall information placed at the start or end of a long input far more reliably than information in the middle.
Effective context
The length over which a model actually maintains reliable performance, as opposed to the maximum length it will accept without error. The first is always shorter.

Frequently asked questions

How many words is a 131K context window?

Roughly 95,000–100,000 words of ordinary English prose, which is a full novel — but that is an estimate, not a conversion factor. Code, tables, non-English text and heavy punctuation all tokenise less efficiently, sometimes dramatically so. And the window covers the system prompt, attachments and the response as well as your text.

Why does the AI forget what I said earlier?

Because it is no longer in the context. Either the conversation exceeded the window and the oldest turns were evicted, or the detail is still present but buried in the middle of a long input where retrieval is weakest. Restating the important constraint fixes both cases immediately.

Is a bigger context window always better?

No. Long context is slower, more expensive and subject to attention degradation, and a large window filled with irrelevant history produces worse answers than a small window filled with exactly the right material. Size raises the ceiling; curation determines the result.

Does the model remember me between conversations?

Not by itself. Any cross-conversation memory is a product feature: the application stores notes or summaries somewhere and re-injects them into the context of new chats. The model is stateless in every case — what varies is what the app chooses to put in front of it.

Do file uploads use context?

Yes, and usually a lot of it. Attached documents are extracted to text and inserted into the input, so a long PDF can consume most of a mid-sized window on its own. Systems that use retrieval instead insert only the relevant fragments, which is cheaper but means the model never sees the whole document at once.

On OpenRogue

OpenRogue's library spans a wide context range — 8K on the smallest fast models up to 131K on the Hermes family — and each model page states the figure, because the right window is task-dependent rather than universally larger-is-better. Switching models mid-conversation carries your existing context across, so a chat that outgrows a small model can move up without starting over.

Further reading

Related

Start free → · All models · Pricing