The Biggest Window I Can Run — And Quality Slips Anyway

I run models with a million-token context window, and I never compact manually. No /compact in between, no deliberate tidying of the session. The window is so large that it feels as if there were no ceiling — I can keep a session running for hours, pull files in, collect tool outputs, hold discussions, and the window fills up without the session ever cutting out.

And yet something happens that irritated me at first: deep inside a long session, quality slips away. Not with a bang, not with an error message. The answers get subtly worse — the model loses a thread I had strung up four hundred lines earlier, or it contradicts a decision that had long since been made. The unsettling part isn't the quality drop itself. The unsettling part is that the model keeps answering with full confidence, as if nothing were wrong. No hesitation, no "I'm not sure here," no hint that the load has grown too large.

The obvious reflex would be: "Then just use a bigger window." But that very reflex falls flat. The setup I'm describing already has the biggest window I can buy. There is no "more" left. The degradation doesn't happen because the window is too small — it happens inside a nominally enormous window.

That led me to a different question than "how much context fits in?" At the end of this article stands an engineering question that shifts the pipeline design — and the answer to it is not "more window." But before I get there, it's worth looking at what the research knows about this phenomenon. Because it's measurable.

Mental Load Is Measurable — Three Findings From the Research

The phenomenon I experience is not a gut feeling. It's documented in several independent works, and it's worth laying three of them side by side.

Finding 1: The middle falls through. Liu et al. show in Lost in the Middle: How Language Models Use Long Contexts (TACL 2023) that a model's performance depends heavily on where in the context the relevant information sits. The curve is U-shaped: information at the beginning or the end of the input context is used reliably, information in the middle drops off. In their words: "performance is often highest when relevant information occurs at the beginning or end of the input context, and significantly degrades when models must access relevant information in the middle of long contexts." This middle effect worsens as the context length grows.

Finding 2: The effective window is drastically smaller than the advertised one. The work Context Is What You Need: The Maximum Effective Context Window for Real World Limits of LLMs (arXiv 2509.21361) introduces a distinction that hits my setup head-on: between the Maximum Context Window (MCW) — the advertised number — and the Maximum Effective Context Window (MECW) — what the model actually processes reliably. The gap is not small: "All models fell far short of their Maximum Context Window by as much as >99%." More than 99 percent. And the early failure sets in disturbingly soon in places: "A few top of the line models in our test group failed with as little as 100 tokens in context; most had severe degradation in accuracy by 1000 tokens in context." That's the hard evidence underneath my hook observation: my 1M window is a marketing number, not a usable workspace.

Finding 3: Content corrupts silently. The most uncomfortable work is probably LLMs Corrupt Your Documents When You Delegate (Microsoft Research, arXiv 2604.15597), which uses the DELEGATE-52 benchmark across 52 professional domains and 19 models to measure what happens to document content in long delegation workflows. The result: "even frontier models (Gemini 3.1 Pro, Claude 4.6 Opus, GPT 5.4) corrupt an average of 25% of document content by the end of long workflows, with other models failing more severely." A quarter of the content. And the mechanism is the treacherous part: "they introduce sparse but severe errors that silently corrupt documents, compounding over long interaction." Silent and compounding — the errors stack up, and nobody raises the alarm. The title page puts a concrete number on it: "current frontier models degrade 25% of document content after just 20 interactions." Agentic tool use does not improve this.

Three findings, three methods, one common denominator. Chroma Research coined a term for it in 2025 that has stuck ever since: Context Rot. In their report Context Rot: How Increasing Input Tokens Impacts LLM Performance (Hong, Troynikov, Huber, July 2025) they sum up the core insight like this: "models do not use their context uniformly; instead, their performance grows increasingly unreliable as input length grows." Context is not used uniformly; reliability falls as the input length grows. That's the bracket around the three findings — and by now an established term that Anthropic has adopted in its engineering material as well.

Bar comparison of two context windows: a tall bar for the advertised Maximum Context Window (MCW) against a tiny bar for the effectively usable Maximum Effective Context Window (MECW); the gap between them is annotated as ">99%". Advertised maximum context window versus the effectively usable window — the gap reaches more than 99 percent.

The Model Can't Tell — And That's Exactly the Problem

I call this load Mental Load, and I'm about to talk about Self-Care — so let me frame this cleanly before the term has to carry weight. "Self-Care" here is an operational metaphor, an engineering frame for externally built-in self-regulation. It's not about care theory, and it's certainly not about the hype that models would "feel" anything or "suffer" under load. An LLM does not experience mental load. The metaphor describes nothing but a system behavior and the question of who regulates it.

With that clarification behind me, the actual provocation: unlike a human, an LLM possesses no self-initiated self-care mechanisms whatsoever. A human working for hours on an overloaded problem will at some point recognize the overload, signal it ("I need a break," "let's split this up"), compensate for it (notes, structure), or delegate on their own initiative. The model does none of this. It does not recognize the growing load. It does not signal it. It does not compensate for it. And it does not delegate on its own.

This is exactly what closes the loop back to my opening observation. The confidence stays unchanged while quality falls — the model has no internal sensor reporting back "I'm becoming unreliable right now." The research sharpens this further: the degradation runs differently across model families. Some models (the Claude line, for instance) tend under uncertainty to refuse an answer, others (GPT models, for instance) show higher hallucination rates in the presence of distractor files. But even the "better" behavior is not a self-initiated self-care signal — it's just a different failure pattern. None of the models actively reports: "I just silently broke half the document."

This is the point where the diagnosis gets uncomfortable. The measurable load from chapter two meets a system that neither perceives this load nor counters it on its own. The gap doesn't lie in the load itself — load is to be expected. The gap lies in the fact that nobody in the system notices it and nobody in the system closes it.

External Self-Care Is a Structural Obligation, Not a Comfort Feature

If the model doesn't perform the regulation itself, then it has to come from outside. This is not a bonus and not a nice optimization for advanced users — it's the direct logical consequence of the gap the previous section opened up. The self-regulation a human brings along has to be built externally into the pipeline for an LLM. Structurally, not optionally.

And this is not my private opinion that I'm imposing on the field. The field has long treated context exactly this way. Anthropic states the framing explicitly in Effective context engineering for AI agents: "Context is a critical but finite resource for AI agents." Context is a critical but finite resource — with diminishing marginal returns the fuller the window gets. And the goal of the discipline is named just as clearly: "finding the smallest possible set of high-signal tokens that maximize the likelihood of some desired outcome." The smallest possible set of high-signal tokens. If context is a finite resource with diminishing marginal returns, then handling it is an engineering task. And external self-care is nothing other than the management of this resource.

What does this external self-care look like concretely? Four patterns have emerged in the field, each serving a different axis of regulation: reduce the load, distribute the load, cap the load, and offload the load. I present them as a convergence that can be observed — not as "the only right four" and not as a complete best-practice catalog. Four axes that together cover what a human would do intuitively and a model simply does not.

A map with four axes of external self-care, each labeled with its corresponding pattern: reduce the load (Compaction), distribute the load (Sub-Agent Isolation), cap the load (Context Budgets), and offload the load (Externalized Memory). Four axes of external self-care: reduce, distribute, cap, and offload the load.

Pattern 1 — Compaction: Actively Reducing the Load

The first axis is the most direct: when the context fills up, reduce the load by summarizing the conversation and making room. That's Compaction.

How this works mechanically is well documented. The Claude Code docs say of the manual step: "/compact summarizes the conversation to free space while keeping key information." The conversation is condensed, key information is preserved, ballast falls away. Crucial for my setup: this also happens automatically. When the window fills up, that doesn't end the session — Claude Code summarizes on its own, and "the automatic pass works the same way as the /compact step in the timeline." So the automatic pass works just like the manual step, only without me having to intervene.

This is exactly where I have to be honest: I don't run a sophisticated compaction policy of my own. I don't compact manually and have no homegrown heuristic for when to condense — I rely on auto-compaction and the large window. In that sense, Compaction is the pattern that, in my daily operation, most happens for me rather than through me. That doesn't make it any less relevant, on the contrary: it's the external self-care that the tooling already takes over — the first load reduction that kicks in before I even think about it. But it's also the one where it's easiest to forget it's happening at all.

Pattern 2 — Sub-Agent Isolation: Distributing the Load

The second axis distributes the load instead of reducing it. The idea: not everything has to go through the same context window. A subtask that produces a lot of verbose output — a sprawling research run, a long test run — would flood the main conversation and bloat exactly the middle that, according to Lost in the Middle, falls through anyway.

Sub-agents solve this through isolation. According to the Claude Code docs: "Each subagent runs in its own context window with a custom system prompt, specific tool access, and independent permissions." Each sub-agent has its own window. The verbose output stays where it arises — "the subagent does that work in its own context and returns only the summary." Only the relevant summary returns to the main conversation. The load is distributed across several isolated windows instead of stacking up in a single one. The effect is exactly the one a human achieves with delegation: the detail noise of a subtask need not burden the main awareness, as long as the result comes back cleanly.

I don't run a sub-agent setup of my own — that's not part of my daily practice. But it's the pattern with, in my opinion, the greatest architectural leverage, and anyone who wants to dig deeper will find more from me on it in a separate piece on sub-agent architectures and multi-agent memory.

Pattern 3 — Context Budgets: Capping the Load

The third axis caps the load before it arises: treat the token volume as a deliberate engineering variable and not as something that simply emerges. This is the direct practical consequence of the Anthropic principle of the smallest possible set of high-signal tokens. Instead of "in it goes, the window is large," the stance is: every token that goes in should justify itself.

That the configuration of the context is to be taken seriously can also be shown from the model side. Anthropic shows in Quantifying infrastructure noise in agentic coding evals that even the infrastructure configuration can shift benchmarks noticeably: "Infrastructure configuration can swing agentic coding benchmarks by several percentage points—sometimes more than the leaderboard gap between top models." That's several percentage points, in places more than the gap between the top models on the leaderboard. If even configuration noise moves that much, then what you deliberately load into the window is all the more a variable you should steer rather than leave to chance.

Here too: I maintain no formal token-budget discipline of my own. I write Pattern 3 from the field, not from a personal routine — but it's the axis that most calls for discipline rather than tooling.

Pattern 4 — Externalized Memory: Offloading the Load

The fourth axis is the one I actually use every day — and the only pattern where I speak from my own practice. The idea: what's important doesn't have to survive in the context window. It can be offloaded, into a persistent layer outside the window.

In my daily-driver operation, this looks concretely like this: what a single session doesn't have to survive but is needed across sessions migrates into a persistent memory — conventions and decisions made, reference pointers to tools and data sources, the state of ongoing work along with open questions, and distilled lessons from earlier sessions. The crucial part is the discipline behind it: the memory is not a dumping ground but curated — one fact per entry, cross-checked against reality rather than blindly trusted, and checked for duplicates before every addition. This is the simplest and, for me, most reliable form of external self-care: I don't rely on the model still finding the important decision reliably four hundred lines later in the dead middle — I take it out of the window.

With that, Pattern 4 solves exactly the gap that chapters two and three diagnosed. If the middle falls through and content corrupts silently, then the safest thing is not to entrust the essential to the window in the first place. Memory as a persistent layer is the insurance against Context Rot. How memory architectures implement this in detail, I've described more thoroughly in my article series on memory architectures.

The New Guiding Question in Pipeline Design

Back to the beginning. I run the biggest window I can get, and I never compact manually — and yet, deep in the session, quality slips away while the model keeps answering with full confidence. The four patterns are the answer to exactly this experience. They build, from the outside, the self-regulation that the model does not have on its own in the 1M window: Compaction reduces the load, Sub-Agent Isolation distributes it, Context Budgets cap it, externalized memory offloads it.

This shifts the question you ask yourself when designing an LLM pipeline. The obvious question is "How much context fits in?" — and that, as the hook showed, is the wrong one. More window doesn't solve the problem, because the effective window only amounts to a fraction of the advertised one anyway and confidence doesn't fall with quality. Whoever optimizes for window size optimizes for the number on the data sheet, not the one the model actually works with. The right question is a different one:

Where in my pipeline is the external self-care?

Where is the load reduced, distributed, capped, offloaded — and that from the outside, because the model does not do it itself? For me, this question belongs in every pipeline design that goes beyond a single short prompt. For practitioners it's a craft check: which of the four patterns kicks in where? For leads it's an architecture question: is the self-regulation even structurally provided for anywhere, or are we implicitly relying on a large window being enough?

Mental Load in LLMs is measurable, and the model can't tell. That's exactly why we have to tell — and build it into the pipeline.

This article was originally published on Medium.