Gelişimİş
0

Context Windows and Cognitive Load: How to Design Prompts That Fit How You Actually Think

TL;DR. The context window is the part of an AI model that holds what you give it in a single request, and in 2026 it is huge. Anthropic, OpenAI and Google all list flagship models with about one million tokens of context, which by Google’s own conversion rate (100 tokens is about 60-80 English words) is roughly 600,000 to 800,000 words. The temptation is to paste everything in and let the model sort it out. The research says that is a mistake. In the RULER benchmark, 17 models all claimed 32K tokens or more, but only half held satisfactory performance at 32K (Hsieh and colleagues, 2024). In NoLiMa, a test that removes literal keyword matches, 11 of 13 models fell below half of their short-context score at 32K, and GPT-4o dropped from 99.3 to 69.7 percent (Modarressi and colleagues, 2025). Anthropic’s own documentation now calls the context window the model’s “working memory” and warns that accuracy and recall degrade as token count grows. Humans have the same constraint at a smaller scale: working memory holds about four chunks of new information (Cowan, 2001), and cognitive load theory has shown since 1988 that how information is presented matters as much as how much there is (Sweller, van Merrienboer and Paas, 2019). The CEO move is to treat context as a budget with a cost per token. The student move is to write prompts you could check yourself, using the Load-Fit Prompt method below.

This article belongs to the Prompt and Context Craft series. If you have not yet written a standing description of your job for AI tools, start with the context file and how to configure custom instructions, projects and memory. This piece is about the single request: what to put in it, in what order, and how much.

What is a context window, in plain terms?

A context window is the maximum amount of text, measured in tokens, that a model can read and write in one request. Everything counts against it: your instructions, the documents you paste, the conversation so far, and the model’s answer. Tokens are not words. Google’s documentation gives the conversion most people need: for Gemini models “a token is equivalent to about 4 characters” and “100 tokens is equal to about 60-80 English words”. Other vendors’ tokenizers differ slightly, so treat any conversion as an estimate.

Anthropic’s documentation describes the window in human terms. It says the context window “represents a ‘working memory’ for the model”, that “more context isn’t automatically better”, and that “as token count grows, accuracy and recall degrade, a phenomenon known as context rot”. That is a vendor framing, not an experiment, but it is the right starting point: the context window is not a hard drive. It is closer to a desk, and a cluttered desk slows everyone down.

How big are context windows in 2026?

The table below lists the documented limits for current flagship models, read from each vendor’s model pages on 7 October 2026. These numbers change often; check the vendor page before relying on them.

Vendor Model Context window Max output per request Source
Anthropic Claude Fable 5.1, Opus 5.5, Sonnet 5.5 1M tokens 128K tokens Anthropic models overview
Anthropic Claude Haiku 4.5 200K tokens 64K tokens Anthropic models overview
OpenAI GPT-6 Astra, GPT-6.1 Sol, GPT-6 Luna 1.05M tokens 128K tokens OpenAI models page
Google Gemini 3.1 Pro Preview 1,048,576 tokens 65,536 tokens Gemini API model page
Google Gemini 3.8 Flash 1,048,576 tokens 65,536 tokens Gemini API model page

Google offers a sense of scale: in practice, it says, one million tokens would look like “50,000 lines of code”, “8 average length English novels” or “transcripts of over 200 average length podcast episodes”. Applying Google’s own ratio, a 1,048,576-token window corresponds to roughly 629,000 to 839,000 English words (CEOtudent calculation: 1,048,576 tokens x 0.60 to 0.80 words per token).

The headline number answers one question: how much can you send? It does not answer the question that matters for your work: how much will the model actually use well?

Does a bigger window mean the model uses all of it?

No, and this is the most consistent finding in long-context research.

Position matters. In “Lost in the Middle”, Liu and colleagues (2024) found that “performance is highest when relevant information occurs at the very beginning (primacy bias) or end of its input context (recency bias), and performance significantly degrades when models must access and use information in the middle”. In their multi-document question answering test, GPT-3.5-Turbo’s performance could “drop by more than 20%”, and in the worst case it was lower than its performance with no documents at all (56.1 percent). These were 2023-era models, so the magnitudes do not transfer directly to 2026 models. The shape of the problem, a weak middle, has kept showing up.

Claimed length is not effective length. RULER (Hsieh and colleagues, 2024) tested 17 long-context models on 13 tasks. Almost all scored nearly perfectly on the simple “needle in a haystack” test, but “only half of them can maintain satisfactory performance at the length of 32K”. NoLiMa (Modarressi and colleagues, 2025) made the test harder by removing literal word overlap between the question and the answer, which is closer to how real questions work. Of 13 models claiming at least 128K tokens, “11 models drop below 50% of their strong short-length baselines” at 32K.

Irrelevant context costs accuracy even when the answer is there. Chroma’s “Context Rot” report (2025), an industry report from a retrieval-database company and not peer reviewed, evaluated 18 models. In one experiment it compared focused prompts of about 300 tokens with the full 113k-token version of the same conversational task, which included irrelevant material, and observed “consistent performance degradation with the full” input.

Vendors say the same thing in their own guides. OpenAI’s GPT-4.1 guide reports very good needle-in-a-haystack performance up to 1M tokens but adds that “long context performance can degrade as more items are required to be retrieved, or perform complex reasoning that requires knowledge of the state of the entire context”. Google’s long-context documentation notes that with multiple “needles” the model “does not perform with the same accuracy”.

Claimed versus effective context: what the benchmarks found

The two benchmarks define “effective length” differently. RULER uses the longest length at which a model still beats a fixed threshold (85.6 percent, the score of Llama2-7B at 4K). NoLiMa uses the longest length at which a model keeps at least 85 percent of its own short-context score. The ratio column is a CEOtudent calculation (effective divided by claimed, treating K as 1,000).

Model (benchmark year) Claimed context Effective context Effective as share of claimed Benchmark
GPT-4 (2024) 128K 64K 50% RULER
GPT-4o (2025) 128K 8K 6.3% NoLiMa
Claude 3.5 Sonnet (2025) 200K 4K 2% NoLiMa
Gemini 1.5 Pro (2025) 2M 2K 0.1% NoLiMa

Original synthesis by CEOtudent editorial framework from RULER Table 3 and the NoLiMa results table. These are older models tested on deliberately hard retrieval tasks; current models are likely better, and no independent benchmark of the 2026 models above was available when this was written. The point is the gap, not the exact number.

The lesson for a knowledge worker is simple: the context window is a ceiling, not a guarantee. If your answer depends on something buried in page 140 of a pasted document, the model may miss it, and you may not notice.

What does cognitive science say about human working memory?

The human side of the problem is older and better studied.

The span is small. George Miller’s 1956 paper observed that “this span is about seven items in length”, but he also pointed out that the span is measured in chunks, not raw bits: “we can increase the number of bits of information that it contains simply by building larger and larger chunks”. Nelson Cowan’s 2001 review argued that Miller’s seven was “more as a rough estimate and a rhetorical device than as a real capacity limit” and proposed “a single, central capacity limit averaging about four chunks”.

Load comes in three kinds. Cognitive load theory, which began with John Sweller’s 1988 work on problem solving, distinguishes three sources of load. The 2019 review by Sweller, van Merrienboer and Paas summarises them:
– Intrinsic load is the real complexity of the task. It “only can be changed by changing what needs to be learned or changing the expertise of the learner”.
– Extraneous load is the cost of poor presentation. It “is not determined by the intrinsic complexity of the information but rather, how the information is presented and what the learner is required to do”.
– Germane load was originally defined as the load “required to learn”. The 2019 review reframes it as working memory devoted to the intrinsic task rather than a separate load.

Familiar material is cheap. The same review notes that working memory “was limited in capacity and duration when dealing with novel information but these limitations effectively disappeared when working memory dealt with information transferred from long-term memory”. Experts can handle long, dense inputs in their field because they read them as a few large chunks.

Split and redundant information hurts. Chandler and Sweller (1991) found in six experiments that “split-source information may generate a heavy cognitive load, because material must be mentally integrated before learning can commence”, and that “seemingly useful but nonessential explanatory material” could have “deleterious effects” even when integrated.

What helps novices can hurt experts. Kalyuga and colleagues (2003) named the expertise reversal effect: techniques that are “highly effective with inexperienced learners can lose their effectiveness and even have negative consequences when used with more experienced learners”.

Where do the model and the human meet?

The two literatures were built separately, but they describe the same design problem: a reader with limited attention, a lot of available text, and a task that depends on finding and combining the right parts. Anthropic’s engineering team made the comparison explicit in September 2025: “Like humans, who have limited working memory capacity, LLMs have an ‘attention budget’”, and “every new token introduced depletes this budget by some amount”. The authors of the IFScale benchmark use the same language when they describe models that “begin to struggle under cognitive load” as instruction counts rise. These are analogies, not proof that models and brains work alike. They are useful because the design advice from both sides converges.

Design problem Human evidence Model evidence What it means for your prompt
Too much at once About four chunks of new information (Cowan, 2001) Half of 17 models degrade by 32K (RULER) Send the smallest set of material that answers the question
Position effects Not covered by the sources reviewed here Best at start or end, weak middle (Liu and colleagues, 2024) Put the material first and the question last, or repeat key instructions
Irrelevant material Nonessential explanation can hurt (Chandler and Sweller, 1991) Full 113k-token input underperforms a 300-token focused version (Chroma, 2025) Remove what you would skip yourself
Conflicting instructions Split sources must be integrated before work starts GPT-5 “expends reasoning tokens searching for a way to reconcile the contradictions” (OpenAI) Resolve contradictions before you send
Too many rules Working memory limits on novel rules Best of 20 models reach 68% accuracy at 500 instructions (IFScale) Keep standing rules few and specific
Expertise Familiar material costs little (Sweller and colleagues, 2019) Anthropic: “XML tags help Claude parse complex prompts unambiguously” Structure inputs so they read as a few large chunks

Comparison table: CEOtudent editorial framework, built from the sources cited in each cell.

The Load-Fit Prompt: a five-step method

The method below is a CEOtudent editorial framework. It is not a tested protocol; it translates the evidence above into steps you can apply to any serious prompt. Its premise is that a prompt should fit two working memories at once: the model’s, so it uses the context well, and yours, so you can check the answer.

Step 1. Name the intrinsic load in one sentence. Write what the task really is before you add anything. “Decide whether clause 7 of this contract conflicts with our data policy” is a task. “Look at this contract” is not. If you cannot write the sentence, the problem is not the model.

Step 2. Cut the extraneous load. Go through every document and instruction you were about to paste and ask: would I need this to do the task myself? Remove the rest. Anthropic’s engineers describe the goal as “the smallest possible set of high-signal tokens that maximize the likelihood of some desired outcome”. In practice that means excerpts instead of whole files, one version of a document instead of three drafts, and no background you would skip if a colleague handed it to you.

Step 3. Chunk what remains. Label each block of material with a short heading or tag so it reads as one chunk, the way an expert reads familiar material. Keep standing rules to a short list. The four-chunk heuristic, no more than about four separate demands per request, is an editorial rule of thumb derived from Cowan’s estimate for humans; it is not a measured model limit.

Step 4. Place by position. For long inputs, Anthropic recommends putting “your long documents and inputs near the top of your prompt, above your query”, and says “queries at the end can improve response quality by up to 30 percent in tests”. Google gives the same advice. OpenAI’s GPT-4.1 guide found that placing instructions “at both the beginning and end of the provided context” worked best. A safe default that satisfies all three: a one-line task statement at the top, the material in the middle, and the full question plus output format at the end.

Step 5. Make the answer checkable. Ask the model to quote the passages it relied on before it answers. Anthropic suggests exactly this for long documents: “ask Claude to quote relevant parts of the documents first”. Quotes turn a long context into a short one for you, the reviewer, and they expose when the model reached into the weak middle and came back with the wrong thing.

The Prompt Load Map

Load type In a prompt, it looks like Signal that it is too high Fix
Intrinsic The real difficulty of the task You cannot state the task in one sentence Split the task into sequential prompts
Extraneous Pasted files, old drafts, conflicting rules, long chat history The answer cites the wrong section or ignores a key fact Cut to excerpts; start a fresh conversation with a summary
Germane (redirected) Labels, examples, a stated output format You spend more time interpreting the answer than reading it Add structure: headings, a template, one worked example

CEOtudent editorial framework, adapted from the three load categories in Sweller, van Merrienboer and Paas (2019).

Why does this matter for your own thinking?

There is a second working memory in every prompt: yours. Long prompts and long answers shift effort from thinking to reviewing, and the evidence suggests people do not always keep up.

  • In a survey of 319 knowledge workers who shared 936 examples, Lee and colleagues (CHI 2025) found that “higher confidence in GenAI is associated with less critical thinking, while higher self-confidence is associated with more critical thinking”. The nature of critical thinking shifted “toward information verification, response integration, and task stewardship”. This is self-reported and correlational.
  • Gerlich (2025) surveyed and interviewed 666 participants and reported “a significant negative correlation between frequent AI tool usage and critical thinking abilities, mediated by increased cognitive offloading”. Also correlational.
  • In an MIT Media Lab preprint that has not yet been peer reviewed, Kosmyna and colleagues (2025) followed 54 participants writing essays. In the first session, 83.3 percent of the AI-assisted group (15 of 18) could not correctly quote their own essay, against 11.1 percent (2 of 18) in each of the other two groups. The sample is small and the setting narrow.
  • In a preregistered field experiment with 758 knowledge workers, Dell’Acqua and colleagues (Organization Science, 2026) found that on 18 tasks inside AI’s capabilities, AI users completed 12.2 percent more tasks, 25.1 percent faster, with significantly better quality. On one task outside those capabilities, they were “19% less likely to produce correct solutions”. Knowing which side of the line a task sits on is a human job.

Risko and Gilbert (2016) define cognitive offloading as using “physical action to alter the information processing requirements of a task so as to reduce cognitive demand”. Offloading is not the problem; it is how people have always extended their minds. The problem is offloading the check. A prompt you cannot review in a few minutes is a prompt you will approve without reviewing. For more on that trade-off, see whether AI is making you worse at thinking and the cognitive load budget.

When should you use the long context anyway?

The goal is not short prompts for their own sake. Long context is the right tool when:
– The task is a search across a large body you cannot pre-filter, such as finding every mention of a supplier across a year of meeting notes. Expect misses, and ask for quotes.
– Whole-document consistency is the point, such as checking a 60-page proposal for contradictions. Break it into sections and ask the model to summarise each before comparing them.
– You are setting up a reusable workspace, such as a project with standing reference documents. Here the structure of the material matters more than the size: a clean, labelled context file beats a folder of raw exports. Choosing the right model for that job is covered in which AI model for which task, and keeping your best prompts reusable in prompt libraries for professionals.

It is the wrong tool when the honest reason you are pasting everything is that you have not decided what the question is.

A 10-minute Load-Fit check

Run this before any prompt that will take you more than five minutes to review:

  1. Can I state the task in one sentence?
  2. Have I removed every document I would not need myself?
  3. Is there only one version of each document?
  4. Have I resolved conflicting instructions?
  5. Are there four or fewer separate demands?
  6. Is each block of material labelled?
  7. Is the material at the top and the question at the end?
  8. Have I stated the output format?
  9. Have I asked for quotes or references to the source passages?
  10. Could I check the answer in the time I have?

Checklist: CEOtudent editorial framework.

Frequently asked questions

Is a 1M-token context window useless then?
No. It removes a hard limit, which is valuable for search and for reusable workspaces. The benchmarks show that the quality of use falls off long before the limit, so the window is best treated as headroom rather than a target.

Do these benchmark results apply to the newest models?
Not directly. RULER and NoLiMa tested models from 2024 and early 2025, and newer models are likely to do better. Vendor guides for current models still recommend focused context and careful placement, and no independent benchmark of the 2026 models listed above was available when this article was written.

Should instructions go at the top or the bottom?
Vendors differ slightly. Anthropic and Google recommend long material first and the question last. OpenAI’s GPT-4.1 guide found instructions at both the start and end worked best, and above the context better than below if you only use them once. A short task line at the top plus the full question at the end covers both.

Is “seven plus or minus two” still the rule for working memory?
Miller himself described seven as approximate and measured in chunks. Cowan’s 2001 review places the limit at about four chunks of new information. The exact number matters less than the principle: fewer, larger, well-labelled chunks.

Does this mean AI makes people think less?
The evidence is mixed and mostly correlational. The safest reading is that confidence in the tool tends to reduce the effort people spend checking it. Designing prompts whose answers you can verify keeps that effort in place.

Sources

  1. Anthropic. Models overview; Context windows; Prompting best practices (long context prompting). Claude Platform Docs. Accessed 7 October 2026.
  2. Anthropic Applied AI team. Effective context engineering for AI agents. Anthropic Engineering, 29 September 2025.
  3. OpenAI. Models. OpenAI API documentation. Accessed 7 October 2026.
  4. OpenAI. GPT-4.1 Prompting Guide; GPT-5 prompting guide. OpenAI Cookbook, 2025.
  5. Google. Gemini 3.1 Pro Preview and Gemini 3.8 Flash model pages; Long context; Understand and count tokens. Gemini API documentation. Accessed 7 October 2026.
  6. Liu NF, Lin K, Hewitt J, Paranjape A, Bevilacqua M, Petroni F, Liang P. Lost in the Middle: How Language Models Use Long Contexts. Transactions of the Association for Computational Linguistics. 2024;12:157-173.
  7. Hsieh CP, Sun S, Kriman S, Acharya S, Rekesh D, Jia F, Zhang Y, Ginsburg B. RULER: What’s the Real Context Size of Your Long-Context Language Models? COLM 2024.
  8. Modarressi A, Deilamsalehy H, Dernoncourt F, Bui T, Rossi R, Yoon S, Schuetze H. NoLiMa: Long-Context Evaluation Beyond Literal Matching. ICML 2025.
  9. Hong K, Troynikov A, Huber J. Context Rot: How Increasing Input Tokens Impacts LLM Performance. Chroma Technical Report, 14 July 2025. Industry report, not peer reviewed.
  10. Jaroslawicz D, Whiting B, Shah P, Maamari K. How Many Instructions Can LLMs Follow at Once? arXiv 2507.11538, 2025. Preprint.
  11. Miller GA. The magical number seven, plus or minus two: Some limits on our capacity for processing information. Psychological Review. 1956;63(2):81-97.
  12. Cowan N. The magical number 4 in short-term memory: A reconsideration of mental storage capacity. Behavioral and Brain Sciences. 2001;24(1):87-114. Abstract.
  13. Sweller J. Cognitive load during problem solving: Effects on learning. Cognitive Science. 1988;12(2):257-285. Abstract.
  14. Sweller J, van Merrienboer JJG, Paas F. Cognitive Architecture and Instructional Design: 20 Years Later. Educational Psychology Review. 2019;31:261-292.
  15. Chandler P, Sweller J. Cognitive Load Theory and the Format of Instruction. Cognition and Instruction. 1991;8(4):293-332. Abstract.
  16. Kalyuga S, Ayres P, Chandler P, Sweller J. The Expertise Reversal Effect. Educational Psychologist. 2003;38(1):23-31. Abstract.
  17. Risko EF, Gilbert SJ. Cognitive Offloading. Trends in Cognitive Sciences. 2016;20(9):676-688. Abstract.
  18. Lee HP, Sarkar A, Tankelevitch L, Drosos I, Rintel S, Banks R, Wilson N. The Impact of Generative AI on Critical Thinking: Self-Reported Reductions in Cognitive Effort and Confidence Effects From a Survey of Knowledge Workers. CHI 2025.
  19. Gerlich M. AI Tools in Society: Impacts on Cognitive Offloading and the Future of Critical Thinking. Societies. 2025;15(1):6. Abstract.
  20. Kosmyna N, Hauptmann E, Yuan YT, Situ J, Liao XH, Beresnitzky AV, Braunstein I, Maes P. Your Brain on ChatGPT: Accumulation of Cognitive Debt when Using an AI Assistant for Essay Writing Task. MIT Media Lab, arXiv 2506.08872, 2025. Preprint, not peer reviewed.
  21. Dell’Acqua F, McFowland E III, Mollick E, Lifshitz-Assaf H, and colleagues. Navigating the Jagged Technological Frontier: Field Experimental Evidence of the Effects of Artificial Intelligence on Knowledge Worker Productivity and Quality. Organization Science. 2026;37(2). Abstract.

The claimed-versus-effective table, the human-model comparison table, the Load-Fit Prompt method, the Prompt Load Map and the 10-minute check are CEOtudent analyses built on the sources above. All other figures are reported as printed in those sources; sources read in abstract form only are marked. Two widely repeated figures were not used because they could not be found in the published texts: a “40 percent higher quality” result attributed to the jagged-frontier study, and a “three quarters of a word per token” rule attributed to OpenAI.


This content was compiled with the support of AI following in-depth research, then written and prepared for publication by the CEOtudent editorial team.

This post is also available in: Türkçe Français Español Deutsch

Benzer içerikler