Productivity 19-Year-Old Founder Emerges From Stealth With $11 Million to Sell You a $3,499 'Brain in a Box' That Runs Your AI Agents at Home AI Agents Half a Million Interviews In: HackerRank's AI Interviewer Chakra Goes GA, and It Wants to Replace Three Hiring Rounds With One Security After Claude Agents Escaped Its Sandbox 3 Times, Anthropic Deploys Real-Time Classifiers to Stop the Next Escape Before It Happens Business Meta Halves Its Internal Claude Users to 30,000 and Microsoft Slashes a $1 Billion Anthropic Budget by More Than a Third Business Sony Innovation Fund Backs Primitive Labs, a Startup That Builds Simulated Crowds to Stress-Test Products Before Launch Apple Intelligence Apple Removed the Apple Intelligence Off Switch in macOS 27 — So a Developer Built a CLI That Deletes It Anyway Security OpenAI Turns On Invisible Text Watermarks for ChatGPT in the EU — and Publishes Exactly How Weak They Are News A Mystery 'Space Bunny Alpha' Model Just Topped OpenRouter's Leaderboard With 38.7 Trillion Tokens a Week Productivity 19-Year-Old Founder Emerges From Stealth With $11 Million to Sell You a $3,499 'Brain in a Box' That Runs Your AI Agents at Home AI Agents Half a Million Interviews In: HackerRank's AI Interviewer Chakra Goes GA, and It Wants to Replace Three Hiring Rounds With One Security After Claude Agents Escaped Its Sandbox 3 Times, Anthropic Deploys Real-Time Classifiers to Stop the Next Escape Before It Happens Business Meta Halves Its Internal Claude Users to 30,000 and Microsoft Slashes a $1 Billion Anthropic Budget by More Than a Third Business Sony Innovation Fund Backs Primitive Labs, a Startup That Builds Simulated Crowds to Stress-Test Products Before Launch Apple Intelligence Apple Removed the Apple Intelligence Off Switch in macOS 27 — So a Developer Built a CLI That Deletes It Anyway Security OpenAI Turns On Invisible Text Watermarks for ChatGPT in the EU — and Publishes Exactly How Weak They Are News A Mystery 'Space Bunny Alpha' Model Just Topped OpenRouter's Leaderboard With 38.7 Trillion Tokens a Week

Context Windows Keep Growing - Here Is What Actually Breaks

Context Windows Keep Growing - Here Is What Actually Breaks

Bigger context does not mean better recall. Models still lose information in the middle of long inputs, and cost scales with everything you paste in.

Advertised context lengths now run into the millions of tokens. In practice, dumping an entire archive into a prompt is one of the fastest ways to get a worse answer at a higher price.

Three effects that show up immediately

  • Position bias. Information at the very beginning and very end of a long input is recalled more reliably than material buried in the middle.
  • Attention dilution. As the input grows, relevance per token falls. Irrelevant passages crowd out the ones that matter.
  • Cost and latency. Both scale with input size, and long prompts make streaming feel sluggish.

What works instead

Retrieval still earns its keep. Split documents into coherent sections, index them, and fetch the handful of passages that match the question. A well-tuned retrieval step routinely outperforms a naive full-document paste, and it costs a fraction as much.

Where full context genuinely helps is when the task depends on the whole document — a contract, a specification, a novel. Even then, structure the input: headings, explicit section markers and a short summary at the top measurably improve results.

A useful mental model

Treat context as a budget rather than a container. Every passage you include should be able to justify its place. If you cannot say why a paragraph is there, removing it will usually help.

Comments (0)

Log in to join the discussion

Log In

No comments yet