Advertised context lengths now run into the millions of tokens. In practice, dumping an entire archive into a prompt is one of the fastest ways to get a worse answer at a higher price.
Three effects that show up immediately
- Position bias. Information at the very beginning and very end of a long input is recalled more reliably than material buried in the middle.
- Attention dilution. As the input grows, relevance per token falls. Irrelevant passages crowd out the ones that matter.
- Cost and latency. Both scale with input size, and long prompts make streaming feel sluggish.
What works instead
Retrieval still earns its keep. Split documents into coherent sections, index them, and fetch the handful of passages that match the question. A well-tuned retrieval step routinely outperforms a naive full-document paste, and it costs a fraction as much.
Where full context genuinely helps is when the task depends on the whole document — a contract, a specification, a novel. Even then, structure the input: headings, explicit section markers and a short summary at the top measurably improve results.
A useful mental model
Treat context as a budget rather than a container. Every passage you include should be able to justify its place. If you cannot say why a paragraph is there, removing it will usually help.
Comments (0)
Log in to join the discussion
Log InNo comments yet