Retrieval systems live or die on how documents are split. Get chunking right and a modest model performs well; get it wrong and no amount of prompt tuning recovers it.
Chunking rules that work
- Split on structure, not character count. Headings, list boundaries and paragraph breaks are natural units.
- Keep headings with their content. Prefix each chunk with the heading path so a retrieved passage carries its context.
- Overlap slightly. A small overlap prevents answers that straddle a boundary from being cut in half.
- Never split a table. Convert it to a compact textual form or keep it whole, even if the chunk grows.
Choosing a size
Aim for chunks in the range of one to three paragraphs. Smaller chunks retrieve more precisely but lose surrounding explanation; larger chunks retrieve noise. Test both on your own question set — the right answer depends on your documents.
Metadata is not optional
Store source file, section title, last-modified date and a URL with every chunk. You need it for citations, for filtering by recency, and for deleting stale content when source documents change.
Refresh on a schedule
An index that is not refreshed slowly becomes wrong. Tie reindexing to your source system's update events, or rebuild nightly if events are unavailable.
Comments (0)
Log in to join the discussion
Log InNo comments yet