Retrieval-augmented generation sounds more complicated than it is. A working version is a few hundred lines, and most of the quality comes from three decisions.
Ingest and chunk
Convert your sources to plain text, then split on structural boundaries — headings, list items, paragraph breaks. Prefix each chunk with its heading path so a retrieved passage carries context. Keep tables intact.
Embed and store
Generate a vector for every chunk and store it with metadata: source, section, date, URL. Any vector store works at this scale; choose one that supports metadata filtering so you can restrict by date or source.
Retrieve hybrid, not pure semantic
Combine keyword search with vector search and merge results. Pure vector search misses exact identifiers, product codes and names that users type verbatim. Hybrid retrieval fixes this with modest extra complexity.
Generate with citations
Pass the top few chunks with their sources, and require the model to cite which chunk each statement came from. If it cannot cite, it should say the information is not available.
The three knobs that matter
- Chunk size. One to three paragraphs. Test both ends on your own questions.
- Number of chunks. Start at five. More is not better — extra chunks add noise and cost.
- Similarity threshold. Below the threshold, refuse rather than guess.
Evaluate continuously
Keep a set of questions with known correct answers. Every time you change chunking or retrieval, re-run it. Without this, improvements are guesses and regressions go unnoticed.
Comments (0)
Log in to join the discussion
Log InNo comments yet