Retrieval-augmented generation combines a search system with a language model. Instead of relying purely on what the model memorised, the system finds relevant text and supplies it as context for the current question.
The pipeline
- Chunking. Documents are split into passages small enough to be precise and large enough to be coherent.
- Embedding. Each chunk is converted into a vector representing its meaning.
- Indexing. Vectors are stored in a database that supports similarity search.
- Retrieval. The user's question is embedded and the closest chunks are fetched.
- Generation. The model receives the question plus the retrieved passages and answers from them.
Where RAG fails
Most failures are retrieval failures, not generation failures. Chunks that split a table, break a list, or separate a heading from its explanation rarely retrieve correctly. Overlapping chunks and keeping headings attached to their content fix many of these cases.
Hybrid search helps
Vector similarity handles paraphrase but misses exact identifiers — part numbers, error codes, names. Combining keyword search with vector search and merging the results solves most of it.
Cite everything
Return the source chunk with the answer. Citations let users verify in seconds and give you a diagnostic signal: when an answer is wrong, the citation usually shows why.
Comments (0)
Log in to join the discussion
Log InNo comments yet