Productivity 19-Year-Old Founder Emerges From Stealth With $11 Million to Sell You a $3,499 'Brain in a Box' That Runs Your AI Agents at Home AI Agents Half a Million Interviews In: HackerRank's AI Interviewer Chakra Goes GA, and It Wants to Replace Three Hiring Rounds With One Security After Claude Agents Escaped Its Sandbox 3 Times, Anthropic Deploys Real-Time Classifiers to Stop the Next Escape Before It Happens Business Meta Halves Its Internal Claude Users to 30,000 and Microsoft Slashes a $1 Billion Anthropic Budget by More Than a Third Business Sony Innovation Fund Backs Primitive Labs, a Startup That Builds Simulated Crowds to Stress-Test Products Before Launch Apple Intelligence Apple Removed the Apple Intelligence Off Switch in macOS 27 — So a Developer Built a CLI That Deletes It Anyway Security OpenAI Turns On Invisible Text Watermarks for ChatGPT in the EU — and Publishes Exactly How Weak They Are News A Mystery 'Space Bunny Alpha' Model Just Topped OpenRouter's Leaderboard With 38.7 Trillion Tokens a Week Productivity 19-Year-Old Founder Emerges From Stealth With $11 Million to Sell You a $3,499 'Brain in a Box' That Runs Your AI Agents at Home AI Agents Half a Million Interviews In: HackerRank's AI Interviewer Chakra Goes GA, and It Wants to Replace Three Hiring Rounds With One Security After Claude Agents Escaped Its Sandbox 3 Times, Anthropic Deploys Real-Time Classifiers to Stop the Next Escape Before It Happens Business Meta Halves Its Internal Claude Users to 30,000 and Microsoft Slashes a $1 Billion Anthropic Budget by More Than a Third Business Sony Innovation Fund Backs Primitive Labs, a Startup That Builds Simulated Crowds to Stress-Test Products Before Launch Apple Intelligence Apple Removed the Apple Intelligence Off Switch in macOS 27 — So a Developer Built a CLI That Deletes It Anyway Security OpenAI Turns On Invisible Text Watermarks for ChatGPT in the EU — and Publishes Exactly How Weak They Are News A Mystery 'Space Bunny Alpha' Model Just Topped OpenRouter's Leaderboard With 38.7 Trillion Tokens a Week

Retrieval-Augmented Generation: How It Works

Retrieval-Augmented Generation: How It Works

RAG fetches relevant passages from your own documents and places them in the prompt, so answers are grounded in material the model was never trained on.

Retrieval-augmented generation combines a search system with a language model. Instead of relying purely on what the model memorised, the system finds relevant text and supplies it as context for the current question.

The pipeline

  • Chunking. Documents are split into passages small enough to be precise and large enough to be coherent.
  • Embedding. Each chunk is converted into a vector representing its meaning.
  • Indexing. Vectors are stored in a database that supports similarity search.
  • Retrieval. The user's question is embedded and the closest chunks are fetched.
  • Generation. The model receives the question plus the retrieved passages and answers from them.

Where RAG fails

Most failures are retrieval failures, not generation failures. Chunks that split a table, break a list, or separate a heading from its explanation rarely retrieve correctly. Overlapping chunks and keeping headings attached to their content fix many of these cases.

Hybrid search helps

Vector similarity handles paraphrase but misses exact identifiers — part numbers, error codes, names. Combining keyword search with vector search and merging the results solves most of it.

Cite everything

Return the source chunk with the answer. Citations let users verify in seconds and give you a diagnostic signal: when an answer is wrong, the citation usually shows why.

Comments (0)

Log in to join the discussion

Log In

No comments yet