Business Sony Innovation Fund Backs Primitive Labs, a Startup That Builds Simulated Crowds to Stress-Test Products Before Launch Apple Intelligence Apple Removed the Apple Intelligence Off Switch in macOS 27 — So a Developer Built a CLI That Deletes It Anyway Security OpenAI Turns On Invisible Text Watermarks for ChatGPT in the EU — and Publishes Exactly How Weak They Are News A Mystery 'Space Bunny Alpha' Model Just Topped OpenRouter's Leaderboard With 38.7 Trillion Tokens a Week News AMD Goes All-In on the 192GB "Agentic PC" Two Days Before the Nvidia RTX Spark Event: 300-Billion-Parameter Models, No Cloud Required ChatGPT OpenAI Will Put Sponsored Images Inside ChatGPT Image Generation — Testing in the US This Month for Its 1.2 Billion Weekly Users Opinion Hinton, Bengio and 20 Other Top Researchers Warn a Year of AI Progress Could Soon Take Five Weeks Business Schneider Electric to Buy PTC for $22.6 Billion in Its Largest-Ever Deal — and Its Stock Dropped 9% on the Price Tag Business Sony Innovation Fund Backs Primitive Labs, a Startup That Builds Simulated Crowds to Stress-Test Products Before Launch Apple Intelligence Apple Removed the Apple Intelligence Off Switch in macOS 27 — So a Developer Built a CLI That Deletes It Anyway Security OpenAI Turns On Invisible Text Watermarks for ChatGPT in the EU — and Publishes Exactly How Weak They Are News A Mystery 'Space Bunny Alpha' Model Just Topped OpenRouter's Leaderboard With 38.7 Trillion Tokens a Week News AMD Goes All-In on the 192GB "Agentic PC" Two Days Before the Nvidia RTX Spark Event: 300-Billion-Parameter Models, No Cloud Required ChatGPT OpenAI Will Put Sponsored Images Inside ChatGPT Image Generation — Testing in the US This Month for Its 1.2 Billion Weekly Users Opinion Hinton, Bengio and 20 Other Top Researchers Warn a Year of AI Progress Could Soon Take Five Weeks Business Schneider Electric to Buy PTC for $22.6 Billion in Its Largest-Ever Deal — and Its Stock Dropped 9% on the Price Tag

Building a Working RAG Pipeline from Scratch

Index, retrieve, generate, evaluate — the minimum viable version, with the tuning knobs that matter most.

Retrieval-augmented generation sounds more complicated than it is. A working version is a few hundred lines, and most of the quality comes from three decisions.

Ingest and chunk

Convert your sources to plain text, then split on structural boundaries — headings, list items, paragraph breaks. Prefix each chunk with its heading path so a retrieved passage carries context. Keep tables intact.

Embed and store

Generate a vector for every chunk and store it with metadata: source, section, date, URL. Any vector store works at this scale; choose one that supports metadata filtering so you can restrict by date or source.

Retrieve hybrid, not pure semantic

Combine keyword search with vector search and merge results. Pure vector search misses exact identifiers, product codes and names that users type verbatim. Hybrid retrieval fixes this with modest extra complexity.

Generate with citations

Pass the top few chunks with their sources, and require the model to cite which chunk each statement came from. If it cannot cite, it should say the information is not available.

The three knobs that matter

  • Chunk size. One to three paragraphs. Test both ends on your own questions.
  • Number of chunks. Start at five. More is not better — extra chunks add noise and cost.
  • Similarity threshold. Below the threshold, refuse rather than guess.

Evaluate continuously

Keep a set of questions with known correct answers. Every time you change chunking or retrieval, re-run it. Without this, improvements are guesses and regressions go unnoticed.

Comments (0)

Log in to join the discussion

Log In

No comments yet