RAG, explained
Retrieval-augmented generation, or RAG, is how an AI answers from your own documents and shows where it looked. Here is how it works, when it beats a chatbot on its own, and why its quality depends mostly on search.
One of the most common things businesses ask of AI is to answer questions from their own documents, such as policies, contracts, product manuals and past proposals. A general chatbot cannot do that well. It has never seen your files, its knowledge stops at the date it was trained, and it will often produce a confident answer anyway.
Retrieval-augmented generation, usually shortened to RAG, is the standard way to fix this. The name comes from a 2020 research paper by Patrick Lewis and colleagues at Facebook AI Research, now Meta AI. The idea is simple. Before the AI answers, look up the relevant passages in your documents and give them to the model to read.
How RAG works
A RAG system has two stages. The first is done once, and repeated whenever your documents change. The second happens every time someone asks a question.
Collect your documentsYou
Gather the files the AI should answer from, such as policies, manuals and help articles.
Shared driveWebsite
Split them into chunksAgent
Each document is cut into short passages, often a few paragraphs, so a search can return just the part that matters.
Tag each chunk by meaningAgent
An embedding model turns each chunk into a list of numbers that captures what it means, so passages about the same idea sit close together even when they use different words.
Store them in a search indexAgent
The chunks and their numbers go into a search index, often called a vector database, ready to be searched.
Search index
Someone asks a questionYou
For example, a customer asks, what is your refund window?
Find the best passagesAgent
The question is tagged by meaning in the same way and the index returns the closest few chunks.
Search index
The model reads them togetherAgent
The question and the retrieved passages are sent to the AI model with an instruction to answer only from what it has been given.
AI model
Answer with sourcesAgent
The reply names the passages it used, so a person can click through and check.
Simplified. Real systems usually retrieve between five and twenty passages, not three.
Chat alone or with RAG?
A chatbot on its own answers from what it learned in training. That is fine for general questions and poor for anything specific to your business. RAG changes four things.
- It knows your documents. The answer comes from your own policies and files, not the public internet.
- It stays up to date. A chatbot stops at its training date. A RAG system is as fresh as the files in its index.
- It shows its sources. Each answer can point to the passage it came from, which makes checking quick.
- Teaching it something new is easy. You add or update a file, instead of retraining the model.
RAG does not make an AI infallible. A 2024 Stanford study of legal research tools built this way found they still produced wrong or unsupported answers between 17% and 33% of the time, fewer than general chatbots but far from zero. Sources make errors easier to catch; they do not remove them.
Search decides the quality
If the right passage is never found, the model has nothing to answer from and will guess. That makes the search step the part most worth improving.
In September 2024 Anthropic published tests of retrieval methods. They measured how often the correct passage was missing from the top 20 results. Searching by meaning alone missed 5.7% of the time. Adding a short note to each chunk before tagging it, saying which document it comes from and what that part covers, cut that to 3.7%. Also searching by exact keywords brought it to 2.9%. A second pass that re-ranks the results brought it to 1.9%, 67% fewer misses than where they started.
Four ways RAG goes wrong, and the fix
- Messy files. Out-of-date versions, duplicates and scanned pages that cannot be read. The fix is to tidy the source first, and decide which document is the official one.
- Bad chunks. Passages cut mid-sentence or split from their heading lose their meaning. The fix is to split by section, and keep each heading with its text.
- Wrong match. Meaning-only search misses exact terms. The fix is to search by meaning and by words, then re-rank.
- No testing. Nobody checks whether answers stay right as files change. The fix is to keep a list of real questions with known answers and rerun it weekly.
Is RAG right for your business?
RAG suits any case where people repeatedly look things up in a large, changing set of documents, such as customer support, internal policy questions, sales teams answering product queries, or staff searching past work. It is less useful when the answers need calculation across a database, where a direct connection to that system does better, or when the documents are few enough to paste straight into a conversation.