RAG, short for retrieval-augmented generation, is a way of making an AI model look things up before it answers. Instead of relying only on what it learned in training, the system first finds the right documents, then hands them to the model with the question, so the answer is grounded in your real information rather than a confident guess.
The problem: a model only knows its training data
Ask a chat model 'What's our vacation policy?' or 'How do I reset the build server?' and it has a problem. Those answers live in your company's documents, which the model never saw during training. It may also have a cut-off date, so even public facts can be out of date.
A model in that position doesn't usually say 'I don't know'. It predicts a plausible answer, which is how you get confident nonsense, the problem covered in AI hallucinations.
RAG's fix is simple to say: don't just answer. Go look it up first.
The three steps
The name lists the steps in order.
- Retrieval. When a question comes in, the system searches the right documents for the passages most likely to answer it: HR policies, support tickets, manuals, legal files, that sacred folder of PDFs.
- Augmentation. It takes the user's question and the useful snippets and puts them together into the prompt. Now the model has context.
- Generation. The model writes an answer grounded in that retrieved information.
The model is still the writer. The documents are the notes on the desk.
Here is one question going through all three steps:
A question going through retrieval, augmentation and generation
Step 1 of 7: You ask what the vacation policy is.
A worked example in Python
This toy version shows the whole idea with no AI libraries at all. Retrieval here just counts the words a document shares with the question; real systems search by meaning, but the shape is the same:
import re
docs = [
"Vacation policy: staff get 25 days of leave a year.",
"Up to 5 unused leave days carry over each year.",
"Build server reset: run reset.sh, then wait 5 min.",
"Expense claims must be filed within 30 days.",
]
def words(text):
return set(re.findall(r"[a-z0-9]+", text.lower()))
def retrieve(question, k=2):
q = words(question)
return sorted(docs, key=lambda d: len(q & words(d)),
reverse=True)[:k]
question = "What is our vacation policy on leave days?"
notes = retrieve(question)
prompt = (
"Answer using only these notes. "
"If they don't cover it, say so.\n\n"
+ "\n".join(f"- {n}" for n in notes)
+ f"\n\nQuestion: {question}"
)retrieve picks the two leave-policy lines and skips the build server and expenses ones. The prompt it builds looks like this:
Answer using only these notes. If they don't cover it, say so.
- Vacation policy: staff get 25 days of leave a year.
- Up to 5 unused leave days carry over each year.
Question: What is our vacation policy on leave days?The last step sends prompt to whichever model you use. Two lines in it do a lot of work. 'Answer using only these notes' keeps the model on your documents, and 'If they don't cover it, say so' gives it permission not to guess.
How real retrieval works
Word matching breaks quickly: a question about 'holiday' won't match a document that says 'annual leave'. Real RAG systems usually search by meaning instead:
- Chunking. Long documents are split into small passages, often a few paragraphs each, so the system can fetch just the part that matters.
- Embeddings. Each chunk is turned into an embedding: a list of numbers that captures its meaning, so passages about similar ideas end up close together.
- Similarity search. The question is embedded the same way, and a vector database finds the chunks closest to it.
Many systems combine this with keyword search, which is still better at exact terms such as error codes and product names.
Why companies use RAG
- More specific. The model can answer about your policies, products and systems, not just the world in general.
- More up to date. Change a document and the next answer changes with it. There is no retraining.
- Less likely to hallucinate. With the facts in front of it, the model has far less reason to invent tuna regulations. It still can, so grounding lowers the risk rather than removing it.
- Checkable. The app knows which passages it used, so it can show the sources and people can check the receipts.
That's why RAG sits behind so many internal chatbots, customer support assistants and tools for document-heavy work. Choosing what goes into the prompt, and how, is a large part of context engineering.
RAG or fine-tuning?
Both make a model better at your work, and they are easy to mix up. Fine-tuning trains the model further on examples, changing its weights, and is best for style, tone and format. RAG leaves the model alone and changes what it reads, and is best for facts, especially facts that change. A support bot often uses both: fine-tuned to sound like your team, with RAG to know today's policies. If you need exact facts from private documents, start with RAG.
When RAG isn't the answer
- The problem is reasoning, not facts. RAG gives the model information; it doesn't make it better at maths or planning.
- The knowledge is small and stable. If it all fits comfortably in the prompt, just include it and skip the search.
- The documents are bad. Out-of-date, contradictory or badly written sources produce answers to match.
Common mistakes
- Blaming the model for bad retrieval. If the right passage never reached the prompt, no model can use it. Check what was retrieved first.
- Chunks of the wrong size. Too big and the prompt fills with noise; too small and a passage loses its meaning.
- No 'say so' instruction. Without it, the model fills gaps with guesses.
- Ignoring permissions. If the search can see every document, the chatbot can leak HR files to anyone who asks. Only retrieve what the person asking is allowed to read.
Key takeaways
- RAG stands for retrieval-augmented generation: don't just answer, look it up first.
- Retrieval finds the right documents, augmentation puts the snippets in the prompt, and generation writes a grounded answer.
- It makes answers more specific, more up to date and less likely to be hallucinated.
- Real systems split documents into chunks and search them by meaning with embeddings.
- The answer is only as good as what was retrieved, so test the search as well as the model.