The architecture behind grounded AI
Retrieval-augmented generation (RAG) is an AI architecture that retrieves relevant information from an external knowledge source before a large language model generates an answer. This allows AI applications to work with current, private, or domain-specific information instead of relying only on the model's training data.
Learn how retrieval connects your data with an LLM at query time
Plain answer
Retrieval-augmented generation (RAG) is an AI architecture that improves large language model responses by retrieving relevant information from an external knowledge base at query time and providing that information to the model as context before it generates an answer. Instead of relying only on information learned during training, the application can use current, private, or domain-specific source content when responding to a user.
Step by step
A typical RAG system retrieves relevant information first and then provides that information to the language model as context for generating the answer.
A user asks a question in natural language.
The query is converted into a numerical representation that captures its semantic meaning.
Search retrieves relevant chunks from a knowledge base or vector database.
Retrieved content is added to the query and supplied to the LLM as context.
The LLM generates a response using the retrieved information, with sources available when the system supports citations.
The decision
RAG, plain language models, and fine-tuning solve different problems. They can also be combined when an application needs both task-specific behavior and access to current or proprietary information.
| Factor | Plain LLM | RAG | Fine-tuning |
|---|---|---|---|
| Knowledge source | Training data | Training data plus retrieved external content | Training data plus task-specific training examples |
| Update process | Model updates require new training or a newer model | Update the external knowledge source without retraining the underlying model | Requires additional training when model behavior needs to be changed |
| Best for | General language and knowledge tasks | Current, domain-specific, or private information | Task-specific behavior, style, or output patterns |
| Traceability | No retrieval source by default | Retrieved sources can be cited | No retrieval source by default |
Core concepts
A RAG application is more than connecting an LLM to a vector database. The quality of retrieval, source preparation, context selection, and evaluation all influence the final response.
Source material such as PDFs, web pages, wikis, databases, and other documents is collected and prepared for retrieval.
Large documents are divided into smaller pieces so the retrieval system can find information relevant to a particular question.
Text can be converted into numerical representations that allow semantically similar queries and documents to be found.
The system searches the knowledge source and selects the information most relevant to the user's question.
Relevant retrieved content is provided to the language model so it can generate an answer using that context.
Retrieval and answer quality should be evaluated against representative questions before the system is relied upon in production.
Why it matters
A language model's training data does not automatically provide access to an organization's private documents, changing product information, internal policies, or other external knowledge. RAG provides a way for an AI application to retrieve relevant information at query time and use it as context when generating a response.
Practical guide
RAG is particularly useful when an AI application needs access to information that is private, frequently changing, or specific to a particular domain.
Use RAG when an assistant needs to answer questions using internal documents, policies, manuals, or company knowledge.
RAG can retrieve updated information without requiring the underlying language model to be retrained whenever documents change.
RAG is useful when users need answers that can be connected back to the documents or sources used by the application.
Build with RAG
If you are planning a production RAG application, you can learn more about the development process, architecture, and engagement options on the dedicated RAG development page.
RAG pipeline
FAQ