Retrieval-Augmented Generation (RAG) combines a language model with an external knowledge source. Documents are indexed, a retriever selects relevant passages for a user query, and those passages are supplied to the model as context before generation. RAG is useful when knowledge changes frequently, must remain traceable to documents, or should not be permanently baked into model weights. With open-weight models, inference can be self-hosted, but privacy still depends on the embedding service, vector store, logs and document pipeline.
What RAG changes compared with a standalone model
A language model stores patterns and some factual knowledge in its parameters, but those parameters are expensive to update and do not provide a reliable database interface. The original RAG research combined parametric model memory with non-parametric retrieved memory so generation could use external evidence.
In a production system, the pattern is broader: documents are ingested, parsed, split into chunks, represented for search, stored in an index and retrieved when a user asks a question. The model receives the top passages as temporary context. Updating the knowledge base can therefore be as simple as re-indexing documents rather than fine-tuning the model.
The RAG pipeline has more components than the model
| Component | Job | Failure mode |
|---|---|---|
| Ingestion | Extract text/structure from source documents | Missing tables, bad OCR, stale files |
| Chunking | Create retrievable units | Context split across chunks or chunks too large |
| Embeddings / search | Represent and retrieve relevant content | Semantically similar but wrong passages |
| Reranking | Improve ordering of candidate passages | Latency or poor ranking model |
| Prompt assembly | Combine evidence, question and instructions | Prompt injection or context overflow |
| Generation | Produce answer from supplied evidence | Hallucination beyond sources |
Permissions are a first-class RAG problem
A vector database that contains every internal document is not automatically safe. Retrieval must respect the user’s authorization. If an employee is not allowed to read a document in the source system, the RAG layer should not surface a chunk from that document merely because its embedding is similar to the query.
Practical designs attach metadata such as tenant, department, ACL or document ID to indexed chunks and apply authorization filters before or during retrieval. The application should also avoid leaking hidden documents through citations, summaries or aggregate answers.
How to improve answer quality without changing the model
RAG quality is often limited by retrieval rather than generation. Before switching to a larger model, inspect whether the correct evidence reached the prompt.
- Improve document parsing and preserve headings, tables and metadata.
- Choose chunk sizes that match the semantic structure of the corpus.
- Use hybrid lexical + vector search where exact names and semantic similarity both matter.
- Rerank a broader candidate set before prompting the model.
- Require citations and let the model say when evidence is insufficient.
- Evaluate retrieval recall separately from answer quality.
Self-hosted RAG and data residency
Open-weight generation can keep inference on chosen infrastructure, but the RAG system may still call external services for embeddings, OCR, reranking, vector search, telemetry or backups. Each component can create a separate data transfer.
For a genuinely self-contained deployment, map the complete path: document ingestion → embedding → index → retrieval → prompt construction → inference → logging → monitoring → backup. Provider origin of the model is less important for data residency than where these services actually process personal or confidential data.
RAG versus fine-tuning
RAG and fine-tuning solve different problems. RAG is usually better for changing factual knowledge, document provenance and access-controlled corpora. Fine-tuning is better for behavioral adaptation: style, format, domain conventions, tool policies or specialized task patterns.
They can be combined. A model can be fine-tuned for a company’s response format while RAG supplies current product manuals and policies. The important design choice is to keep volatile facts outside the weights when they need frequent updates or source-level traceability.
Frequently asked questions
Does RAG train the model on my documents?
No. Standard RAG retrieves documents at request time and places them in the prompt/context. The model weights do not need to change.
Do I need a vector database for RAG?
Not always. Retrieval can be lexical, vector or hybrid. A dedicated vector database is common at scale but not a requirement for every corpus.
Is self-hosted RAG automatically private?
No. Embeddings, OCR, vector storage, logs, monitoring and backups must also be self-hosted or otherwise governed if the goal is a contained data path.
Should I fine-tune instead of using RAG?
Use RAG for changeable knowledge and provenance; use fine-tuning for behavior and task adaptation. Many production systems use both.
Apply this knowledge
Move from the concept to a concrete deployment shortlist.
Primary sources and technical references
OpenWeightModels prefers publisher documentation, standards bodies, official repositories and original research papers. The source material remains authoritative where it changes.