Learn / Basics
What is retrieval-augmented generation (RAG)?
Short answer
Retrieval-augmented generation (RAG) is a way of making an AI model answer from your own documents. Before the model writes a reply, a search step finds the most relevant passages and hands them to the model as context, so the answer is grounded in real sources instead of the model's memory.
With troveGEN
troveGEN is RAG as an API. You send documents; it reads, structures, protects and indexes them, and your app gets accurate, permission-aware search with sources.
Why RAG exists
A large language model only knows what was in its training data, and that data has a cut-off date. It does not know your price list, your contracts, your patient guidelines or last week's circular. Asked about them, it may guess, and a confident guess that is wrong is called a hallucination.
RAG fixes this without retraining the model. The knowledge stays in your documents, where you can update, correct or delete it. The model is only asked to read what the search step found and to answer from it.
How RAG works, step by step
Every RAG system has two halves: getting documents ready, and answering questions.
- Ingest: documents are read, including scans, and turned into clean text.
- Chunk: the text is split into passages that are small enough to retrieve precisely and large enough to make sense on their own.
- Embed and index: each passage is converted into a vector, a list of numbers that captures its meaning, and stored in a vector database next to the original text.
- Retrieve: a question is converted the same way and the closest passages are found, usually combined with keyword search and re-ranked.
- Generate: the model receives the question and the retrieved passages and writes an answer, ideally with citations pointing back to the sources.
What RAG is good for
RAG suits any situation where the answer lives in documents that change or are private: customer support over product manuals, staff questions over HR policy, analysts searching filings, clinicians checking guidelines, lawyers searching contracts, citizens asking about a government scheme.
It also gives you something a plain model cannot: a source for every answer. A reader can open the cited passage and check it.
Where RAG goes wrong
Most RAG failures are retrieval failures, not model failures. If the right passage is never found, the model cannot use it.
- Poor reading of the source: a scanned page that returns no text, or a table flattened into a stream of numbers.
- Bad chunking: a clause or a list cut in half, so neither piece answers the question.
- Weak search: meaning-only search misses exact codes and names; keyword-only search misses different wording.
- No access control: the system retrieves a document the asker should never see.
- No way to measure: changes are made by feel and quality drifts.
Key takeaways
- RAG grounds an AI model in your documents at question time, without retraining it.
- The quality of a RAG answer is mostly decided before the model is called, by how documents were read, split and searched.
- Good RAG shows its sources, respects permissions, and can say "I do not know".
How troveGEN helps with RAG
troveGEN is the ingestion and search half of RAG as an API: it reads documents (including scans and web pages), structures and protects them, indexes them, and serves hybrid search with re-ranking, filters and permission checks. Cited answers are available when you want them, and the evaluation tools let you measure quality on your own documents.
What troveGEN provides
- Ingestion from uploads, web crawls and connectors, including scanned pages
- Hybrid search with reranking, filters and per-user permissions
- Optional cited answers that refuse when the documents do not cover the question
- Your own vector database, or fully managed hosting
- Built-in evaluation tools to measure quality on your own documents
Frequently asked questions
Is RAG the same as a chatbot?
No. A chatbot is an interface. RAG is the method that lets the chatbot answer from your documents. You can use RAG behind a chatbot, a search box, an internal tool or an AI agent.
Does RAG remove hallucinations?
It reduces them a great deal by giving the model real sources, but it does not remove them. Good systems also require citations, check that each claim is supported by the retrieved text, and refuse when the documents do not cover the question.
Do I need a vector database for RAG?
Almost always, because it is what makes finding passages by meaning fast. Some systems also keep the text in a regular database for keyword search, which is why many RAG systems use both.
How does troveGEN help with RAG?
troveGEN is RAG as an API. You send documents; it reads, structures, protects and indexes them, and your app gets accurate, permission-aware search with sources. It provides: Ingestion from uploads, web crawls and connectors, including scanned pages; Hybrid search with reranking, filters and per-user permissions; Optional cited answers that refuse when the documents do not cover the question; Your own vector database, or fully managed hosting; Built-in evaluation tools to measure quality on your own documents.