Managed RAG, or bring your own stack

Turn any document into accurate, private AI search.

Upload, crawl or sync PDFs, scans, spreadsheets and web pages. troveGEN reads them, protects sensitive data, and gives your app or AI agent the right passages, with sources, only for people allowed to see them — fully managed, or inside your own vector database.

18 languages verified6 industries coveredAir-gapped capableYour vectors, or ours
From documents to cited answersFour kinds of source documents flow into troveGEN, which turns them into structured chunks for finance, hospitality, government and healthcare. A question about payments to one vendor then returns the matching rows with citations.Statement.pdftext drawn as shapesRate card.xlsxtables, 3 sheetsNotice scan.jpgread with OCRWebsitecrawled with its PDFsparsechunksecureTransaction rows24 rows, each one searchablefinanceCancellation policycomplete, never cut offhospitalityEligibilityscheme rules, completegovernmentMedicationsfull list, with doseshealthcareshow all payments to Demo Software3 PAYMENTS FOUND · CITED04 AugDEMO SOFTWARE SUBSCRIPTION−$100.00[1]04 SepDEMO SOFTWARE USAGE−$300.00[2]04 OctDEMO SOFTWARE SUBSCRIPTION−$100.00[3]
Documents of every kind go in. Structured, domain-aware chunks come out. Ask a question and get the exact rows back, with citations.

What is troveGEN

Your documents, ready for AI search.

troveGEN turns your documents into something your apps and AI agents can search accurately. You send it files, web pages or feeds. It reads them, including scans, makes sense of their structure, protects sensitive details and indexes them. Your app then asks questions through one API and gets back the right passages, with their sources, from only what each person is allowed to see.

FOR

Teams building AI products

Add document search to your app without building the plumbing.

Reading scans, structuring documents, ranking results and enforcing permissions normally take months. Here it is one API, an SDK and an MCP server for your agents.

FOR

Regulated industries

Finance, healthcare, legal, hospitality and government.

Statements, clinical records, contracts, policies and circulars handled with the care each needs, with sensitive data protected and every answer traceable to its source.

FOR

Platform and IT teams

Keep control of your data.

Use your own vector database, models and keys, or run fully air-gapped. Permissions, audit and deletion that actually removes the data.

FOR

Everyone who needs answers

Ask your documents, get cited answers.

Invite colleagues to chat with the projects you choose. They get answers with sources, and see only what they are permitted to see.

More about troveGEN →

How it works

Connect. Understand. Search.

Three steps, one API. Everything that normally takes a year of glue code is already between them.

1. Connect

UPLOAD · CRAWL · SYNC

Drop in files, post text, crawl a website together with its PDFs, or sync from S3, Azure Blob, Google Drive and any HTTP source. Scanned pages are read in the language you set.

2. Understand

PARSE · CHUNK · SECURE

Structure survives: headings, tables, clauses and policies stay intact. Sensitive values are protected before anything is embedded, then indexed in the vector database you choose.

3. Search

HYBRID · RERANK · CITE

One call combines keyword and meaning search, reranks, applies filters and permissions, and returns passages with sources. Add cited answers when you want them.

Ingestion that understands structure

A bank statement, understood.

Most pipelines see a wall of text. troveGEN sees a table, and every row becomes something you can search and add up.

rebuilt as a table · one chunk per row

DateDescriptionDebitCreditBalance
01 AugSAMPLE CAFE250.0050,000.00
04 AugDEMO SOFTWARE SUBSCRIPTION100.0049,900.00
05 AugPAYROLL SAMPLE LTD10,000.0059,900.00
09 AugSAMPLE CLOUD HOSTING1,200.0058,700.00
12 AugDEMO SOFTWARE USAGE300.0058,400.00
15 AugSAMPLE UTILITY CO900.0057,500.00

Debits and credits are worked out for you. Anything uncertain is flagged, never guessed.

A statement whose text has no columns becomes a real table. Every row is its own searchable chunk, and the amounts can be summed exactly.
  • Real structure, not blind slicing. Tables stay tables, headings give every chunk a breadcrumb, and code and formulas are never split.
  • Scans and drawn text are read. Pages that are only pictures, or text drawn as vector shapes, are rendered and read with OCR instead of failing silently.
  • Exact numbers. “Total spent with one vendor” is computed from the table, not guessed by a language model.

Search that earns trust

Right passage first. Wrong people never see it.

Hybrid retrieval, reranking and permissions run inside one query, so the answer is both relevant and allowed.

How a search is answeredA question is matched by keywords and by meaning, the two result lists are fused, a reranker puts the cancellation policy on top, a permission check removes a restricted document, and the answer cites its source.QUESTIONrefund if I cancel late?Keywordexact wordsMeaningvectors, any wordingFUSEmergeRooms · DeluxeRs. 8,500 per nightDining · MenuStarters, mains, dessertsCancellation policyFree up to 48 hours before arrival#1Staff handbookNot permitted for this userRERANKED · PERMISSION-CHECKEDCITED ANSWERCancellation is free up to 48 hours before arrival; after that the first night is charged. [1]
Keyword and meaning search run side by side, their results are fused, a reranker puts the best passage first, and permissions are enforced inside the query before an answer is written with its source.
  • Both ways of matching. Exact words catch codes and names; meaning catches different wording. They are fused, not chosen.
  • Permissions inside the query. Access rules filter results before ranking, so a restricted document cannot leak through a clever question.
  • Says “I don’t know”. Questions the documents do not cover are refused, and every citation is checked against a real passage.

Web crawling

Crawl the web, PDFs included.

Circulars, filings, guidelines and judgments live in documents, not web pages. The crawler collects both, downloads the linked files and files each one with its details.

A website crawled into documentsA crawler visits a website, collects its useful pages, skips the login page and images, downloads the linked PDFs and files each one with its details.example.organy website/documents/notices/reports/login/logo.pngPDFPDFFILED WITH THEIR DETAILSCircularRef. SAMPLE/2026/001 · 12 Mar 2026PDFAnnual reportFY 2025 · 84 pagesPDFPolicy pageText and tables, with its addressPAGENoise skipped; PDFs collected.
The crawler visits a site, skips what is not content, downloads the linked PDFs and files every document with its details.
  • Finds the content — beyond the home page, into the sections and documents that matter.
  • Skips the noise — logins, images, tracking links and duplicates.
  • Safe by default — stays on your hosts, respects robots.txt, and never fetches private addresses.

Measured, not claimed

Numbers you can check.

18languages verified end to end, with more to come
4vector databases supported: Qdrant, Pinecone, pgvector, Postgres
400+automated end-to-end checks across the product
0raw sensitive values in your vector store

We publish no leaderboard score on purpose: a number on someone else’s documents says little about yours. The evaluation harness ships inside the product so you can measure your own.

Two ways to run it

Same engine. You choose where it runs.

Most platforms make this decision for you. troveGEN treats it as a setting, and you can move between the two.

troveGEN-Z · FULLY MANAGED

We host everything

No vector database to run, no model to deploy. Your documents get a dedicated schema and vector namespace, isolated from every other customer.

  • Hosted vectors, embeddings, reranking and OCR included
  • Nothing to operate, patch or scale
  • From $25/month; most teams start here

See troveGEN-Z plans →

BYO VECTOR DB · SOVEREIGN

You host the data

Point troveGEN at your own Qdrant, Pinecone, pgvector or Postgres. It writes into your database and never keeps a copy of your embeddings.

  • Vectors never leave your infrastructure
  • Your models are metered at zero, so you pay your provider
  • Runs fully air-gapped; from $49/month

How sovereignty is enforced →

Capabilities

Everything a production retrieval stack needs.

Not a roadmap. Every card names something that ships today.

Any document in

INGESTION

PDF, DOCX, PPTX, XLSX, HTML, Markdown, EPUB, RTF, ODT, email (.eml) and crawled web pages. Docling for hard layouts, with a built-in parser fallback.

Structure-aware chunking

CHUNKING

Chunks carry a heading breadcrumb, page number and element kind. Tables stay whole or serialise per row; code fences never split; formulas are preserved verbatim.

Hybrid search + reranking

RETRIEVAL

Dense vectors and BM25 fused with RRF, diversified with MMR, then reranked by a cross-encoder — Cohere, a self-hosted TEI model, or an LLM.

Permission-aware retrieval

ACL

ACLs are pre-filtered inside the vector query, never post-filtered. Signed end-user assertions and IdP group closure mean the caller cannot escape their own scope.

PII/PHI vault

PRIVACY

Detected before embedding and replaced with reversible tokens; the mapping lives in a separate encrypted vault. Raw PII never reaches your vector store.

Prove it on your corpus

EVALUATION

Golden question sets, scored runs and A/B compare with hit@k, MRR, nDCG and context recall — so a config change is measured, not assumed.

See all 17 capabilities

OCR that refuses to guess

OCR

Scanned pages are read in the language you choose, including Hindi, Tamil, Arabic and other non-Latin scripts. With OCR off, a scan is rejected with a clear reason rather than indexed as an empty shell.

Tuned for your industry

VERTICALS

Finance, healthcare, legal, hospitality, government and education documents are handled the way that industry needs, so answers come back complete and correct.

Many languages

MULTILINGUAL

A multilingual embedder plus per-script keyword search. Ask in English or in the document's own language and retrieve from Hindi, Tamil, Arabic, Russian, Spanish and more. Verified end to end in 18 languages today; Chinese, Japanese, Korean and Thai are not supported yet.

Filters that hold on both halves

FILTERING

One Mongo-style filter language compiled per backend and enforced on the vector AND keyword halves. A filter can never widen what a caller sees.

Related-chunk expansion

RECALL

When one fact is split across chunks, return the neighbours, the whole section, the parent group, or chunks sharing an entity — measured by a context-recall metric.

Knowledge graph

GRAPHRAG

Entities, co-occurrence edges, typed relations and communities, with k-hop traversal at query time and a visual explorer in the console.

Point-in-time retrieval

VERSIONING

Re-ingest supersedes instead of overwriting, so you can ask what a document said last quarter. Deletion still purges the whole lineage.

Cited answers, optional

GENERATION

Every citation is validated against a real chunk or stripped. Out-of-context questions are refused, and each claim is scored for groundedness by a local NLI model.

Agentic retrieval

AGENTIC

A budgeted plan-execute-reflect loop that returns its full plan. Every hop runs through the same permission-scoped search — the planner cannot escape ACLs.

Continuous sync

CONNECTORS

S3, Azure Blob, Google Drive, any OAuth2 HTTP source, or a plain HTTP manifest for air-gapped deployments. Deletions at the source propagate.

Hosted, if you prefer

MANAGED

troveGEN-Z runs the vector store, embeddings, reranking and OCR for you, with a dedicated Postgres schema and vector namespace per tenant. Same API, same governance, nothing to operate.

# Search every document you are allowed to see curl -X POST https://trovegen.com/api/v1/projects/$PROJECT/search \ -H "Authorization: Bearer trv_..." \ -H "Content-Type: application/json" \ -d '{"query":"cancellation policy for late bookings","mode":"hybrid","topK":5}'

For developers

One request. Everything above it handled.

A REST API, a typed TypeScript SDK and an MCP server, so your app or your AI agent gets grounded access to your documents in minutes.

Read the API →

Sovereign by architecture

Your data stays where you put it.

Choose hosted for speed, or keep everything inside your own walls. Either way, the same protections apply.

  • Your vector database, models and LLM. Bring any of them, including a model on your own GPU.
  • Credentials encrypted with AES-256-GCM and never returned by the API once stored.
  • Sensitive data tokenised before it is embedded; the mapping lives in a separate encrypted vault.
  • Deletion propagates to chunks, vectors in your store and vault entries, with a receipt of what was removed.
  • Air-gapped is a supported mode, not a workaround.

How sovereignty is enforced →

The honest comparison

Where troveGEN stands.

Rows we would lose are included. Verify every one of them.

CapabilitytroveGENManaged RAG platformsVector DB alone
Vectors CAN stay in your own database✓—✓
Bring your own embedding model✓~✓
Parsing, OCR and chunking included✓✓—
Tuned for regulated industries✓——
Hybrid search + cross-encoder rerank✓✓~
Permission-aware retrieval (pre-filtered)✓~—
PII tokenised before embedding✓——
Point-in-time retrieval✓——
Knowledge graph + traversal✓~—
Evaluation harness in the product✓~—
Runs fully air-gapped✓—~
Managed hosting if you want it✓✓✓
Mature ecosystem and integrations~✓✓
Third-party compliance certifications—✓✓

Questions we actually get

Straight answers.

What makes troveGEN different for finance, healthcare, legal, hospitality or government documents?

troveGEN recognises what kind of document it is reading and handles it accordingly, so statements can be searched and added up, medication lists and clauses come back complete, and policies and circulars are found with their details. The same tuning applies when it crawls a website and when it understands your search terms.

Can it crawl a website and its PDFs?

Yes. Point the crawler at a site and it collects the useful pages, skips the noise, downloads linked PDF, Word and Excel files and files each one with its details. Scanned PDFs found during a crawl are not read yet; scans you upload are read with OCR.

Where do my vectors actually live?

In your own database by default — Qdrant, Pinecone, pgvector or plain Postgres. troveGEN stores only bookkeeping: chunk text for the keyword half of hybrid search, extracted tables, and usage. If you would rather not run a vector database, the troveGEN-Z plans host it for you in a dedicated schema and namespace.

Should I use the managed tier or bring my own vector database?

Take troveGEN-Z if you would rather not operate a vector database — it is the faster start and what most teams choose; we host the vectors, embeddings, reranking and OCR, and your corpus gets a dedicated Postgres schema and vector namespace. Take BYO if your data cannot leave your infrastructure, you already run Qdrant or pgvector, or you want to pay your own model provider directly. The API, the console and the governance features are identical, so this is an infrastructure decision rather than a feature one, and it is not permanent — moving means re-ingesting into the new target.

Can it run without internet access?

Yes. With a local embedding model, a local reranker and a local LLM, no request leaves your network. The keyless local stack ships in the repository, and the HTTP manifest connector exists precisely so an air-gapped deployment can still sync from an internal system. Nothing is paywalled that an air-gapped install needs.

How is this different from a vector database?

A vector database stores and searches vectors. troveGEN is everything around it: parsing, OCR, structure-aware chunking, embedding, hybrid retrieval, reranking, ACL enforcement, evaluation — and it writes into the vector database you already chose rather than replacing it.

What does it cost, and what is a "page"?

Ingestion is billed per page, because a document is unbounded in size — a one-page note and a 400-page manual are both "one document" but differ by orders of magnitude in cost. A page is the parser's real page count for paginated formats, otherwise about 3,000 characters of extracted text. Plans start at $0 and run to $999/month, with a quoted Sovereign tier.

Do I pay for embeddings and reranking?

Only when troveGEN supplies the model. Bring your own OpenAI or Cohere key and you pay your provider directly — those calls are metered at zero here, which is the concrete value of bring-your-own rather than a slogan.

Which languages are supported?

Verified end to end today: Hindi, Marathi, Tamil, Telugu, Bengali, Urdu, Arabic, Hebrew, Russian, Greek, Spanish, French, German, Polish, Turkish, Indonesian, Vietnamese and English. Questions in the document's own language are answered accurately, and an English question can retrieve a passage in another language, most reliably with reranking switched on. The embedding model understands about 100 languages, but Chinese, Japanese, Korean and Thai documents are not accepted yet. To read scans in another language, set the OCR language for the project.

Stop building retrieval from scratch.

The free tier is a one-time budget of 500 pages and 2,000 searches — enough to load a real corpus and judge the quality properly. No card required.