Managed RAG, or bring your own stack
Turn any document into accurate, private AI search.
Upload, crawl or sync PDFs, scans, spreadsheets and web pages. troveGEN reads them, protects sensitive data, and gives your app or AI agent the right passages, with sources, only for people allowed to see them — fully managed, or inside your own vector database.
What is troveGEN
Your documents, ready for AI search.
troveGEN turns your documents into something your apps and AI agents can search accurately. You send it files, web pages or feeds. It reads them, including scans, makes sense of their structure, protects sensitive details and indexes them. Your app then asks questions through one API and gets back the right passages, with their sources, from only what each person is allowed to see.
FOR
Teams building AI products
Add document search to your app without building the plumbing.
Reading scans, structuring documents, ranking results and enforcing permissions normally take months. Here it is one API, an SDK and an MCP server for your agents.
FOR
Regulated industries
Finance, healthcare, legal, hospitality and government.
Statements, clinical records, contracts, policies and circulars handled with the care each needs, with sensitive data protected and every answer traceable to its source.
FOR
Platform and IT teams
Keep control of your data.
Use your own vector database, models and keys, or run fully air-gapped. Permissions, audit and deletion that actually removes the data.
FOR
Everyone who needs answers
Ask your documents, get cited answers.
Invite colleagues to chat with the projects you choose. They get answers with sources, and see only what they are permitted to see.
How it works
Connect. Understand. Search.
Three steps, one API. Everything that normally takes a year of glue code is already between them.
1. Connect
UPLOAD · CRAWL · SYNC
Drop in files, post text, crawl a website together with its PDFs, or sync from S3, Azure Blob, Google Drive and any HTTP source. Scanned pages are read in the language you set.
2. Understand
PARSE · CHUNK · SECURE
Structure survives: headings, tables, clauses and policies stay intact. Sensitive values are protected before anything is embedded, then indexed in the vector database you choose.
3. Search
HYBRID · RERANK · CITE
One call combines keyword and meaning search, reranks, applies filters and permissions, and returns passages with sources. Add cited answers when you want them.
Ingestion that understands structure
A bank statement, understood.
Most pipelines see a wall of text. troveGEN sees a table, and every row becomes something you can search and add up.
rebuilt as a table · one chunk per row
Debits and credits are worked out for you. Anything uncertain is flagged, never guessed.
- Real structure, not blind slicing. Tables stay tables, headings give every chunk a breadcrumb, and code and formulas are never split.
- Scans and drawn text are read. Pages that are only pictures, or text drawn as vector shapes, are rendered and read with OCR instead of failing silently.
- Exact numbers. “Total spent with one vendor” is computed from the table, not guessed by a language model.
Search that earns trust
Right passage first. Wrong people never see it.
Hybrid retrieval, reranking and permissions run inside one query, so the answer is both relevant and allowed.
- Both ways of matching. Exact words catch codes and names; meaning catches different wording. They are fused, not chosen.
- Permissions inside the query. Access rules filter results before ranking, so a restricted document cannot leak through a clever question.
- Says “I don’t know”. Questions the documents do not cover are refused, and every citation is checked against a real passage.
Built for your industry
Tuned for the documents your industry runs on.
Statements, clinical records, contracts, policies and circulars each need different care. Anything else with words in it works too, text or scanned.
Finance
Statements you can search and add up. Filings that stay in one piece.
Learn more →Healthcare
Clinical lists come back complete, never a fragment.
Learn more →Legal
A clause stays with its sub-clauses. A judgment stays in context.
Learn more →Hospitality
Policies, rates and menus answered accurately, every time.
Learn more →Government
Every circular found by its reference number, date and department.
Learn more →Any other document
Text or scanned, structured or messy. If it holds words, troveGEN can search it.
Learn more →Web crawling
Crawl the web, PDFs included.
Circulars, filings, guidelines and judgments live in documents, not web pages. The crawler collects both, downloads the linked files and files each one with its details.
- Finds the content — beyond the home page, into the sections and documents that matter.
- Skips the noise — logins, images, tracking links and duplicates.
- Safe by default — stays on your hosts, respects robots.txt, and never fetches private addresses.
Measured, not claimed
Numbers you can check.
We publish no leaderboard score on purpose: a number on someone else’s documents says little about yours. The evaluation harness ships inside the product so you can measure your own.
Two ways to run it
Same engine. You choose where it runs.
Most platforms make this decision for you. troveGEN treats it as a setting, and you can move between the two.
troveGEN-Z · FULLY MANAGED
We host everything
No vector database to run, no model to deploy. Your documents get a dedicated schema and vector namespace, isolated from every other customer.
- Hosted vectors, embeddings, reranking and OCR included
- Nothing to operate, patch or scale
- From $25/month; most teams start here
BYO VECTOR DB · SOVEREIGN
You host the data
Point troveGEN at your own Qdrant, Pinecone, pgvector or Postgres. It writes into your database and never keeps a copy of your embeddings.
- Vectors never leave your infrastructure
- Your models are metered at zero, so you pay your provider
- Runs fully air-gapped; from $49/month
Capabilities
Everything a production retrieval stack needs.
Not a roadmap. Every card names something that ships today.
Any document in
INGESTION
PDF, DOCX, PPTX, XLSX, HTML, Markdown, EPUB, RTF, ODT, email (.eml) and crawled web pages. Docling for hard layouts, with a built-in parser fallback.
Structure-aware chunking
CHUNKING
Chunks carry a heading breadcrumb, page number and element kind. Tables stay whole or serialise per row; code fences never split; formulas are preserved verbatim.
Hybrid search + reranking
RETRIEVAL
Dense vectors and BM25 fused with RRF, diversified with MMR, then reranked by a cross-encoder — Cohere, a self-hosted TEI model, or an LLM.
Permission-aware retrieval
ACL
ACLs are pre-filtered inside the vector query, never post-filtered. Signed end-user assertions and IdP group closure mean the caller cannot escape their own scope.
PII/PHI vault
PRIVACY
Detected before embedding and replaced with reversible tokens; the mapping lives in a separate encrypted vault. Raw PII never reaches your vector store.
Prove it on your corpus
EVALUATION
Golden question sets, scored runs and A/B compare with hit@k, MRR, nDCG and context recall — so a config change is measured, not assumed.
See all 17 capabilities
OCR that refuses to guess
OCR
Scanned pages are read in the language you choose, including Hindi, Tamil, Arabic and other non-Latin scripts. With OCR off, a scan is rejected with a clear reason rather than indexed as an empty shell.
Tuned for your industry
VERTICALS
Finance, healthcare, legal, hospitality, government and education documents are handled the way that industry needs, so answers come back complete and correct.
Many languages
MULTILINGUAL
A multilingual embedder plus per-script keyword search. Ask in English or in the document's own language and retrieve from Hindi, Tamil, Arabic, Russian, Spanish and more. Verified end to end in 18 languages today; Chinese, Japanese, Korean and Thai are not supported yet.
Filters that hold on both halves
FILTERING
One Mongo-style filter language compiled per backend and enforced on the vector AND keyword halves. A filter can never widen what a caller sees.
Related-chunk expansion
RECALL
When one fact is split across chunks, return the neighbours, the whole section, the parent group, or chunks sharing an entity — measured by a context-recall metric.
Knowledge graph
GRAPHRAG
Entities, co-occurrence edges, typed relations and communities, with k-hop traversal at query time and a visual explorer in the console.
Point-in-time retrieval
VERSIONING
Re-ingest supersedes instead of overwriting, so you can ask what a document said last quarter. Deletion still purges the whole lineage.
Cited answers, optional
GENERATION
Every citation is validated against a real chunk or stripped. Out-of-context questions are refused, and each claim is scored for groundedness by a local NLI model.
Agentic retrieval
AGENTIC
A budgeted plan-execute-reflect loop that returns its full plan. Every hop runs through the same permission-scoped search — the planner cannot escape ACLs.
Continuous sync
CONNECTORS
S3, Azure Blob, Google Drive, any OAuth2 HTTP source, or a plain HTTP manifest for air-gapped deployments. Deletions at the source propagate.
Hosted, if you prefer
MANAGED
troveGEN-Z runs the vector store, embeddings, reranking and OCR for you, with a dedicated Postgres schema and vector namespace per tenant. Same API, same governance, nothing to operate.
For developers
One request. Everything above it handled.
A REST API, a typed TypeScript SDK and an MCP server, so your app or your AI agent gets grounded access to your documents in minutes.
Sovereign by architecture
Your data stays where you put it.
Choose hosted for speed, or keep everything inside your own walls. Either way, the same protections apply.
- Your vector database, models and LLM. Bring any of them, including a model on your own GPU.
- Credentials encrypted with AES-256-GCM and never returned by the API once stored.
- Sensitive data tokenised before it is embedded; the mapping lives in a separate encrypted vault.
- Deletion propagates to chunks, vectors in your store and vault entries, with a receipt of what was removed.
- Air-gapped is a supported mode, not a workaround.
The honest comparison
Where troveGEN stands.
Rows we would lose are included. Verify every one of them.
| Capability | troveGEN | Managed RAG platforms | Vector DB alone |
|---|---|---|---|
| Vectors CAN stay in your own database | ✓ | — | ✓ |
| Bring your own embedding model | ✓ | ~ | ✓ |
| Parsing, OCR and chunking included | ✓ | ✓ | — |
| Tuned for regulated industries | ✓ | — | — |
| Hybrid search + cross-encoder rerank | ✓ | ✓ | ~ |
| Permission-aware retrieval (pre-filtered) | ✓ | ~ | — |
| PII tokenised before embedding | ✓ | — | — |
| Point-in-time retrieval | ✓ | — | — |
| Knowledge graph + traversal | ✓ | ~ | — |
| Evaluation harness in the product | ✓ | ~ | — |
| Runs fully air-gapped | ✓ | — | ~ |
| Managed hosting if you want it | ✓ | ✓ | ✓ |
| Mature ecosystem and integrations | ~ | ✓ | ✓ |
| Third-party compliance certifications | — | ✓ | ✓ |
Questions we actually get
Straight answers.
What makes troveGEN different for finance, healthcare, legal, hospitality or government documents?
troveGEN recognises what kind of document it is reading and handles it accordingly, so statements can be searched and added up, medication lists and clauses come back complete, and policies and circulars are found with their details. The same tuning applies when it crawls a website and when it understands your search terms.
Can it crawl a website and its PDFs?
Yes. Point the crawler at a site and it collects the useful pages, skips the noise, downloads linked PDF, Word and Excel files and files each one with its details. Scanned PDFs found during a crawl are not read yet; scans you upload are read with OCR.
Where do my vectors actually live?
In your own database by default — Qdrant, Pinecone, pgvector or plain Postgres. troveGEN stores only bookkeeping: chunk text for the keyword half of hybrid search, extracted tables, and usage. If you would rather not run a vector database, the troveGEN-Z plans host it for you in a dedicated schema and namespace.
Should I use the managed tier or bring my own vector database?
Take troveGEN-Z if you would rather not operate a vector database — it is the faster start and what most teams choose; we host the vectors, embeddings, reranking and OCR, and your corpus gets a dedicated Postgres schema and vector namespace. Take BYO if your data cannot leave your infrastructure, you already run Qdrant or pgvector, or you want to pay your own model provider directly. The API, the console and the governance features are identical, so this is an infrastructure decision rather than a feature one, and it is not permanent — moving means re-ingesting into the new target.
Can it run without internet access?
Yes. With a local embedding model, a local reranker and a local LLM, no request leaves your network. The keyless local stack ships in the repository, and the HTTP manifest connector exists precisely so an air-gapped deployment can still sync from an internal system. Nothing is paywalled that an air-gapped install needs.
How is this different from a vector database?
A vector database stores and searches vectors. troveGEN is everything around it: parsing, OCR, structure-aware chunking, embedding, hybrid retrieval, reranking, ACL enforcement, evaluation — and it writes into the vector database you already chose rather than replacing it.
What does it cost, and what is a "page"?
Ingestion is billed per page, because a document is unbounded in size — a one-page note and a 400-page manual are both "one document" but differ by orders of magnitude in cost. A page is the parser's real page count for paginated formats, otherwise about 3,000 characters of extracted text. Plans start at $0 and run to $999/month, with a quoted Sovereign tier.
Do I pay for embeddings and reranking?
Only when troveGEN supplies the model. Bring your own OpenAI or Cohere key and you pay your provider directly — those calls are metered at zero here, which is the concrete value of bring-your-own rather than a slogan.
Which languages are supported?
Verified end to end today: Hindi, Marathi, Tamil, Telugu, Bengali, Urdu, Arabic, Hebrew, Russian, Greek, Spanish, French, German, Polish, Turkish, Indonesian, Vietnamese and English. Questions in the document's own language are answered accurately, and an English question can retrieve a passage in another language, most reliably with reranking switched on. The embedding model understands about 100 languages, but Chinese, Japanese, Korean and Thai documents are not accepted yet. To read scans in another language, set the OCR language for the project.
Stop building retrieval from scratch.
The free tier is a one-time budget of 500 pages and 2,000 searches — enough to load a real corpus and judge the quality properly. No card required.