About troveGEN
Your documents, ready for AI search.
troveGEN turns your documents into something your apps and AI agents can search accurately. You send it files, web pages or feeds. It reads them, including scans, makes sense of their structure, protects sensitive details and indexes them. Your app then asks questions through one API and gets back the right passages, with their sources, from only what each person is allowed to see.
What it does
From a pile of documents to answers you can trust.
Three things happen between your files and your users, and troveGEN does all of them.
Reads and organises
UPLOAD · CRAWL · SYNC
PDFs, spreadsheets, scans, emails and whole websites go in. Scanned pages are read, tables stay tables, and sensitive details are protected before anything is stored.
Finds the right passage
SEARCH · RANK · FILTER
Keyword and meaning search work together, results are ranked, and each person only sees what they are allowed to see. It works in 18 languages verified end to end, including Hindi, Tamil, Arabic, Russian and Spanish.
Answers with sources
CITED · MEASURED
Add cited answers when you want them, and measure quality on your own documents with the built-in evaluation tools, so changes are proven rather than assumed.
Who it is for
Built for teams that depend on their documents.
From a developer adding search to an app, to a regulated organisation that has to show where every answer came from.
FOR
Teams building AI products
Add document search to your app without building the plumbing.
Reading scans, structuring documents, ranking results and enforcing permissions normally take months. Here it is one API, an SDK and an MCP server for your agents.
FOR
Regulated industries
Finance, healthcare, legal, hospitality and government.
Statements, clinical records, contracts, policies and circulars handled with the care each needs, with sensitive data protected and every answer traceable to its source.
FOR
Platform and IT teams
Keep control of your data.
Use your own vector database, models and keys, or run fully air-gapped. Permissions, audit and deletion that actually removes the data.
FOR
Everyone who needs answers
Ask your documents, get cited answers.
Invite colleagues to chat with the projects you choose. They get answers with sources, and see only what they are permitted to see.
Where it comes from
Built because we needed it in production.
troveGEN is made by Intellara Technologies, an AI-native product studio in Bengaluru. It started as the retrieval subsystem behind a live product serving millions of users in twenty languages, then was extracted, generalised and pointed at whatever vector database you already run.
Why it exists
Retrieval quality is where RAG actually fails.
Not the LLM. Not the vector database. The unglamorous middle — parsing a scanned page correctly, keeping a table intact, not splitting a clause from its sub-clauses, knowing whether a 0.85 score means anything on your embedding model.
We hit these problems first
Running RAG over thirteen classical texts in twenty languages surfaced every one of them: OCR that silently produced nothing, chunks that lost the heading they belonged to, scores that were not comparable between models.
So the fixes are measured, not assumed
The evaluation harness exists because we needed to prove a chunking change helped. It ships in the product for the same reason — a benchmark on our corpus tells you nothing about yours.
And nothing is locked in
We wanted to swap embedding models without re-platforming, and to keep customer data in customer infrastructure. That is why every boundary in troveGEN is a connection you own.
The company
Intellara Technologies
- AI-native product studio building domain-intelligent systems — RAG pipelines, LLM platforms and knowledge-intensive products.
- Bengaluru, Karnataka, India. Founded 2025. DPIIT-recognised startup.
- AstroReeti — an AI-native Vedic astrology platform, live in production, and where most of troveGEN’s hard lessons were learned.
- troveGEN — the retrieval stack, made general and sold on its own.
How we write about the product
Every claim maps to something that ships.
This site carries no roadmap items written in the present tense. Where something is partial or missing — third-party certifications, for instance — the comparison table and the sovereignty page say so directly. If you find a claim that does not hold, tell us and we will correct it.
Talk to us.
Questions about a sovereign deployment, data residency, or whether troveGEN fits your corpus — we would rather have the conversation than have you guess.