Learn / Security and privacy

What is permission-aware retrieval?

Updated 3 October 2026 · 2 min read

Short answer

Permission-aware retrieval means the search itself enforces who may see what, so a question can only retrieve passages the asker is entitled to. It is applied before ranking, inside the query, rather than as a clean-up step after results come back.

With troveGEN

troveGEN pre-filters every search by who is asking, so a question can only retrieve passages that person is allowed to see.

See what troveGEN provides ↓

Why it is needed

A retrieval system does not understand confidentiality. If every document is in one index, a question from an intern can pull in a board paper, and the model will happily summarise it. Source systems such as drives and wikis have permissions; a copy in the index does not inherit them unless you make it.

Pre-filtering versus post-filtering

Post-filtering retrieves the top results and then removes the forbidden ones. It can leave the user with too few results, and any mistake leaks data. Pre-filtering puts the permission condition into the search itself, so forbidden passages are never candidates. It is the safer and more accurate design.

How to model permissions

  • Attach an access list to every passage, copied from the source document.
  • Resolve the asker to a set of identities: the user and the groups they belong to, including nested groups.
  • Allow a passage only when the two sets overlap, and treat a missing list as "public only" under enforcement (fail closed).
  • Compose filters with AND so a user-supplied filter can narrow results but never widen what they can see.

Common mistakes

Trusting a user identity sent by the client without a signature; forgetting that neighbouring passages or graph links may cross a boundary; applying access checks to search but not to answer generation or caching; and not updating the index when permissions change at the source.

Key takeaways

  • Enforce permissions inside the query, before ranking.
  • Fail closed: no identity means public documents only.
  • Cover everything that can surface text: expansion, graphs, answers and caches.

How troveGEN helps with access control for RAG

troveGEN copies each document's access list onto its passages, pre-filters both vector and keyword retrieval with the caller's identities, supports signed end-user assertions and group inheritance from your directory, and fails closed. Related-passage expansion and graph links are re-checked against the same permissions. Team roles control who can use which projects in the console.

What troveGEN provides

  • Each document's access list copied onto its passages
  • Filtering inside the vector and the keyword search, before ranking
  • Signed end-user assertions and group inheritance from your directory
  • Fail-closed behaviour: no identity means public documents only
  • Team roles for who can use which projects in the console

See access control Start free — 500 pages

Frequently asked questions

Can I use my existing groups?

Yes. Sync your group structure and let nested groups inherit, so you do not maintain a second set of permissions by hand.

What happens when a document's permissions change?

The access list on its passages should be updated promptly. Re-syncing the source updates it.

Is a filter in the prompt enough?

No. Telling the model not to reveal something is not access control. The forbidden text must never be retrieved.

How does troveGEN help with access control for RAG?

troveGEN pre-filters every search by who is asking, so a question can only retrieve passages that person is allowed to see. It provides: Each document's access list copied onto its passages; Filtering inside the vector and the keyword search, before ranking; Signed end-user assertions and group inheritance from your directory; Fail-closed behaviour: no identity means public documents only; Team roles for who can use which projects in the console.

Keep reading