Learn / Basics

What is a vector database, and what are embeddings?

Updated 3 October 2026 · 2 min read

Short answer

An embedding is a list of numbers that represents the meaning of a piece of content, and a vector database is a store built to find the embeddings closest to a query very quickly. Together they let software search by meaning instead of by exact words.

With troveGEN

troveGEN works with the vector database you already run, or hosts one for you, and with the embedding model of your choice, including a multilingual one.

See what troveGEN provides ↓

What an embedding is

An embedding model reads text and outputs a vector, typically hundreds to a few thousand numbers. The model is trained so that texts with similar meaning end up as nearby points. "How do I cancel my booking" and "refund policy for reservations" land close together even though they share few words.

What a vector database does

Finding the nearest vectors among millions by comparing against every one is too slow, so vector databases build special indexes that find close neighbours approximately but fast. They store the vector together with the original text and metadata such as the source, date and permissions, and they let you filter on that metadata while searching.

Well-known options include Qdrant, Pinecone and the pgvector extension for PostgreSQL. Some are standalone services and some are add-ons to databases you already run.

How similarity is measured

The usual measure is cosine similarity, which compares the direction of two vectors. A score near one means very similar, near zero means unrelated. Scores are only comparable within one embedding model, which is why switching models means re-embedding your documents.

What vectors cannot do alone

Meaning search is strong on paraphrase and weak on exact strings. A part number, a drug code or a person's name may not be well represented, so a question containing one can retrieve something merely similar. That is why serious systems combine vector search with keyword search, and why filters and permissions are applied alongside similarity rather than after it.

Choosing an embedding model

Consider the languages you need, the maximum text length, the vector size (larger costs more to store and search) and whether it can run where your data must stay. A multilingual model lets a question in one language find a document in another.

Key takeaways

  • Embeddings encode meaning as numbers; a vector database finds the closest ones quickly.
  • Similarity scores are only meaningful within one embedding model.
  • Combine vector search with keyword search and filters for reliable results.

How troveGEN helps with vector databases and embeddings

troveGEN writes into the vector database you already run, including Qdrant, Pinecone, pgvector and plain PostgreSQL, so your vectors stay under your control. If you would rather not run one, the managed option hosts it for you with isolation per customer. You can bring your own embedding model or use ours, including a multilingual model.

What troveGEN provides

  • Connections to Qdrant, Pinecone, pgvector and PostgreSQL, so your vectors stay yours
  • Managed hosting with separate storage for each customer if you prefer not to run one
  • A multilingual embedding model included, or bring OpenAI, Azure, Pinecone or any compatible endpoint
  • Keyword and vector search together, with filters applied on both
  • Credentials stored encrypted and never returned by the API

See how your data stays yours Start free — 500 pages

Frequently asked questions

Is a vector database the same as a normal database?

No. A normal database finds rows that match exact conditions. A vector database finds items that are closest in meaning to a query. Many teams use both, and some databases offer both.

Can I change the embedding model later?

Yes, but vectors from different models are not comparable, so the documents must be re-embedded and re-indexed with the new model.

Do embeddings contain my text?

An embedding is numbers, not text, but it can leak information about the text, so treat vectors as sensitive and protect sensitive values before embedding.

How does troveGEN help with vector databases and embeddings?

troveGEN works with the vector database you already run, or hosts one for you, and with the embedding model of your choice, including a multilingual one. It provides: Connections to Qdrant, Pinecone, pgvector and PostgreSQL, so your vectors stay yours; Managed hosting with separate storage for each customer if you prefer not to run one; A multilingual embedding model included, or bring OpenAI, Azure, Pinecone or any compatible endpoint; Keyword and vector search together, with filters applied on both; Credentials stored encrypted and never returned by the API.

Keep reading