What is a vector database, and when do you actually need one?

A vector database stores embeddings, the numerical fingerprints AI models make for text and images, and finds the ones closest in meaning to a query. That nearest-neighbour search is what powers semantic search and RAG. Here is what it does, and when a library or a Postgres extension will do the same job.

By Yash Malviya

Published

Detailed image of a server rack with glowing lights in a modern data center
Photo: panumas nikhomkhai / Pexels

What a vector database actually stores

A vector database is a system built to store and search one specific thing: long lists of numbers called vectors. Those vectors are almost always embeddings, the numerical fingerprints a model produces for a piece of text, an image or a sound. IBM's definition is about as plain as it gets: "A vector database stores, manages and indexes high-dimensional vector data." Where a normal database is good at exact matches, find the row where the email equals this, a vector database is built for a fuzzier question: find the items whose meaning is closest to this one.

That difference is the whole point. Traditional search matches keywords. Vector search matches meaning. If you embed the sentence "how do I reset my password" and compare it against a library of embedded help articles, the closest vectors will include a document titled "recovering your account," even though the two share no keywords. The database is comparing positions in space, not strings of text.

Similarity search and nearest neighbours

The core operation is nearest-neighbour search. Every stored item is a point in a high-dimensional space, often hundreds or thousands of dimensions wide. When a query arrives, it is embedded into that same space, and the database returns the points closest to it, ranked by a distance measure such as cosine similarity. IBM describes the mechanic directly: the database "compares the query vector against indexed vectors and calculates similarity scores to identify the nearest neighbors."

Doing this exactly, by measuring the distance to every stored vector, is simple but slow once you have millions of them. So real systems use approximate nearest neighbour, or ANN, search, which trades a sliver of accuracy for a large gain in speed. The most widely used method is HNSW, from a 2016 paper by Yu Malkov and Dmitry Yashunin. It builds a layered graph you can walk quickly, and the authors report that starting the search at the top layer "boosts the performance compared to NSW and allows a logarithmic complexity scaling." Logarithmic is the word that matters, because it is why a good vector index can answer a query over millions of items in milliseconds. AWS notes that production systems lean on exactly these algorithms, using "k-nearest neighbor (k-NN) indexes powered by algorithms like HNSW and IVF."

“Vector databases typically manage large collections of embedding vectors.”

Douze et al., The Faiss library, arXiv 2024
Networking equipment with connected cables, showcasing modern technology infrastructure
Approximate nearest neighbour methods such as HNSW build a navigable graph, which is why a vector index can search millions of items in milliseconds. Photo: Vladimir Srajber / Pexels

Why it underpins RAG and semantic search

This is the machinery behind two things you have probably used. Semantic search is the obvious one: results ranked by meaning rather than keyword overlap. The other is retrieval-augmented generation, or RAG, the standard way to give a language model access to facts it was not trained on. In a RAG system your documents are chunked, embedded and stored in a vector database. When a user asks a question, the question is embedded, the database returns the most similar chunks, and those chunks are pasted into the model's prompt as context. The Faiss library paper states the relationship plainly: "Vector databases typically manage large collections of embedding vectors." The vector database is the retrieval half of retrieval-augmented generation.

Because embeddings collapse meaning into distance, the same index also powers recommendation, the "items like this one" list, along with deduplication, clustering and multimodal search, where a text query can retrieve an image because both live in the same space. It is one primitive, nearest-neighbour search over embeddings, wearing a lot of different product names.

When you actually need one, and when you do not

Here is the part the marketing tends to skip. A dedicated vector database is not a requirement for doing this work. If you have a few thousand or even a few hundred thousand vectors, you can hold them in memory and search them with a library such as Faiss, the open-source toolkit Meta built for exactly this, or with a vector extension bolted onto a database you already run, such as pgvector for Postgres. No new service, no new bill, no new thing to operate.

You start needing a purpose-built vector database when scale and operations catch up with you: tens of millions of vectors and beyond, high query volume, frequent updates, filtering by metadata alongside the vector search, and the ordinary production demands of replication, persistence and access control. Those are real problems, and managed vector databases solve them well. But they are infrastructure problems, not intelligence problems. Buying one does not make your retrieval smarter. The quality of your results is set upstream, by the embedding model you choose and by how you chunk your documents.

The honest summary

The skeptic's read is short. A vector database is a specialised index for one operation, finding the nearest vectors to a query vector, and it is genuinely useful because so many AI features reduce to that single operation. It is also easy to over-buy. The intelligence lives in the embeddings, not in the database that stores them, and for smaller projects a library or a Postgres extension will do the same job without a separate system to run. Reach for a dedicated vector database when your scale, update rate or operational needs demand it, not because a diagram told you that every AI app needs one.

Frequently asked questions

What is a vector database in simple terms?

It is a database built to store embeddings, which are high-dimensional lists of numbers that represent the meaning of text, images or audio, and to find the ones closest to a query. IBM describes it as a system that stores, manages and indexes high-dimensional vector data. Instead of matching keywords, it ranks results by how near their vectors are to yours.

How does similarity search work in a vector database?

Your query is turned into a vector in the same space as the stored items, and the database returns the nearest points, ranked by a distance measure such as cosine similarity. Checking every vector is slow at scale, so most systems use approximate nearest neighbour methods like HNSW to find close matches in milliseconds while trading a little accuracy for speed.

Why do vector databases matter for RAG?

Retrieval-augmented generation feeds a language model relevant context at query time. A vector database is the retrieval part: your documents are chunked, embedded and stored, and when a question comes in the database returns the most similar chunks, which are pasted into the prompt. That is why the same store powers both semantic search and RAG.

Do I always need a dedicated vector database?

No. For a few thousand to a few hundred thousand vectors you can search in memory with a library such as Faiss, or add a vector extension like pgvector to a database you already run. A purpose-built vector database earns its place at larger scale, with high query volume, frequent updates and metadata filtering alongside the vector search.

Is the vector database what makes my AI results good?

No. The database is an index for one operation, finding the nearest vectors. The quality of your results is set upstream, by the embedding model you choose and how you chunk your documents. Buying a managed vector database solves infrastructure problems like scale and reliability, but it does not make retrieval smarter.

Sources

What each one is, and whose it is.

  1. PaperIndependent of the vendor
  2. 2

    The Faiss library, Douze et al., arXiv (January 16, 2024)

    PaperThe vendor’s own
  3. DocumentationIndependent of the vendor
  4. 4

    What is a Vector Database?, Amazon Web Services

    Documentation