A vector database stores numerical representations of data, called embeddings, and finds the stored items whose representations sit closest to the representation of your query. That is the whole trick: instead of matching the exact words you typed, the system ranks records by how near they are in a mathematical space. Closeness is a useful ranking signal, but it is not proof that a result is relevant, and most of the practical work lies in handling that gap.
The core intuition: search by nearness
Traditional databases answer questions like “find rows where category equals shoes and price is under 50.” Vector databases answer a different question: “which stored items are most similar to this one?” Similarity is expressed as position. Each item is turned into a list of numbers, a vector, and items with similar meaning or features end up near each other in that space.
As an Amazon Associate I earn from qualifying purchases.
A simple way to picture it: imagine every document plotted as a point on a map. A query about “trail running shoes for wet ground” lands somewhere on that map, and the nearest points are documents about waterproof trail footwear, even if they never use the phrase “wet ground.” The map has hundreds or thousands of dimensions rather than two, but the principle is the same.
Quick wins for a faster PC:
Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Repair Windows errors before they cause bigger problemsFix Now →How the data gets in: embeddings
The vectors are produced by an embedding model, a separate component that reads a piece of content and outputs an array of numbers. The vector database does not understand the content on its own; it stores and searches what the model produced. A typical ingestion flow looks like this:
#1 Best Overall
- Split the source material into units that suit the model, such as paragraphs, product records, or images.
- Pass each unit through an embedding model to produce a vector.
- Store the vector in the database together with a reference back to the original content, such as a document ID or URL.
- Attach metadata where useful, such as document type, publication date, language, category, or access permissions.
Pinecone describes this sequence of converting content to vectors, storing them, and querying them. The reference step matters because embeddings are not the original content. A search result is a pointer to a record, and the application uses that pointer to fetch the text, image, or row it actually needs to show.
How a query runs
When a user asks something, the application runs the same kind of pipeline on the query:
- Convert the query into a vector with a compatible embedding model.
- Apply any metadata filters, such as “only documents from 2025 or later” or “only records this user may read.”
- Compare the query vector with the stored vectors using a distance or similarity measure. Common choices include cosine similarity and Euclidean distance; the right one depends on how the embeddings were trained.
- Return the nearest records, usually ranked and limited to a set number such as the top 10.
- Use the results directly, combine them with other retrieval methods, or pass them to a generative model as context.
The last step is where many modern applications use vector search: retrieval-augmented generation, where a language model answers using the retrieved material. The database supplies candidates; the rest of the application decides what to do with them.
Recommended Free Tools
Why compatible embeddings matter
A query vector can only be compared meaningfully with stored vectors produced by the same embedding model, or at least a model that places content in the same space. Vectors from different models are not interchangeable, even when they have the same length. Weaviate documents this at the collection level: changing the configured vectorizer for a collection requires creating a new collection and migrating the data, and using vectors from a different model risks incompatibility. Plan model changes as data migrations, not configuration tweaks.
Indexes: how the search stays fast
Comparing a query against every stored vector is exact but expensive at scale. Vector databases therefore use indexes to narrow the work. Two broad approaches matter.
Exact nearest-neighbor search
Exact search compares the query with every candidate and returns the true nearest results. pgvector performs exact nearest-neighbor search by default, which gives perfect recall for that search. It is the simplest correct baseline, and for small collections it may be fast enough. As collections grow, the cost grows with them.
Approximate nearest-neighbor search
Approximate nearest-neighbor (ANN) techniques reduce the number of comparisons by organizing vectors so the search only examines a promising subset. Google Cloud notes that this can reduce computation but can trade some recall, meaning some true nearest neighbors may be missed, for speed. Milvus explains that the index type affects throughput, memory use, and search correctness, so the index is a tuning decision rather than a background detail.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Index types compared in pgvector
pgvector documents two approximate index types, HNSW and IVFFlat. Its comparison states that HNSW offers a better speed-recall trade-off than IVFFlat, while HNSW builds more slowly and uses more memory. This is guidance from pgvector’s own documentation for its implementation, not a universal benchmark across products or workloads.
Rank #3
| Approach | Recall | Query speed | Build time and memory |
|---|---|---|---|
| Exact search (pgvector default) | Perfect recall for that search | Compares against every stored vector; cost grows with collection size | No index to build |
| HNSW (pgvector) | Better speed-recall trade-off than IVFFlat in pgvector’s comparison | Faster than exact search at scale; see trade-off above | Slower build and more memory than IVFFlat, per pgvector’s documentation |
| IVFFlat (pgvector) | Lower speed-recall trade-off than HNSW in pgvector’s comparison | Approximate; exact figures not stated in the cited documentation | Faster build and less memory than HNSW, per pgvector’s documentation |
Index behavior varies across products. Measure recall and latency on your own data before choosing settings.
Keyword search and hybrid search
Vector search is good at matching meaning across different wording, but it can be weaker at exact terms. Product codes, personal names, error identifiers, and quoted phrases often need literal matching. Keyword search preserves that precision.
Hybrid search combines the two, and Weaviate documents it as doing exactly that. The practical rule: if your queries include names, identifiers, or exact phrases, test whether keyword or hybrid retrieval improves results before assuming vectors alone are enough.
Closeness is a ranking signal, not proof of relevance
This is the point most often lost in explanations. A vector search always returns something: the nearest records, even if none of them answers the question. Weaviate’s search documentation notes that even a nearest-neighbor result can be a bad match. The score tells you how close the representation is, according to the model, not whether the record is correct, current, authorized, or useful.
Several failure patterns follow from this:
- Plausible but wrong: a record on a related topic outranks the one the user needs because the model treats the two as similar.
- Missing exact terms: a query for a specific model number returns general product pages without the matching part.
- Stale results: outdated documents are close in meaning but superseded by newer ones, unless metadata filters or freshness rules exclude them.
- Overconfident generation: in RAG, a language model may present weakly relevant context as a firm answer. Retrieval grounds an answer in material, but does not guarantee that the answer is correct.
Guard against these with evaluation sets of real queries and known good results, filters for required attributes, hybrid retrieval where exact terms matter, and human review for high-stakes outputs.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Common use cases
Google Cloud lists several patterns for vector databases. Treat them as patterns, not automatic outcomes; results still depend on the data, the embedding model, retrieval configuration, and evaluation.
- Semantic search: find documents with related meaning even when the wording differs from the query.
- Multimodal search: search across represented media, such as finding images from text descriptions, where the chosen models and data support it.
- Retrieval-augmented generation: retrieve relevant documents or records and supply them as context for an LLM-based answer.
- Recommendations: retrieve items similar to a given item or match content to a user’s preference representation.
- Anomaly and fraud detection: compare a record’s representation with patterns in a dataset to help surface unusual cases.
Standalone vector database or an extension?
You do not always need a dedicated vector database. pgvector adds vector search to PostgreSQL, so teams already running PostgreSQL can store embeddings next to their relational data. Dedicated services, including managed offerings such as Pinecone and open-source systems such as Milvus and Weaviate, are built around vector workloads. The right choice depends on operations, data location, filtering, indexing, and scale. Compare these axes:
| Axis | Questions to answer |
|---|---|
| Deployment and operations | Do you want a managed service, a self-hosted service, or an extension inside an existing database? |
| Existing data stack | Does your system already run PostgreSQL or another platform with vector capabilities? |
| Retrieval quality | How do exact and approximate search perform on a representative set of real queries? What recall loss is acceptable? |
| Filtering and hybrid search | Can you apply required metadata or permission filters, and combine keyword matching with vectors? |
| Index resources | What are the query-speed, memory, and index-build trade-offs of the chosen index? |
| Updates and lifecycle | How are vectors refreshed, deleted, backed up, and migrated when the embedding model changes? |
These axes establish what to evaluate, not which product wins. Capabilities and trade-offs are documented by each project, but performance for your workload has to be measured on your own data.
Where to go next
For the primary explanations behind this article, read Google Cloud’s overview at https://cloud.google.com/discover/what-is-a-vector-database, Pinecone’s guide at https://www.pinecone.io/learn/vector-database/, Weaviate’s documentation on vector search and search more broadly, Milvus’s guide to basic vector search, and the pgvector project documentation. Product features and index behavior change between releases, so confirm current details in those sources before making a design decision.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




