What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
To build semantic search with pgvector and Python, generate embeddings for your documents and queries with a compatible embedding model, store document vectors in PostgreSQL, and order SQL results by the matching vector-distance operator. Start with exact nearest-neighbor search; add HNSW or IVFFlat only when measurements on your workload justify approximate retrieval.
How semantic search works with PostgreSQL
An embedding model converts text into a vector: a list of numbers that represents the text in a model-specific vector space. Semantic search embeds a query in the same compatible space, then finds stored document vectors close to it. pgvector provides PostgreSQL’s vector storage, distance operations, and indexes; it does not generate embeddings. Choose an embedding model and decide how to split and represent your text separately.
The basic flow is:
- Generate and store an embedding for each searchable document or text segment.
- Generate an embedding for the incoming query using the compatible model.
- Ask PostgreSQL to order stored vectors by distance from the query vector and return the closest matches.
The examples below use Psycopg 3. pgvector also documents integrations for Psycopg 2, asyncpg, SQLAlchemy, SQLModel, and Django; follow the integration’s own registration and adaptation instructions for your chosen driver or framework. See the pgvector Python package documentation.
How do I store embeddings in PostgreSQL?
Install pgvector for your PostgreSQL environment, then enable the extension in the database where you will use it. Create a vector column with a dimension that matches the embeddings produced by your chosen model. The dimension of 3 below is just a compact example, not a recommendation for production.
#1 Best Overall
CREATE EXTENSION IF NOT EXISTS vector;
CREATE TABLE documents (
id bigint GENERATED ALWAYS AS IDENTITY PRIMARY KEY,
content text NOT NULL,
embedding vector(3) NOT NULL
);
In Python, register the vector type on the Psycopg connection before passing vectors as query parameters. Replace the illustrative three-element vector with embeddings generated by your application.
import psycopg
from pgvector.psycopg import register_vector
conn = psycopg.connect("postgresql://user:password@localhost:5432/mydb")
register_vector(conn)
embedding = [0.1, 0.2, 0.3] # Illustrative only; generate this with your model.
with conn:
conn.execute(
"INSERT INTO documents (content, embedding) VALUES (%s, %s)",
("Example document text", embedding),
)
The schema should also carry whatever your application needs to identify and filter results: for example, a document ID, a text value or reference, tenant or category metadata, and the embedding model or version used. There is no single schema that suits every application. Consult the Python package’s driver examples for its documented setup pattern.
Rank #2
How do I query similar vectors with pgvector?
Embed the user’s query with the same compatible model, then use a distance operator in ORDER BY and limit the result count. Here is a Psycopg query using L2 distance, the metric represented by <->:
query_embedding = [0.12, 0.18, 0.31] # Illustrative only.
rows = conn.execute(
"""
SELECT id, content
FROM documents
ORDER BY embedding <-> %s
LIMIT 5
""",
(query_embedding,),
).fetchall()
Choose the distance measure that fits your embedding model and application, then keep the query operator and any vector index operator class aligned. pgvector’s Python documentation covers L2, inner product, and cosine-distance examples. Using a different operator from the one supported by the index can prevent that index from serving the query.
The Tool Desk
Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →When should I add a vector index?
The pgvector project documentation says, “By default, pgvector performs exact nearest neighbor search, which provides perfect recall.” Exact search is a useful correctness baseline. Consider an approximate nearest-neighbor index when measured query latency and corpus size make exact retrieval unsuitable, and compare the results on representative data before choosing one.
| Index | How it works | Build and workload considerations |
|---|---|---|
| HNSW | Organizes vectors in a multilayer graph. | The project describes a stronger speed/recall tradeoff than IVFFlat, with slower index builds and higher memory use. It can be created before data is loaded because it does not require IVFFlat-style training. |
| IVFFlat | Partitions vectors into lists. | It requires data for training, so the project advises building it after loading initial data. Query-time probes affect the speed/recall tradeoff. |
Neither index is universally best. Evaluate exact and approximate results with representative queries, an application-appropriate recall measure, and realistic latency measurements. Include memory use, index build time, data-loading and update patterns, and operational complexity in the decision; the documentation does not establish a universal corpus-size threshold or speedup.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.How filters affect approximate search
With an approximate index, a SQL filter such as WHERE tenant_id = ... is applied after the index scan. As a result, the scan may return fewer matching rows than the requested limit. In the pgvector documentation’s illustrative example, a filter that matches 10% of rows combined with the default HNSW hnsw.ef_search value of 40 yields four matching rows on average. That is an example, not a guarantee for every dataset or query.
For filtered workloads, pgvector documents iterative index scans, which can scan further to find enough qualifying results. It also suggests considering partial indexes when there are few distinct filter values, or partitioning when there are many. Test these choices with the selectivity and metadata filters your application actually uses.
Best Value
Tune against your application, not a copied example
- Measure exact-search latency first to establish a baseline.
- Compare approximate results with exact results on representative queries and assess recall as well as latency.
- Check filtered queries separately; their qualifying-result counts may differ from unfiltered search.
- Review query plans and test index settings using your own data and update pattern.
- Treat parameter values in examples, such as HNSW construction settings or IVFFlat list counts, as examples rather than universal recommendations.
The pgvector documentation does not establish one best index or parameter set for all hardware and datasets. If you use managed PostgreSQL, verify the provider’s supported pgvector version, limits, and configuration. For example, Google Cloud documents using pgvector with Cloud SQL for PostgreSQL, including an HNSW example.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




