Hardware FixRecommendedDevice not working? Your driver may be the problemCheck updates for common hardware issues.Fix DriversOctober DealsAmazon USOctober deal check: compare before you payAmazon US: current deals, useful picks and tech finds.Check DealsSlow PC?RecommendedPC slow today? Run a repair scan before it gets worseResolve common Windows issues and optimize system performance.Scan Now×
Skip to content
World desk4 min

How to Build Semantic Search with pgvector and Python

A practical guide to storing text embeddings in PostgreSQL and retrieving similar results from Python with pgvector, including exact search, HNSW, IVFFlat, and filtering.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

To build semantic search with pgvector and Python, generate embeddings for your documents and queries with a compatible embedding model, store document vectors in PostgreSQL, and order SQL results by the matching vector-distance operator. Start with exact nearest-neighbor search; add HNSW or IVFFlat only when measurements on your workload justify approximate retrieval.

How semantic search works with PostgreSQL

An embedding model converts text into a vector: a list of numbers that represents the text in a model-specific vector space. Semantic search embeds a query in the same compatible space, then finds stored document vectors close to it. pgvector provides PostgreSQL’s vector storage, distance operations, and indexes; it does not generate embeddings. Choose an embedding model and decide how to split and represent your text separately.

The basic flow is:

  1. Generate and store an embedding for each searchable document or text segment.
  2. Generate an embedding for the incoming query using the compatible model.
  3. Ask PostgreSQL to order stored vectors by distance from the query vector and return the closest matches.

The examples below use Psycopg 3. pgvector also documents integrations for Psycopg 2, asyncpg, SQLAlchemy, SQLModel, and Django; follow the integration’s own registration and adaptation instructions for your chosen driver or framework. See the pgvector Python package documentation.

How do I store embeddings in PostgreSQL?

Install pgvector for your PostgreSQL environment, then enable the extension in the database where you will use it. Create a vector column with a dimension that matches the embeddings produced by your chosen model. The dimension of 3 below is just a compact example, not a recommendation for production.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
CREATE EXTENSION IF NOT EXISTS vector;

CREATE TABLE documents (
    id bigint GENERATED ALWAYS AS IDENTITY PRIMARY KEY,
    content text NOT NULL,
    embedding vector(3) NOT NULL
);

In Python, register the vector type on the Psycopg connection before passing vectors as query parameters. Replace the illustrative three-element vector with embeddings generated by your application.

import psycopg
from pgvector.psycopg import register_vector

conn = psycopg.connect("postgresql://user:password@localhost:5432/mydb")
register_vector(conn)

embedding = [0.1, 0.2, 0.3]  # Illustrative only; generate this with your model.

with conn:
    conn.execute(
        "INSERT INTO documents (content, embedding) VALUES (%s, %s)",
        ("Example document text", embedding),
    )

The schema should also carry whatever your application needs to identify and filter results: for example, a document ID, a text value or reference, tenant or category metadata, and the embedding model or version used. There is no single schema that suits every application. Consult the Python package’s driver examples for its documented setup pattern.

How do I query similar vectors with pgvector?

Embed the user’s query with the same compatible model, then use a distance operator in ORDER BY and limit the result count. Here is a Psycopg query using L2 distance, the metric represented by <->:

query_embedding = [0.12, 0.18, 0.31]  # Illustrative only.

rows = conn.execute(
    """
    SELECT id, content
    FROM documents
    ORDER BY embedding <-> %s
    LIMIT 5
    """,
    (query_embedding,),
).fetchall()

Choose the distance measure that fits your embedding model and application, then keep the query operator and any vector index operator class aligned. pgvector’s Python documentation covers L2, inner product, and cosine-distance examples. Using a different operator from the one supported by the index can prevent that index from serving the query.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

When should I add a vector index?

The pgvector project documentation says, “By default, pgvector performs exact nearest neighbor search, which provides perfect recall.” Exact search is a useful correctness baseline. Consider an approximate nearest-neighbor index when measured query latency and corpus size make exact retrieval unsuitable, and compare the results on representative data before choosing one.

Index How it works Build and workload considerations
HNSW Organizes vectors in a multilayer graph. The project describes a stronger speed/recall tradeoff than IVFFlat, with slower index builds and higher memory use. It can be created before data is loaded because it does not require IVFFlat-style training.
IVFFlat Partitions vectors into lists. It requires data for training, so the project advises building it after loading initial data. Query-time probes affect the speed/recall tradeoff.

Neither index is universally best. Evaluate exact and approximate results with representative queries, an application-appropriate recall measure, and realistic latency measurements. Include memory use, index build time, data-loading and update patterns, and operational complexity in the decision; the documentation does not establish a universal corpus-size threshold or speedup.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

How filters affect approximate search

With an approximate index, a SQL filter such as WHERE tenant_id = ... is applied after the index scan. As a result, the scan may return fewer matching rows than the requested limit. In the pgvector documentation’s illustrative example, a filter that matches 10% of rows combined with the default HNSW hnsw.ef_search value of 40 yields four matching rows on average. That is an example, not a guarantee for every dataset or query.

For filtered workloads, pgvector documents iterative index scans, which can scan further to find enough qualifying results. It also suggests considering partial indexes when there are few distinct filter values, or partitioning when there are many. Test these choices with the selectivity and metadata filters your application actually uses.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Tune against your application, not a copied example

  • Measure exact-search latency first to establish a baseline.
  • Compare approximate results with exact results on representative queries and assess recall as well as latency.
  • Check filtered queries separately; their qualifying-result counts may differ from unfiltered search.
  • Review query plans and test index settings using your own data and update pattern.
  • Treat parameter values in examples, such as HNSW construction settings or IVFFlat list counts, as examples rather than universal recommendations.

The pgvector documentation does not establish one best index or parameter set for all hardware and datasets. If you use managed PostgreSQL, verify the provider’s supported pgvector version, limits, and configuration. For example, Google Cloud documents using pgvector with Cloud SQL for PostgreSQL, including an HNSW example.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Leave a Reply

Your email address will not be published. Required fields are marked *

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

More from the Wire

  1. Shenzhen desk3 min
    HONOR Expands Beyond Smartphones With Humanoid Robot RevealHONOR said it unveiled its first humanoid robot at MWC 2026 and named shopping assistance, workplace inspections, and supportive companionship as intended uses. Later Robotics D1 claims and a reported…
  2. Cupertino desk5 min
    Apple Unveils AirPods Max 2: The Upgrade That Should Have Happened Years AgoAirPods Max 2 adds H2-powered audio features and Apple claims up to 1.5Ă— more effective ANC, but its design, Smart Case, and 20-hour battery rating are unchanged. Wired lossless audio…
  3. Cupertino desk4 min
    Apple’s OLED Touch MacBooks Are Coming—but the Dynamic Island Is the Real GambleApple has not announced an OLED touchscreen MacBook, but reports point to high-end models arriving in late 2026 or early 2027. The reported Mac Dynamic Island could be useful, but…
Recommended PC Tool
Recommended PC Tool
PC Slower Than It Used to Be?Free scan - under a minute
Outdated Drivers Are Slowing You DownFree scan - exact matches

Two free Windows tools

One Free Minute Could Fix That PC

Before you go - each of these free tools takes about a minute and tackles what quietly slows a Windows PC down.

Special offer. View Outbyte info, uninstall instructions, EULA, and Privacy Policy.