To use pgvector, install the extension on your PostgreSQL server, enable it in the database where you need it, then create a vector column and an index suited to your distance metric. The SQL below walks through a small working example; the three-dimensional vectors are for demonstration only.
1. Install pgvector on the PostgreSQL server
Installing the extension’s server files is separate from enabling it in a database. The pgvector project documents package-manager routes including Docker, Homebrew, PGXN, APT, and Yum, as well as building from source. Package names and supported PostgreSQL major versions vary, so choose the documented route for your operating system and server version rather than assuming one installation command works everywhere.
For a source build, the project’s current README gives Linux and Mac instructions for PostgreSQL 13 and later. Its example checks out the v0.8.7 branch, then runs make and make install; installation may require elevated privileges. See the pgvector project README for the applicable installation steps. On a managed PostgreSQL service, first check that provider’s current documentation for pgvector availability, supported versions, and permissions; the project README does not establish those details for each provider.
2. Enable the extension in the database
Connect to the specific database that will store vectors and run:
Recommended Free Tools
#1 Best Overall
CREATE EXTENSION vector;
Extension availability is database-specific: run this once in every database where you need pgvector. The connected role must have enough privileges to create the extension; the exact permission requirements can depend on your PostgreSQL setup or hosting provider.
3. Create a vector column and insert sample data
Declare the vector’s dimension in the column type. In this example, vector(3) means each value has three elements:
Rank #2
CREATE TABLE items (
id bigserial PRIMARY KEY,
embedding vector(3)
);
INSERT INTO items (embedding)
VALUES ('[1,2,3]'), ('[4,5,6]');
In an application, set the dimension to match the vectors produced by your embedding model or other vector source. Stored values and query vectors must match the column’s declared dimension. The tiny three-element values here simply make the SQL easy to follow; they are not representative application embeddings.
4. Run a nearest-neighbor query
Start with an exact nearest-neighbor query. This example orders rows by L2 distance from the query vector and returns at most five:
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Rank #3
SELECT *
FROM items
ORDER BY embedding <-> '[3,1,2]'
LIMIT 5;
pgvector’s distance operators include <-> for L2 distance, <#> for negative inner product, <=> for cosine distance, and <+> for L1 distance. PostgreSQL supports ascending-order index scans on operators, so <#> returns negative inner product; multiply that result by -1 if you need the positive inner product value.
5. Create an approximate vector index
Without an approximate index, pgvector uses exact nearest-neighbor search, which provides perfect recall. Approximate indexes can improve query speed but trade away some recall, so results may differ from exact search. For the L2 query above, a direct HNSW index is:
CREATE INDEX ON items USING hnsw (embedding vector_l2_ops);
Choose the operator class to match the query’s distance operator:
- For L2 distance with
<->, usevector_l2_ops. - For cosine distance with
<=>, usevector_cosine_ops. - For inner product with
<#>, usevector_ip_ops.
Keep the metric used by the query and the index operator class aligned. For production workloads, the project recommends creating indexes after initial bulk loading for best performance, and using concurrent index creation to avoid blocking writes. Consult the project README for the relevant index-creation guidance and syntax.
Quick wins for a faster PC:
Clear out junk files and repair common Windows errorsFree Scan →Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →HNSW or IVFFlat?
These are qualitative tradeoffs described by the pgvector project, not independent benchmark results. The right choice depends on your data, workload, and tolerance for recall loss.
| Consideration | HNSW | IVFFlat |
|---|---|---|
| Speed/recall tradeoff in project guidance | Better query performance than IVFFlat in the project’s speed-recall comparison | Lower query performance than HNSW in that comparison |
| Build and memory | Slower to build; uses more memory | Faster to build; uses less memory |
| Building on an empty table | Can be created before the table has data | Build after the table has some data for good recall |
| Basic index form | CREATE INDEX ON items USING hnsw (embedding vector_l2_ops); |
CREATE INDEX ON items USING ivfflat (embedding vector_l2_ops) WITH (lists = ...); |
IVFFlat starting points and tuning
The project README suggests starting with rows / 1000 lists for tables up to one million rows, and sqrt(rows) lists for larger tables. It suggests starting with sqrt(lists) probes. These are tuning starting points, not performance guarantees: measure against your own data and queries. Increasing probes favors recall over speed.
When filtered searches return too few rows
With approximate indexes, filtering is applied after the index scan. A selective WHERE condition can therefore leave fewer matching rows than the requested LIMIT. Depending on the workload, the project README points to iterative index scans, ordinary indexes on filter columns, partial indexes, or partitioning as ways to address this. Choose based on the filter and data layout, then verify that the query returns enough relevant results.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.
The Tool Desk
Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →




