October DealsAmazon USOctober deal check: compare before you payAmazon US: current deals, useful picks and tech finds.Check DealsPC HealthRecommendedCrashes, freezes, slowdowns? Check your PC nowSpot repairable issues before they interrupt work.Check PCOctober DealsAmazon USDeal season is back - check today's better picksAmazon US: current deals, useful picks and tech finds.See Picks×
Skip to content
World desk4 min

What Is pgvector? PostgreSQL’s Built-In Route to Vector Search

pgvector adds vector storage and similarity search to PostgreSQL. Here’s how exact search, HNSW, IVFFlat, filtering, and hybrid retrieval affect the choice.
Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

pgvector is a PostgreSQL extension that adds vector data types and similarity-search operators to a database you may already run. It lets applications store embeddings beside regular records and query them with SQL, without requiring a separate vector database for every use case. The station-wagon analogy has limits: whether pgvector suits your workload depends on measured recall, latency, filtering, memory, and operating scale.

What is pgvector?

pgvector extends PostgreSQL; it is not a standalone database. The extension adds vector types and distance operators, so embeddings can live in the same tables as the records they describe. That makes it possible to use familiar PostgreSQL features such as transactions, joins, replication, and point-in-time recovery alongside vector search. See the pgvector project documentation.

Vector search compares numerical representations of items—often called embeddings—to find items that are close under a chosen distance or similarity measure. The project documents vector, halfvec, bit, and sparsevec representations, as well as L2, inner-product, cosine, L1, Hamming, and Jaccard operators. Choose an index operator class that matches the distance function used by your query; otherwise, the index may not serve the search as intended.

How search works: exact results first, approximate indexes when needed

Without an approximate index, pgvector performs exact nearest-neighbor search. That is a useful baseline: it gives you a reference for checking which results an approximate index misses. HNSW and IVFFlat can speed up searches, but approximate search trades some recall—the share of relevant neighbors found—for speed, and can change which results appear.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

There is no documented universal vector-count threshold at which you must move to a specialized database. Compare exact and approximate search using representative embeddings, query patterns, filters, concurrency, and production-relevant latency and memory limits.

Should you use HNSW or IVFFlat?

The project describes these as general tradeoffs, not benchmark results for your data. Measure both against exact search and tune them with your actual query and filter distribution.

Decision point HNSW IVFFlat
Speed/recall tradeoff Generally better Generally lower
Index build Slower Faster
Memory use Higher Lower
When to build Can be created on an empty table Build after loading data
Main tuning concepts m, ef_construction, and hnsw.ef_search lists and ivfflat.probes

HNSW’s graph-based index generally offers the stronger speed/recall tradeoff, at the cost of more memory and slower builds. IVFFlat generally builds faster and uses less memory, but needs tuning and tends to offer a weaker speed/recall tradeoff. The pgvector documentation explains the index options and their parameters.

How many dimensions can you index in pgvector?

The project documents these indexing limits: 2,000 dimensions for vector, 4,000 for halfvec, 64,000 for bit, and 1,000 non-zero elements for sparsevec. These are documented limits, not recommended workload sizes or guarantees of useful performance. Check the current project documentation for details.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Why can filters return too few rows?

With an approximate index, filtering is applied after the index scan. The scan may therefore produce fewer rows that satisfy a filter than your application expects. In its README, pgvector illustrates the effect with a filter matching 10% of rows and the default HNSW ef_search value of 40: the example yields about four matching rows on average. That is an illustration from the project, not a guarantee for other query distributions.

For filtered approximate searches, pgvector documents several approaches:

  • Iterative scans: Available starting in pgvector 0.8.0; they can continue scanning to find more qualifying rows.
  • Partial indexes: Consider them when a filter has only a few distinct values.
  • Partitioning: Consider it when a filter has many distinct values.

For multitenant workloads, a shared approximate index can let one tenant’s vectors affect another tenant’s recall and speed. The project suggests list partitioning or separate tables when tenant isolation is important. Review its guidance on filtering and approximate indexes, then test with the actual tenant and filter distribution.

Can pgvector support hybrid search?

Yes. PostgreSQL full-text search can be combined with pgvector search. The project describes approaches such as reciprocal rank fusion and cross-encoders for combining candidate lists; these are ranking techniques to implement, not one-click built-in strategies. Hybrid retrieval can be useful when a query needs both semantic similarity and text matching. See the project’s hybrid-search guidance.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

When does keeping vectors in PostgreSQL make sense?

An existing PostgreSQL deployment is a reasonable place to evaluate vector search when keeping embeddings, relational records, joins, transactions, and database operations together is more valuable than the performance or operational benefits of a separate vector system. The integration alone cannot establish that pgvector will meet a particular workload’s requirements.

Before choosing, test representative embeddings and filters, query concurrency, recall against exact search, latency, index-build time, and memory use. Consider a specialized system if those measurements or operational requirements show that one database is not the right fit; there is no universal capacity cutoff supported by the project documentation.

Version, installation, and security checks

The project documentation describes support for PostgreSQL 13 and later. Version information in the project’s public materials is inconsistent: the GitHub tags page lists v0.8.6, dated 2026-07-29, as its newest visible tag, while the README installation command refers to v0.8.7. Check the current release artifact and installation instructions before pinning a version; the available project materials do not resolve that discrepancy. See the pgvector tags and the project README.

In a notice dated 2026-02-26, PostgreSQL said pgvector 0.8.2 fixes CVE-2026-3172, a buffer overflow in parallel HNSW index builds that could leak data from other relations or crash the database server. The notice encouraged users to upgrade. That advisory does not establish that later versions have no subsequent issues, so check current advisories and release notes before deploying. Read the PostgreSQL security notice.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

For managed hosting, Amazon Web Services documents pgvector support in Aurora PostgreSQL. Aurora is one deployment option, not a requirement for using pgvector. AWS also claims “up to 9x” more vector-search queries per second for workloads exceeding available instance memory with Aurora optimized reads; that is an AWS claim about Aurora, not an independent benchmark or a general result for pgvector deployments. See AWS Aurora PostgreSQL vector-search documentation.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Leave a Reply

Your email address will not be published. Required fields are marked *

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

More from the Wire

  1. World desk4 min
    How to Spot an AI Voice Scam Before Sending MoneyDon’t rely on how a caller sounds. Pause, call back through a known number, and verify the emergency with another trusted person before sending money.
  2. Mountain View desk4 min
    Google’s SynthID Detector: How to Check AI-Generated Images, Video and AudioGoogle’s SynthID Detector looks for an embedded watermark in supported images, video and audio. Here is what its results do—and do not—show.
  3. Redmond desk20 min
    How to create a link to File or Folder in Windows 11Windows 11 gives you several ways to point to a file or folder without moving or duplicating it. You can create a desktop shortcut,…
Recommended PC Tool
Recommended PC Tool
Crashes, No Sound, or Screen Glitches?Free driver scan
Windows Errors? Fix Them Before They SpreadFree repair scan

Two free Windows tools

One Free Minute Could Fix That PC

Before you go - each of these free tools takes about a minute and tackles what quietly slows a Windows PC down.

Special offer. View Outbyte info, uninstall instructions, EULA, and Privacy Policy.