AI Engineering

Choosing a Vector Database for RAG

Dipankar Sarkar · · 5 min read

The right vector database for a production RAG system is usually the one that adds the fewest new moving parts to your stack, not the one with the best benchmark numbers. For most teams below tens of millions of vectors, that means starting with pgvector on Postgres you already operate, and only moving to a dedicated vector database once a specific, measured limitation shows up — query latency at your real concurrency, filtering performance on metadata, or index build time on your update frequency.

What a Vector Database Actually Needs to Do

Strip away the marketing and a vector database has three jobs: store embeddings alongside enough metadata to filter on, build an approximate nearest-neighbor index so similarity search does not degrade to a linear scan, and serve queries fast enough that retrieval is not the bottleneck in your RAG pipeline’s latency budget. Every option below does all three; they differ in how they are operated, how they scale, and how much they cost you in engineering time rather than in the search algorithm itself, since most of these systems converge on the same family of approximate nearest-neighbor algorithms (HNSW being the most common).

The Real Decision: Embedded vs. Managed vs. Self-Hosted

Embedded in your existing database. pgvector adds vector columns and HNSW indexing directly to Postgres. If your application data already lives in Postgres, this means one fewer system to operate, one fewer network hop per query, and transactional consistency between your vectors and the rows they describe — a property that is easy to take for granted until you need it, such as when a document is deleted and its embedding must disappear atomically with it. Supabase ships pgvector as a first-class extension, which is why it shows up as the default recommendation in a lot of the Supabase and AI codebase to production work in this practice.

Managed vector databases. Services that run a purpose-built vector store for you, exposed over an API, remove the operational burden of running the database yourself — no index tuning, no capacity planning, no upgrade windows. The trade-off is a new external dependency in your critical path, a data-residency question if the service is not in a region your compliance posture requires, and less control over exactly how the index is tuned. This category makes the most sense for teams without the operational capacity to run stateful infrastructure themselves, or for use cases where vector search volume dwarfs everything else the application does.

Self-hosted, purpose-built vector databases. Running a dedicated vector database yourself — as opposed to a managed version of one — gives you the performance characteristics of a system built specifically for this workload, plus full control over deployment, without handing an external vendor your data path. The cost is operational: you now run and monitor a new stateful service, with its own upgrade cadence and failure modes, on top of everything else your team already operates.

Comparison

Factorpgvector (embedded)Managed vector DBSelf-hosted vector DB
New systems to operateNone (uses existing Postgres)None (vendor-operated)One
Latency to your app dataLowest (same database)Extra network hopExtra network hop
Transactional consistency with app dataYesNoNo
Scaling ceilingHigh, but shares Postgres resourcesVendor-dependent, usually highHigh, requires your own capacity planning
Operational burdenLow (part of existing Postgres ops)LowestHighest
Data residency controlFull (your Postgres instance)Depends on vendor regionsFull

A Worked Example: When to Move Off pgvector

A reasonable default path looks like this. Start with pgvector because your application data is already in Postgres and the dataset is in the low millions of vectors. Monitor three things as usage grows: p99 query latency under real concurrent load, not a synthetic benchmark; index build and rebuild time, which matters if your corpus updates frequently; and whether metadata filtering (searching within a subset of vectors by category, tenant, or date) is fast enough for your access patterns, since filtered vector search is harder to optimize than unfiltered search.

If p99 latency degrades past your budget under real load, first check whether the HNSW index parameters are tuned for your data before assuming you need a different system — a mistuned ef_search or m parameter accounts for a large share of “pgvector is too slow” reports. If tuning does not close the gap, or if your dataset moves past what a single Postgres instance can reasonably hold in memory for fast search, that is the point to evaluate a dedicated vector database, at which point the choice between managed and self-hosted comes down to whether your team already operates stateful infrastructure and whether data residency requirements permit a managed vendor.

Limitations

None of this advice replaces benchmarking against your own data and query patterns — vector database performance is sensitive enough to embedding dimensionality, corpus size, and filter selectivity that a generic recommendation can be wrong for a specific workload. This article also deliberately avoids naming prices or licence terms for the managed and self-hosted options, since pricing changes faster than an article can track it and a licence claim about a third-party product is worth verifying against that product’s own documentation rather than taking secondhand.

FAQ

Is pgvector as fast as a dedicated vector database?

For most production RAG workloads — low millions of vectors, moderate query volume — a properly tuned pgvector index is fast enough that the difference is not the bottleneck in your pipeline. At very large scale or very high query-per-second requirements, purpose-built vector databases can out-perform it, but that ceiling is higher than most teams’ actual traffic.

Do I need a vector database at all for a small RAG system?

Below roughly tens of thousands of documents, a brute-force similarity search (comparing the query against every stored vector) can be fast enough without any approximate-nearest-neighbor index at all. Reach for indexing once linear scan time shows up in your latency budget, not before.

Can I switch vector databases later without starting over?

Usually yes, if you keep your embedding generation and chunking logic separate from the storage layer. The embeddings themselves are portable; what is expensive to redo is the chunking and metadata design, so invest the design effort there rather than in the storage choice.

Bottom Line

Pick the option that adds the least new operational surface area for the scale you are actually at, start with pgvector if your data already lives in Postgres, and move to a managed or self-hosted vector database only once a measured limitation — not a benchmark you read somewhere — tells you to. The Python & ML engineering and Supabase practices both cover this decision as part of production RAG builds, and the broader question of getting an AI-generated or AI-assisted RAG pipeline production-ready is covered in the AI codebase to production guide.

Dipankar Sarkar

Dipankar Sarkar

AI Enablement, Contract AI Engineering & Delivery

Related Articles