Choosing a Vector Database for RAG
The right vector database for a production RAG system is usually the one that adds the fewest new moving parts to your stack, not the one with the best benchmark numbers. For most teams below tens of millions of vectors, that means starting with pgvector on Postgres you already operate, and only moving to a dedicated vector database once a specific, measured limitation shows up — query latency at your real concurrency, filtering performance on metadata, or index build time on your update frequency.
What a Vector Database Actually Needs to Do
Strip away the marketing and a vector database has three jobs: store embeddings alongside enough metadata to filter on, build an approximate nearest-neighbor index so similarity search does not degrade to a linear scan, and serve queries fast enough that retrieval is not the bottleneck in your RAG pipeline’s latency budget. Every option below does all three; they differ in how they are operated, how they scale, and how much they cost you in engineering time rather than in the search algorithm itself, since most of these systems converge on the same family of approximate nearest-neighbor algorithms (HNSW being the most common).
The Real Decision: Embedded vs. Managed vs. Self-Hosted
Embedded in your existing database. pgvector adds vector columns and HNSW indexing directly to Postgres. If your application data already lives in Postgres, this means one fewer system to operate, one fewer network hop per query, and transactional consistency between your vectors and the rows they describe — a property that is easy to take for granted until you need it, such as when a document is deleted and its embedding must disappear atomically with it. Supabase ships pgvector as a first-class extension, which is why it shows up as the default recommendation in a lot of the Supabase and AI codebase to production work in this practice.
Managed vector databases. Services that run a purpose-built vector store for you, exposed over an API, remove the operational burden of running the database yourself — no index tuning, no capacity planning, no upgrade windows. The trade-off is a new external dependency in your critical path, a data-residency question if the service is not in a region your compliance posture requires, and less control over exactly how the index is tuned. This category makes the most sense for teams without the operational capacity to run stateful infrastructure themselves, or for use cases where vector search volume dwarfs everything else the application does.
Self-hosted, purpose-built vector databases. Running a dedicated vector database yourself — as opposed to a managed version of one — gives you the performance characteristics of a system built specifically for this workload, plus full control over deployment, without handing an external vendor your data path. The cost is operational: you now run and monitor a new stateful service, with its own upgrade cadence and failure modes, on top of everything else your team already operates.
Comparison
| Factor | pgvector (embedded) | Managed vector DB | Self-hosted vector DB |
|---|---|---|---|
| New systems to operate | None (uses existing Postgres) | None (vendor-operated) | One |
| Latency to your app data | Lowest (same database) | Extra network hop | Extra network hop |
| Transactional consistency with app data | Yes | No | No |
| Scaling ceiling | High, but shares Postgres resources | Vendor-dependent, usually high | High, requires your own capacity planning |
| Operational burden | Low (part of existing Postgres ops) | Lowest | Highest |
| Data residency control | Full (your Postgres instance) | Depends on vendor regions | Full |
A Worked Example: When to Move Off pgvector
A reasonable default path looks like this. Start with pgvector because your application data is already in Postgres and the dataset is in the low millions of vectors. Monitor three things as usage grows: p99 query latency under real concurrent load, not a synthetic benchmark; index build and rebuild time, which matters if your corpus updates frequently; and whether metadata filtering (searching within a subset of vectors by category, tenant, or date) is fast enough for your access patterns, since filtered vector search is harder to optimize than unfiltered search.
If p99 latency degrades past your budget under real load, first check whether the HNSW index parameters are tuned for your data before assuming you need a different system — a mistuned ef_search or m parameter accounts for a large share of “pgvector is too slow” reports. If tuning does not close the gap, or if your dataset moves past what a single Postgres instance can reasonably hold in memory for fast search, that is the point to evaluate a dedicated vector database, at which point the choice between managed and self-hosted comes down to whether your team already operates stateful infrastructure and whether data residency requirements permit a managed vendor.
Limitations
None of this advice replaces benchmarking against your own data and query patterns — vector database performance is sensitive enough to embedding dimensionality, corpus size, and filter selectivity that a generic recommendation can be wrong for a specific workload. This article also deliberately avoids naming prices or licence terms for the managed and self-hosted options, since pricing changes faster than an article can track it and a licence claim about a third-party product is worth verifying against that product’s own documentation rather than taking secondhand.
FAQ
Is pgvector as fast as a dedicated vector database?
For most production RAG workloads — low millions of vectors, moderate query volume — a properly tuned pgvector index is fast enough that the difference is not the bottleneck in your pipeline. At very large scale or very high query-per-second requirements, purpose-built vector databases can out-perform it, but that ceiling is higher than most teams’ actual traffic.
Do I need a vector database at all for a small RAG system?
Below roughly tens of thousands of documents, a brute-force similarity search (comparing the query against every stored vector) can be fast enough without any approximate-nearest-neighbor index at all. Reach for indexing once linear scan time shows up in your latency budget, not before.
Can I switch vector databases later without starting over?
Usually yes, if you keep your embedding generation and chunking logic separate from the storage layer. The embeddings themselves are portable; what is expensive to redo is the chunking and metadata design, so invest the design effort there rather than in the storage choice.
Bottom Line
Pick the option that adds the least new operational surface area for the scale you are actually at, start with pgvector if your data already lives in Postgres, and move to a managed or self-hosted vector database only once a measured limitation — not a benchmark you read somewhere — tells you to. The Python & ML engineering and Supabase practices both cover this decision as part of production RAG builds, and the broader question of getting an AI-generated or AI-assisted RAG pipeline production-ready is covered in the AI codebase to production guide.
Related Articles
MLOps for Small Teams: A Playbook
A practical MLOps playbook for teams too small for a platform team: what to automate first, what to skip, and when to add each piece.
AI Agent Observability in Production
What to log, trace, and alert on when an AI agent runs in production, and the three signals that catch failures a reviewer would otherwise miss.
Building Forward-Deployed AI Systems
Forward deployment engineering builds AI systems that hold up in production, not demos. The patterns and guardrails that separate prototype AI from production AI.