Pinecone vs Pgvector vs In-Browser Embeddings: Choosing the Right Vector Database for AI Applications
Comparing query latency, scale limitations, operational costs, and privacy tradeoffs across Pinecone Cloud, Pgvector HNSW indexes, and client-side vector search.
By Uttam Thapa · · AI/ML
⚡ Executive Summary (TL;DR)
The vector store is rarely what makes retrieval good or bad — your embedding model and chunking strategy decide that. What the store decides is operational
cost and how much of your data you can filter and join in one place. Use pgvector if you already run PostgreSQL, a managed
vector service once scale or index tuning becomes the job itself, and in-browser search when the data should never leave the device.
Figure 1: Retrieval quality is set upstream of the database. The database sets the bill.
What a Vector Database Is Actually For
An embedding model turns text into a dense array of floats positioned so that semantically similar text lands nearby. A vector store keeps those arrays and,
given a query vector, returns its nearest neighbours by cosine distance. That is the whole job. Everything else — filtering, hybrid search, index tuning — is
about doing it fast enough over enough vectors.
The Three Options, Compared Honestly
|
pgvector |
Managed vector service |
In-browser |
| Comfortable scale | Up to a few million vectors | Hundreds of millions | Thousands |
| Filter + join with your data | Plain SQL, one transaction | Metadata filters only | Whatever you write |
| Operational burden | A database you already run | None | None |
| Cost model | Existing instance | Per vector and per query | Free |
| Privacy | Your infrastructure | Third party | Never leaves the device |
pgvector: One Database, One Transaction
If PostgreSQL is already in your stack, adding pgvector means embeddings live beside the rows they describe — so a similarity search can join to permissions,
tenancy and status in a single query. That is the advantage that gets undersold, because the alternative is fetching ids from a vector service and then filtering
them against your database, which is two round trips and a correctness problem when the two drift.
CREATE EXTENSION IF NOT EXISTS vector;
CREATE TABLE document_chunks (
id UUID PRIMARY KEY DEFAULT gen_random_uuid(),
tenant_id UUID NOT NULL,
content TEXT NOT NULL,
embedding vector(1536)
);
-- HNSW: fast approximate search, slower to build than IVFFlat, better recall.
CREATE INDEX ON document_chunks
USING hnsw (embedding vector_cosine_ops);
-- The query a dedicated vector service cannot express in one hop.
SELECT c.id, c.content
FROM document_chunks c
JOIN documents d ON d.id = c.document_id
WHERE c.tenant_id = $1
AND d.status = 'published'
ORDER BY c.embedding <=> $2 -- cosine distance
LIMIT 8;
🚨 Filtered vector search is the hard case everywhere
Approximate indexes search a neighbourhood graph. Apply a restrictive filter and the nearest neighbours may all be filtered out, so you get fewer results
than requested — or the planner abandons the index and scans. This is not a pgvector quirk; every approximate index has the same tension. Test recall with
your real filters, not with an unfiltered benchmark.
In-Browser: Small, Private, Instant
For personal-scale corpora, a brute-force scan in JavaScript is genuinely competitive. There is no index to build, no approximation to tune, no network hop, and
the data never leaves the device.
export function computeCosineSimilarity(vectorA: Float32Array, vectorB: Float32Array): number {
let dot = 0.0, normA = 0.0, normB = 0.0;
for (let i = 0; i < vectorA.length; i++) {
dot += vectorA[i] * vectorB[i];
normA += vectorA[i] * vectorA[i];
normB += vectorB[i] * vectorB[i];
}
return normA && normB ? dot / (Math.sqrt(normA) * Math.sqrt(normB)) : 0;
}
Store the vectors in IndexedDB rather than localStorage — a single 1,536-dimension embedding is around 6 KB, so a few hundred documents exceed
the localStorage quota. Normalise vectors at write time and the two square roots disappear from the loop entirely.
What Actually Determines Retrieval Quality
Teams switch vector databases hoping for better answers and get the same answers faster. The levers that matter sit upstream:
Chunking
Chunks that split mid-idea retrieve fragments that answer nothing. Respect document structure and overlap slightly.
Embedding model
Domain fit beats dimension count. More dimensions cost storage and query time without guaranteeing relevance.
Hybrid search
Combine vector similarity with keyword matching. Exact terms — product codes, names — are where pure vectors are weakest.
💡 Start where your data already is
Begin with pgvector in the database you already operate. You will learn what your queries and filters actually look like at no additional infrastructure
cost, and migrating to a dedicated service later is a data move — not a rewrite. Choosing a specialised store on day one usually means optimising a
bottleneck you have not measured.
✅ Key takeaways
- ✓The store sets cost and joins, not relevance. Chunking and the embedding model set relevance.
- ✓pgvector's real advantage is SQL. Filtering and joining in one transaction avoids a whole class of drift.
- ✓Test recall with your real filters. Approximate indexes degrade exactly where benchmarks do not look.
- ✓In-browser search wins on privacy and latency up to a few thousand documents.
- ✓Add keyword search alongside vectors. Exact identifiers are where embeddings fail hardest.
- ✓Migrate when you have measured a limit, not in anticipation of one.
See the browser-side approach end to end in building LifeOS, and
scaling PostgreSQL for keeping the database underneath pgvector healthy.
Frequently asked questions
Should I use Pinecone or pgvector?
Use pgvector when your vectors belong beside relational data you already query and the corpus fits comfortably in your database. Use a dedicated vector service when scale, index tuning or query volume become the primary concern.
Is in-browser vector search a real option?
For personal-scale corpora, yes — it is private, offline-capable and has no network latency. It stops making sense once a linear scan over the corpus exceeds a frame budget.
What actually determines vector search quality?
The embedding model and your chunking strategy, far more than the database. A better index makes the same mediocre retrieval faster, not more relevant.
Home · Projects · Blog · Services · Résumé · Contact