Building LifeOS: Implementing In-Browser Vector Search and Semantic Embeddings with React & IndexedDB
Architecting an offline-first AI productivity app: computing vector embeddings, storing high-dimensional vectors in IndexedDB, and executing cosine similarity search.
By Uttam Thapa · · AI/ML
⚡ Executive Summary (TL;DR)
Search "financial planning" in a keyword-based notes app and a note called "Monthly Budget & Savings Strategy" does not come back — same meaning, no shared
words. LifeOS runs semantic vector search entirely inside the browser: embeddings stored as
Float32Array in IndexedDB, cosine similarity computed locally, results in under 12 ms across thousands of notes.
Nothing leaves the device.
Figure 1: Results ranked by meaning rather than by string overlap — computed on the device, offline.
Why Keyword Search Fails on Personal Notes
Personal knowledge bases are the worst case for lexical search. You write a note in one vocabulary and look for it months later in another. The idea is stable;
the words are not. Every miss trains you to distrust the search box, and an untrusted search box turns a knowledge base into a folder.
| You search for |
The note is called |
Keyword |
Vector |
| "financial planning" |
"Monthly Budget & Savings Strategy" |
Miss |
Match |
| "how to sleep better" |
"Evening routine experiments" |
Miss |
Match |
| "deploy notes" |
"Deploy checklist" |
Match |
Match |
What Vector Search Actually Computes
An embedding model turns text into a dense array of floats — often 1,536 dimensions — positioned so that semantically similar text lands nearby. Search stops
being string comparison and becomes geometry: embed the query, then measure the angle between it and every stored document vector.
function cosineSimilarity(vecA, vecB) {
let dotProduct = 0;
let normA = 0;
let normB = 0;
for (let i = 0; i < vecA.length; i++) {
dotProduct += vecA[i] * vecB[i];
normA += vecA[i] * vecA[i];
normB += vecB[i] * vecB[i];
}
return dotProduct / (Math.sqrt(normA) * Math.sqrt(normB));
}
Cosine similarity measures direction, not magnitude, which is what you want: a long note and a short note about the same subject should score alike. One fused
loop computes all three sums in a single pass over the arrays — worth doing, because this function runs once per stored document per keystroke.
💡 Normalise once, at write time
If you unit-normalise every embedding before storing it, cosine similarity collapses to a plain dot product — the two square roots disappear from the hot
loop entirely. Same ranking, meaningfully less work per query.
Why IndexedDB, Not localStorage
A single 1,536-dimension embedding is roughly 6 KB as Float32Array. A few hundred notes will therefore blow through
localStorage's ~5 MB quota, and localStorage is synchronous — every read blocks the main thread and the interface stutters.
📝 Notes store
Title, Markdown body, timestamps and tag metadata — the things you display.
🧮 Embeddings store
Float32Array buffers keyed by document id — the things you compute over. Kept separate so a search never deserialises note bodies.
Splitting the stores is the performance decision that matters most. Ranking touches only the embeddings store; note bodies are fetched for the handful of results
that actually made the top ten.
Offline-First Is a Feature, Not a Constraint
🔒
Privacy
Personal notes and their vectors never leave the device. There is no server to breach.
⚡
<12ms queries
No network round trip, so search can run on every keystroke instead of on submit.
✈️
Works offline
Aeroplane mode, bad hotel wifi, a dead API key — search keeps working.
The honest caveat: a brute-force scan is linear in document count. Thousands of notes are comfortable; hundreds of thousands are not, and at that point you want
an approximate index or a server-side vector database. For a personal workspace, the linear scan is the right trade — it has no index to rebuild, no
approximation to tune, and no recall cliff.
If your corpus is heading past that ceiling, Pinecone vs pgvector vs in-browser embeddings
walks through where each option stops making sense.
✅ Key takeaways
- ✓Store vectors as
Float32Array in IndexedDB. Async, transactional, and gigabytes of headroom.
- ✓Separate the embeddings store from the notes store. Ranking should never touch note bodies.
- ✓Normalise at write time. Turns cosine similarity into a dot product in the hot loop.
- ✓Local search changes the interaction. Sub-frame latency means search-as-you-type instead of search-on-submit.
- ✓Know the ceiling. A linear scan is ideal up to low tens of thousands of documents, and wrong beyond it.
Frequently asked questions
What is in-browser vector search?
Text is converted into embeddings — dense arrays of floats positioned so that similar meanings land near each other — and stored locally in IndexedDB. A query is embedded the same way, and cosine similarity ranks the stored documents without any network call.
Why use IndexedDB instead of localStorage for embeddings?
A single 1,536-dimension embedding is roughly 6 KB, so a few hundred notes exceed localStorage's ~5 MB quota. localStorage is also synchronous, so every read blocks the main thread.
How many documents can a browser-side vector search handle?
A brute-force scan is linear, so thousands of documents are comfortable and sub-frame fast. Beyond low tens of thousands you want an approximate index or a server-side vector database.
Home · Projects · Blog · Services · Résumé · Contact