Our v0 was ChromaDB. Our v1 was FAISS behind a thin Python wrapper. Our v2 was a fork of FAISS with custom shard management. Our v3 — which is what ships today — is our own vector index, our own storage layer, and zero external dependencies in the hot path.
This is a short note on why we made that journey. It's not a critique of either project. Both are excellent at what they were designed for. They just weren't designed for what we needed.
The Chroma chapter
We started with Chroma because it's the fastest path from zero to a working RAG demo. It still is. For tenants with under 100k chunks, Chroma is genuinely great and our recommendation if you don't need anything we offer.
The break for us came around 2M chunks per tenant. Specifically:
- ›Cold start. Our per-tenant collection model meant cold tenants paid a startup cost for index loading that we couldn't amortise.
- ›Operational opaqueness. When a query was slow, we couldn't tell whether the bottleneck was disk, RAM, or the kernel. Chroma's profiling story wasn't there yet.
- ›Schema evolution. Adding a new metadata field meant reindexing.
We hit the first two limits within a quarter. The third one was the breaking change.
The FAISS chapter
We moved to FAISS for the index, with a Postgres-backed metadata store on the side. This solved the cold-start problem and gave us first-class observability — FAISS's IVF indexes are well-instrumented and we could finally see where queries were spending time.
The new problem: operational complexity. FAISS gives you index primitives. It does not give you sharding, replication, online retraining, or multi-tenant isolation. We built all of those.
After six months we'd written 11,000 lines of glue code around FAISS. Most of it was load balancing across shard replicas, cache eviction policies, and pre-warming logic for tenants we knew were about to wake up. None of it was the interesting part of the problem.
The breaking change here was the realisation that we were building a database, not using one.
Building our own
The v3 system is purpose-built for our workload, which has three specific shapes:
- ›Multi-tenant. Tens of thousands of small-to-medium collections, not a handful of giant ones. Per-tenant isolation is non-negotiable.
- ›Burst-and-idle. A tenant might do nothing for 12 hours and then 50,000 queries in 5 minutes. We need to spin up fast, spin down faster.
- ›Hybrid retrieval. Vector search is one signal of seven (see the Engram post). The index needs to surface candidates that get heavily re-ranked downstream.
Building for those constraints let us make decisions FAISS and Chroma can't make. Our index format is unique to our workload. Our shard placement algorithm is unique to our workload. Our query planner cuts work the moment Engram tells it to.
What we kept from FAISS
We didn't reinvent the math. The actual vector arithmetic in our index uses ideas from FAISS (HNSW graph construction, product quantisation for high-dim vectors). What changed was everything around the math: the layer that decides which shard to query, how to merge results, when to evict cached buffers.
What we kept from Chroma
The shape of the developer API. Chroma's mental model — collections, metadata filters, simple add/query/delete operations — is the right one. Our SDK ergonomics owe a lot to it.
What we tell early-stage teams
Use Chroma. Honestly. If you're building your first RAG product, the right vector store is the one that lets you ship a demo this week. You won't hit our problems until you're way past where Chroma stops being good.
The moment you find yourself writing more than 1,000 lines of glue around your vector store, that's the signal you might need something purpose-built. Until then, Chroma is fine.
Tomás Reyes
April 8, 2026 · 6 min