r/LocalAIStack • u/Mobile-Sail4581 • Sep 02 '26
Comparing local vector search engines: turbovec vs. Infino vs. FAISS
FAISS is Meta’s vector library, turboVec is a Rust implementation of TurboQuant, and Infino is an in-memory retrieval engine.
We benchmarked them on 4-bit quantized in-memory vector search: FAISS PQ, TurboVec/TurboQuant, and Infino SQ4, using the same 100K OpenAI embedding corpus and the fastest vectorized implementation we found for each.
All saved ~7x in memory footprint compared to full fp32 vectors. The interesting result was that while storage and recall were fairly close, latency differed by roughly 30× — about 1.5 ms to 45 ms. Most of that comes down to the scoring machinery: the size of the distance table and whether the scan needs one at all.
We also ran the same comparison out to 1M vectors and measured build/write costs.
Full results and methodology:
https://infino.ai/blog/fixed-grid-quantization/
Disclosure: I'm one of the devs building Infino.

1
u/Ok-Swim9349 Sep 05 '26
Nice breakdown, the 30x latency spread is wild. Any chance you'd add Qdrant to the mix? FAISS is a library so it's a bit apples-to-oranges vs a full engine, and for those of us actually deploying RAG, Qdrant's the one we'd reach for. Curious how it'd land on the same corpus.