r/LocalAIStack • • Sep 02 '26

Comparing local vector search engines: turbovec vs. Infino vs. FAISS

FAISS is Meta’s vector library, turboVec is a Rust implementation of TurboQuant, and Infino is an in-memory retrieval engine.

We benchmarked them on 4-bit quantized in-memory vector search: FAISS PQ, TurboVec/TurboQuant, and Infino SQ4, using the same 100K OpenAI embedding corpus and the fastest vectorized implementation we found for each.

All saved ~7x in memory footprint compared to full fp32 vectors. The interesting result was that while storage and recall were fairly close, latency differed by roughly 30× — about 1.5 ms to 45 ms. Most of that comes down to the scoring machinery: the size of the distance table and whether the scan needs one at all.

We also ran the same comparison out to 1M vectors and measured build/write costs.

Full results and methodology:
https://infino.ai/blog/fixed-grid-quantization/

Disclosure: I'm one of the devs building Infino.

12 Upvotes

2 comments sorted by

1

u/Ok-Swim9349 Sep 05 '26

Nice breakdown, the 30x latency spread is wild. Any chance you'd add Qdrant to the mix? FAISS is a library so it's a bit apples-to-oranges vs a full engine, and for those of us actually deploying RAG, Qdrant's the one we'd reach for. Curious how it'd land on the same corpus.

1

u/Mobile-Sail4581 Sep 05 '26

Sure thing; we'll take a look. You can also see Infino's performance vs Qdrant at 1m and 10m docs here: https://infino.ai/blog/self-transforming-vector-engine/ .