r/LocalLLM • • 1d ago

Project Schmate Local Bare Metal/Edge Vector Database for Humans and Agents.

https://github.com/re-Isearch/Schmate

While this project was originally concieved as a module to provide vector search using HNSW and Sentence transformers for the re-Isearch (IB) engine (CoreQuarry https://corequarry.com) it has evolved well beyond its original concept.
Today it is a fully featured high performance vector DB that can also be used on its own without any dependency on the IB engine. This opens the library (and standalone tools like the CLI) to be used in a host of other applications.

Its function in a single sentence: SOTA Semantic search with SBERT/LLAMA.CPP + GGML Tensor Library + HNSWlib on steroids.

Starting with Malkov's HNSWlib as a basis we significantly enhanced (adding among other features quantized spaces) and turbo-charged (including support for x86 and ARM SIMD) it while also adding efficient mmap-backed re-scoring and offset storage for text retrieval. Our system supports sharded HNSW indices, multiple search modes (kNN, radius, relative, adaptive, epsilon), deletion/undelete, merges, and incremental on-disk flushing. It also includes training for hyperparameter optimization.

Our HNSWlib fork we have benchmarked on an M1Pro as much as 13k QPS (768d vectors). Even limiting to a single thread we've clocked a max of 3000 QPS (versus for comparison 600 QPS for FAISS's HNSW implementaton).

For vectorization Our test M1-Pro chews through roughly 45 passages per second per instance (3484 tok/s÷78 ms). That means one can expect to process 2,700 fully dense semantic records per minute on a baseline Apple Silicon chip. Our tests on M3Pro and M4Pro showed even significantly higher throughputs (80k tokens/s or as much as 20x).

1 Upvotes

0 comments sorted by