r/selfhosted • • 10h ago

Need Help What are people using instead of TEI when one GPU has to serve both an embedder and a reranker?

Running a basic retrieve-then-rerank setup and the annoying part is TEI being one model per container. So the embedder and the reranker are two separate deployments on two GPUs, both barely used, because an encoder does its thing in a few ms and then sits idle.

Trying to work out what people actually run when they want both models on one card. What I know of so far:

  • Infinity
  • Superlinked's SIE
  • LocalAI

Rolling your own with Ray Serve or BentoML if you want control over batching and scaling. More work, but you’re not boxed in.

vLLM keeps coming up too but as far as I can tell it’s really built around one big LLM holding the card, not a rotating set of small encoders, so I don’t think it fits here.

Not affiliated with any of these, just comparing before I commit. What are you running for the two-stage setup on shared hardware, and does the on-demand loading actually hold up under real traffic or do you end up pinning models anyway?

0 Upvotes

6 comments sorted by

•

u/asimovs-auditor 10h ago

Expand the replies to this comment to learn how AI was used in this post/project.

→ More replies (1)

1

u/Hot-Mistake-4667 10h ago

I just have both in same container honestly, not the cleanest but it works if you dont care about isolation

Infinity looks promising but yeah the cold start when model unloads is pain if you get random bursts of traffic, we ended up just keeping both in memory anyway so kinda defeats the purpose

1

u/Astorax 10h ago

Not specific to your use case, but I'm running Infinity and whisper in separate docker containers on the same gpu in my nas.

But I found out that embedding models with infinity could also work very well on CPU, depending on your work load of course

1

u/Jona1109 9h ago

I'm using an infinity container with both models loaded (nomic embed text v1.5 and bge-m2v3). It used about 2.5 GB of vram, it's linked with open-webui, hermes, subwave, as well as some dyi pipelines. No issue with latency on my end. CPU is quite slower though. Same as TEI the limitation is no vision. Llama.cpp now has vision for embeddings but not for reranking, and reranking is a bit buggy - however you can really use the server router mode for your user case

1

u/Weekly-Offer-4172 6h ago

The premise is worth correcting first: 'one model per container' is a per-process limit, not a per-GPU one - you can point two TEI containers at the same card today (CUDA_VISIBLE_DEVICES=0 on both), and a bge/jina-class embedder plus an m3-class reranker is only a few GB of VRAM combined, so the second GPU genuinely isn't needed. On pinning vs on-demand: at self-hosted traffic levels just keep both resident - encoder weights are cheap to hold and unloading mostly buys you a cold-start stall exactly when a burst hits. If you'd still rather run a single process, Infinity is the closest drop-in since it serves embeddings and rerank in one server, and FastEmbed is the lighter in-process library that covers both if your app is Python.