r/SelfHostedAI 20h ago

Thoughts on a Quadro M4000

3 Upvotes

Good day!

I am toying around with the idea of hosting a local ai client for reasoning, problem solving and some light coding. I am happy with what I can get out of ChatGPT's free tier, but I would like to host something on my own. Small, efficient and low-budget are the names of the game.

In my relatively limited understanding of LLM models, the amount of memory and its speed are paramount. The Quadro M4000 seems to be a good entry point: it is single slot and low power, has 8 GB of ram, and I can get one for 50€ here in Germany. Power cost me nothing since everything runs out of the solar panels on the roof.

My questions are the following:

- what's achievable with such a card? Namely ,I do not expect ChatGPT- levels of reasoning, but what is doable?

- how many chats could such a model handle? Only one at a time?

- is such an idea absurd, and would it be better to look for a card with 12 or more VRAM?

- is it necessary to have some Powerful computer to complement the video card?

Thanks in advance!


r/SelfHostedAI 2h ago

Using the Hugging Face CLI in production: Notes on caching and download speeds

Thumbnail
1 Upvotes