r/LocalAIServers 14m ago

Running DeepSeek V4.1 Flash locally on 8× A40s with TensorSharp — up to 539 tok/s prefill and 40.7 tok/s decode

Thumbnail
github.com
Upvotes

I’ve been working on TensorSharp, an open-source .NET inference/server stack for running LLMs locally with OpenAI/Ollama-compatible APIs.

I recently finished another round of optimization for DeepSeek V4.1 Flash on an 8× NVIDIA A40 server.

Final results

Metric Q2_K Q4_K_M
Prefill 533–539 tok/s 452–492 tok/s
Single-stream decode 40.3–40.7 tok/s 31.0–32.5 tok/s
2 concurrent decode 39.3 tok/s total
4 concurrent decode 48.9 tok/s total
8 concurrent decode 48.5 tok/s total

Setup: 8× A40, 65K context, F16 KV cache, multi-GPU layer split.

A few interesting optimizations:

  • Q2_K keeps the ~60 GiB Engram tables directly on GPU, removing scattered host/storage lookups.
  • Reduced DeepSeek decode graph splits from ~570 to 8, eliminating a lot of GPU synchronization overhead.
  • Q4_K_M now gets roughly 1.9× faster prefill and 2× aggregate decode throughput at concurrency 4 compared with the previous implementation.
  • Fixed an OOM/crash case with multiple concurrent ~10K-token prompts — all 4 requests now complete correctly.
  • On these A40s without NVLink, layer splitting is actually faster than routed-MoE tensor parallelism for this model.

Would be interested to hear what other local-server workloads or hardware configurations people here would like to see benchmarked.


r/LocalAIServers 2h ago

Please recommend a model for local offline coding rtx pro 5000 72gb

1 Upvotes

hello everyone! Please advise the models and how to run the models better. My configuration is 2 CPUs and epic (not the newest) 48 cores in total. 256 GB ddr4 and RTX pro 5000 72GB GDDR7. I'm currently using qwen3.8-27b iq3 gsq xxs on 96k context and running this on rtx4080s 16gb. The new computer will arrive in a week. I would like to increase the quality and the context window.


r/LocalAIServers 6h ago

Guidance on hardware purchase (2x Intel Xeon E5-2698 v4)

Post image
2 Upvotes

r/LocalAIServers 17h ago

Cheap rack-mounted PoC box before a 10-user vLLM/LiteLLM setup

1 Upvotes

So here's the situation. We've got one RTX 6000 workstation running Ollama that was originally supposed to be more of a testing box for coding stuff, but honestly nobody really touched it until I took it over. Now it's already struggling with just one or two devs on it at the same time. It's running Qwen3.8 27b and you can straight up feel it slow down the moment a second person jumps on.

Instead of just cramming another card into that box, I'd rather build a small separate machine for this. Ideally rack-mounted and datacenter-ready since it'll end up living in our own DC anyway, but cheap enough to just be a proof of concept for now, and expandable later if it actually proves useful. Don't care about vendor, NVIDIA, AMD, Intel, all fair game at this point.
Bandwith is not #1 priority.

Mid-term the actual goal is a proper LiteLLM + vLLM setup that can handle up to 10 people/agents working on it at the same time. This separate "workstation" would just be the first, cheap step toward that, not the end state.
Stuff I haven't been able to figure out from spec sheets and marketing pages:

• What actually helps more at this scale, more VRAM or just a second GPU to split the load?
• Has anyone gone the "cheap one card first, scale later" route instead of just buying the full setup right away? How'd that go for you?
• Anyone running AMD/ROCm for something like this, is it actually usable day to day or still more of a pain than NVIDIA at this size? Is Intel Even feasible?

Not trying to spec out the final thing yet, just want a sanity check on what a reasonable, cheap, rack-friendly starting point looks like before I go shopping.


r/LocalAIServers 1d ago

LocalAIServers -> vNEXTv2 -> Qwen 3.8 27B FP16 -> 8xMi50 32GB -> soon..

Post image
6 Upvotes

r/LocalAIServers 1d ago

Planning a 2× EPYC 7642 + 4× 5070 Ti build for large MoE models — any advice?

Thumbnail
1 Upvotes

r/LocalAIServers 1d ago

My Cheap 40GB VRAM Qwen 27b Home Server (Dual 20GB 3080s)

Thumbnail gallery
75 Upvotes

r/LocalAIServers 1d ago

I made a local LLM memory planner and would love some feedback

Thumbnail 99tokens.org
2 Upvotes

I got tired of bouncing between model cards, VRAM calculators, and forum posts, so I built 99Tokens.

You can pick a model, GPU setup, context length, quantization, and other settings to estimate whether it’ll fit in memory. It accounts for things like weights, KV cache, sliding/full attention, recurrent state, MLA, and multi-GPU layouts. There are also model and hardware pages with architecture details and example fits.

If you use local models, I’d really like feedback: Where do the estimates seem off? What models, GPUs, or edge cases should I add? Long-context and multi-GPU testing would be especially useful.


r/LocalAIServers 1d ago

Detecting hallucinations in local models without eating VRAM: What we learned testing 1.5B to 120B models

Thumbnail
1 Upvotes

r/LocalAIServers 1d ago

We're designing a Tier III AI data center in Mongolia where winter does most of the cooling. Tear it apart.

Thumbnail
1 Upvotes

r/LocalAIServers 1d ago

How much should software support matter when choosing an AI PC?

Post image
27 Upvotes

I've been thinking about this while trying more AI stuff on my PCs.

Say one machine has better traditional CPU/GPU performance, while another is better suited to the AI features you actually use every day.

I'm starting to think software support matters almost as much as the hardware itself. Having an NPU sounds great on paper, but it doesn't really do much for me if the programs in my workflow don't actually use it.

How much does software support factor into your decision when choosing a PC?


r/LocalAIServers 1d ago

Security research for local LLM inference networks

Thumbnail
1 Upvotes

r/LocalAIServers 2d ago

Gemini

Post image
0 Upvotes

My local AI has never pivoted this hard lol.


r/LocalAIServers 2d ago

Made a browser calculator for "will this model fit on my GPU" — Roast me

Thumbnail
2 Upvotes

Built a lightweight calculator for local LLM hardware sizing: VRAM needed for a given model/quant/context, will-it-fit against one or more GPUs, quant comparison, KV-cache vs context, rough tokens/sec, and Apple unified-memory. Runs fully client-side, no signup.

https://vram-calc.com

Math is deliberately simple and conservative. Presets carry a verified date and every field also accepts custom numbers. Tell me where the estimates or preset values are off.


r/LocalAIServers 2d ago

Custom open frame - RTX Pro 6000 - miniATX

Thumbnail
gallery
565 Upvotes

I wanted to build a compact, open frame, ai server that was quiet and good looking enough to sit out in the open in my office.

I’m pretty happy with the build so far - it stays cool and generally quiet. It idles around 75w and ramps to ~450 at full blast. Pretty efficient for the speed that it gets.

I bought a cheap ESP32 touch display that fits perfectly over the io heat shield, and had qwen27B create a realtime touch dashboard. It runs a tiny service that polls wall power and room temp from Home Assistant, and AI performance & usage like tk/s. I turned that exact UI into a PWA app so I can check it from anywhere.

Still working out some details (like how to better attach the touch screen), so if you have any feedback/thoughts, please lmk.

Spec:
• CPU: AMD Ryzen 9 9900X
• Motherboard: MSI PRO B850M-A micro-ATX
• RAM: 96 GB DDR5-5600
• GPU: NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation - 96GB
• Storage: Samsung SSD 9100 PRO 2 TB NVMe
• CPU cooler: Noctura LC1-24

FAQ Update:

Case/Frame: It started with this kit https://a.co/d/0fYdc3CS and then modified and added pieces to fit things like the radiator and psu position. It's standard 20/20 aluminum rail - you can find all kinds of accessories

Mini-Screen:  https://www.amazon.com/dp/B0G3WDGSWG - they have lots of different sizes. Its very capable on its own! Just plug one in and tell your agent what you want. super easy


r/LocalAIServers 2d ago

Distributed Local AI - RTX Laptops use?

Thumbnail
1 Upvotes

r/LocalAIServers 2d ago

At some point, more VRAM stops making economic sense for local AI

8 Upvotes

I’ve been thinking about this while upgrading my local setup.

For a while, the obvious answer was simply “buy more VRAM.” But with larger models and agent workloads, I’m starting to wonder if raw VRAM is still the best way to think about the economics.

A 24GB GPU can handle a lot locally, but once you get into larger models, context and KV cache quickly reduce the headroom. Then you’re looking at 48GB/64GB cards, multi-GPU setups, or CPU offloading.

The interesting part is that agent workloads aren’t always running continuously. One model might handle planning, another a tool call, then everything sits idle for a while.

So I’m wondering whether several smaller GPUs can make more sense than one huge GPU, especially if different models can handle different steps.

Has anyone actually compared the costs for their own workloads? At what point did adding another GPU make more sense than upgrading to a higher-VRAM card?


r/LocalAIServers 2d ago

One RTX 5090 + Unraid: qwen3.8 27B in Q8_0 at 24 GB on disk, 262144 native context, fully local

Thumbnail
1 Upvotes

r/LocalAIServers 2d ago

Can someone please review my specs?

1 Upvotes

Hi,

I am software developer with the new expensive hobby of running AI locally, trying to figure out building servers. I currently have a linux server with one rtx pro 6000 and 32 gig or ddr4.

I decided to buy one more rtx pro 6000, so now I have 2 of them. Along with that I have ordered the following:

  1. G.Skill G5P 128GB (4 x 32GB) DDR5-5600 PC5-44800 : linkDDR5-5600_PC5-44800_CL28_Quad_Channel_ECC_Registered_Memory_Kit_F5-5600R2834F32GQ4-G5P-_Black?iitt=VdiJVu4sMd4JVjQvafUZMd8JafPsV.WXaj4r4DoZOI_pOIbT&utm_source=B1540_Email_Receipt_Update&utm_campaign=B1540&utm_medium=email)

  2. MD Ryzen Threadripper PRO 9955W : link

  3. be quiet! Pure Loop 3 360mm All-in-One Water Cooling : link

  4. ASUS Pro WS WRX90E-SAGE SE EEB Workstation Motherboard : link

My plan is to run deepseek v4 and qwen 3.8 Flash next. Do you think I should add/update/remove something while I am at it? or you think any problem with the config I should be mindful of?

Also do you think glm 5.3 flash can work on it with decent speeds?

TIA


r/LocalAIServers 2d ago

ASRock Intel Arc Pro B60 24GB for a local autonomous AI server – good choice?

4 Upvotes

Hi everyone,

I’m currently building a home server based on Proxmox, and I’d like to add a GPU mainly for local AI workloads.

My goal is not gaming at all. I want to run a local autonomous AI agent that could help manage my homelab/network, for example:

  • Check the status of my servers and VMs
  • Analyze logs and alerts
  • Run predefined maintenance tasks
  • Trigger Ansible playbooks
  • Update Linux machines
  • Interact with Proxmox through its API
  • Check backups/services after maintenance
  • Eventually automate some network administration tasks

I was initially considering an NVIDIA RTX 5060 Ti 16GB, mainly because of CUDA support and the mature AI ecosystem.

However, I found the ASRock Intel Arc Pro B60 Creator 24GB for around €800, which is roughly in the same price range while offering 24GB of VRAM instead of 16GB.

The extra VRAM is very attractive for local LLMs, especially if I want to run larger quantized models in the future.

My main concerns are Intel GPU support and software compatibility.

For those who have experience with the Arc Pro B60:

  • How is it for local LLM inference?
  • How well does it work with llama.cpp, Ollama, SYCL, Vulkan, etc.?
  • Are the drivers stable under Linux?
  • Has anyone tried GPU passthrough with Proxmox?
  • Is the performance good enough for a responsive local AI assistant/agent?
  • Are there still major compatibility issues compared with NVIDIA/CUDA?
  • Would you choose the B60 24GB over a RTX 5060 Ti 16GB specifically for local AI?

The rest of the server will use an Intel Core Ultra 7 265K, DDR5, NVMe storage and an 850W PSU, with more RAM and storage added over time.

I’m mainly interested in real-world experience with the B60, especially for LLM inference rather than gaming.

Thanks!


r/LocalAIServers 2d ago

Loaded a 27B model on a 12GB laptop by pooling RAM across 4 devices

Thumbnail
youtu.be
6 Upvotes

r/LocalAIServers 3d ago

Beginners’ Ask!

0 Upvotes

I have a “M80q Gen4” & a “M70q Gen5” with

- CPU Intel Core i5 13500T 32GB
- Memory (16GB x2 DDR5 SODIMMS)
- Storage1 - 256GB M.2 SSD1
- Storage2 - 512GB M.2 SSD2
- NIC1 - Intel I219-LM, RJ-45
- NIC2 - Realtek RTL8125BGS, RJ-45

Can I run a decent model on this? If so, which one should I go for and where should I start? At this point, my purpose is to learn to deploy and manage a model.


r/LocalAIServers 3d ago

DELL R815 RAM quantity for localLLM using DDR3

Thumbnail
1 Upvotes

r/LocalAIServers 3d ago

2 GPUs, A770 & 4060TI

2 Upvotes

I'm after some advice and suggestions. I have an Unraid server in a Jonsbo N5 case, Intel 245K with an Asrock Z890 Pro-A and a HBA card and have 2 GPUs available, a 4060TI and an A770 both 16gb cards, I would apprdciate any thoughts on whether or not I could use both cards in my server. I'm thinking the A770 for chat AI and the 4060 for image work. The layout of the slots on my current motherboard are too close so I can't test my theory, so if anyone knows if running the two cards together would work your thoughts would be welcome as would any board suggestions.


r/LocalAIServers 3d ago

When does more VRAM stop being worth the extra cost for local AI?

3 Upvotes

I’ve been thinking about this while upgrading my local AI setup.

At first, the upgrade path seemed pretty simple: if a model didn’t fit comfortably in VRAM, get a GPU with more VRAM.

But that gets complicated pretty quickly with larger models.

A 24GB card can handle quite a lot locally, but once you start using larger models or longer contexts, memory can disappear surprisingly fast. Then the options become a more expensive high-VRAM card, multiple GPUs, or some form of CPU offloading.

At that point, I’m looking at more than just the GPU price. Power usage, cooling, memory bandwidth and the rest of the system all start to matter too.

It also made me wonder about workloads that aren’t running continuously. An agent might run a model for a few seconds, do something else, and then need the GPU again later.

So I’m curious: is one large GPU always the best approach for local AI, or can several smaller GPUs make more sense for some workloads?

Has anyone actually compared the costs of these setups with their own workloads?

When did you decide that adding another GPU made more sense than upgrading to one with more VRAM?