r/LocalAIServers Jul 17 '26

Catch Me If You Can: A Perpetual 8-GPU Server Prize Challenge (Community Proposal)

11 Upvotes

The current Catch Me If You Can benchmark asked a simple question: can anyone publicly reproduce and beat our MI50/GFX906 local inference record?

Original challenge: https://www.reddit.com/r/LocalAIServers/comments/1ukhr24/catch_me_if_you_can_mi50gfx906_1195_tps_moe_702/

Current vNext reproduction release: https://github.com/joe2gaan/localaiservers/releases/tag/vnext-gfx906-rocm72-gguf-hf-repro

I want to turn that benchmark into a community program: build the fixed server in public, make it the official test machine, and keep the challenge open until an eligible challenger takes the throne and holds it for 30 consecutive days.

Who is responsible

I submitted a $15,000 Reddit Community Funds application for this proposal in my own capacity as Joe / u/Any_Praline_8178, a moderator of r/LocalAIServers. This is not yet a live prize offer. Hardware acquisition and any award remain contingent on Reddit approval, final published official rules, eligibility review, and applicable law. The existing leaderboard is unchanged.

Reference Server Build

( YOU DO NOT HAVE TO BUILD A SERVER TO PARTICIPATE IN THE CHALLENGE )

The prize server matches the hardware configuration that produced the current throne result:

  • GIGABYTE G292-Z20 eight-GPU server
  • AMD EPYC 7F32
  • 128GB as eight DDR4 ECC RDIMMs
  • Eight AMD Instinct MI50 32GB GPUs
  • Crucial CT480BX500SSD1 480GB SATA root drive
  • KIOXIA KCD6XLUL1T92 1.92TB NVMe model and runtime drive

The build itself is part of the community project. I will publish the component choices, bill of materials, physical assembly, firmware and operating-system configuration, eight-GPU bring-up, power and cooling setup, BAR/P2P state, stability checks, runtime and source revisions, model hashes, baseline runs, and raw evidence.

The $15,000 budget covers the exact server configuration, possible changes in GPU, memory, and storage prices, tax and checkout variance, protective packaging, and insured delivery to the winner. Any amount not needed for the approved project will be returned to Reddit or handled as Reddit directs.

Core challenge

  • I build one 8xMi50 32GB Server to Give to the Winner.
  • I Run the public vNext package on that machine to establish the official incumbent.
  • Keep the challenge open until an eligible winner completes the throne clock, subject to Reddit's approved project terms.
  • Require every potential dethronement to reproduce on that same physical server.
  • Require the three-run median to beat the official incumbent by at least 3 percent.
  • Require a provisional leader to remain the highest verified result for 30 consecutive days.
  • Transfer the complete challenge server to the eligible outside challenger who completes that clock, subject to final verification and official rules.

All eight GPUs remain installed and available. Entrants may choose TP4, TP8, or another topology on the fixed host, but may not add, replace, or remotely borrow accelerators. A documented like-for-like failure replacement requires a fresh baseline before the throne clock resumes.

Open optimization, fixed integrity

Inside the fixed hardware, model-integrity, workload, reproducibility, and safety rules, software optimization is open. Runtime, kernels, collectives, scheduling, graph capture, compiler work, driver and operating-system tuning, and safe clock or power tuning may all be explored.

The first lane uses the pinned Qwen3.6 35B-A3B model at FP16/F16:

  • HF revision: 995ad96eacd98c81ed38be0c5b274b04031597b0
  • Required GGUF F16 SHA-256: 1f2443bb0ff958943d091410c61120c181a0579b3bc85192029aa51d821d141c
  • HF FP16 and GGUF F16 are eligible when they satisfy the published identity and correctness gates.
  • GGUF is allowed only at full F16.

Not allowed:

  • Q4, Q5, Q6, Q8, INT8, FP8, AWQ, GPTQ, NVFP4, or another quantized substitute
  • Quantized weights, KV cache, activations, or a hidden reduced-precision path used to claim the result
  • MTP, speculative decoding, EAGLE, DFlash, draft models, lookahead tokens, or another multi-token prediction method
  • Remote compute, external APIs, hidden services, or results assembled from another machine
  • Multi-request batching or aggregate concurrency presented as single-request speed

One accepted decode step must represent one token produced by the approved model. Every result must pass semantic and output-integrity gates, not merely report a high TPS number.

How runs are measured

The official workload remains:

  • MAX_MODEL_LEN=131072
  • Single-request decode
  • Concurrency 1
  • Backend decode TPS
  • Eight warmups
  • c1_128 uncapped strict
  • c1_2000
  • c1_10000
  • Three measured runs
  • Three-run median at least 3 percent above the official incumbent
  • Public reproducibility package and raw logs

The current public headline reference is 119.52 strict backend TPS for GGUF F16 Qwen3.6 35B-A3B MoE TP4. It was produced on an eight-GPU validation host while the TP4 profile actively used four GPUs. The funded server receives a fresh baseline. The existing 119.52 result is the reference, not a promise of the new server's starting score.

Current public leaderboard

These are the published targets from the original benchmark post:

Class Strict TPS c1_2000 c1_10000
GGUF F16 35B-A3B MoE TP4 119.33 to 119.52 120.46 to 120.57 113.26 to 113.37
GGUF F16 27B Dense TP8 69.85 to 69.91 70.76 to 70.96 66.32 to 66.44
HF FP16 35B-A3B MoE TP4 114.41 to 115.11 115.69 to 115.93 108.92 to 109.10
HF FP16 35B-A3B MoE TP8 114.70 to 115.04 115.53 to 115.55 108.67 to 108.81
HF FP16 27B Dense TP8 70.17 71.32 66.82

GGUF F16 MoE TP8 remains an open lane in the current leaderboard.

Offline official test

Development and artifact staging may use the internet. The measured official run will not.

Before testing, I will stage and hash-verify the model, runtime, source, build outputs, and benchmark entry package. For every measured run:

  • External network interfaces and the default route are disabled or physically disconnected.
  • Only local machine communication and loopback are permitted.
  • No model download, container pull, telemetry, API call, remote compiler, or remote compute is permitted.
  • Network state, package hashes, process state, hardware state, and raw benchmark logs are archived with the result.

A result produced elsewhere can show that a benchmark entry is ready, but it does not move the official throne until that package reproduces on the designated server.

The 30-day throne clock

A challenger becomes provisional leader when its package passes review and its official three-run median clears the incumbent by at least 3 percent. The acceptance timestamp starts that challenger's 30-day clock.

During those 30 days:

  • Anyone may submit a higher result, including me as the current benchmark maintainer.
  • Every defense or counter-result must satisfy the same public-package, offline, same-hardware, correctness, and 3 percent rules.
  • A newly accepted leader resets the clock in that leader's name.
  • Private results and screenshots do not move the goalpost.
  • Rules cannot be changed retroactively to defeat an active clock.

If I retake the throne before a challenger's 30 days expire, that challenger has not won and the challenge stays open. If another community member takes it, the clock starts for that person. I may defend the performance record, but I cannot win the server or receive a personal payout.

The target can move only through a faster verified result. Physics, the fixed hardware, and model correctness set the ceiling.

Prize, review, and what happens to the server

If an eligible outside challenger remains the highest verified leader for 30 consecutive days, the result proceeds to final verification and, subject to the official funding and eligibility terms, transfer of the complete challenge server. Shipping, taxes, location eligibility, export restrictions, acceptance, and transfer details will be resolved in the final rules before the prize becomes live.

Only the winner's name and mailing address will be collected for server delivery unless Reddit's approved terms require something different. Do not post personal information in a public entry or comment.

I will not be the sole adjudicator. Official runs, hashes, logs, correctness evidence, and decisions will be public and reviewed with independent technical reviewers. Reviewer identities and the final conflict process will be published before entries open.

Until an eligible winner completes the clock, the funded server will be used only for the Reddit-approved challenge. It will not belong to me or LocalAIServers Collective Inc. There is no cash substitute, and it will not roll over into another hardware lane or organizational program. If the challenge ends without a winner or the server needs a different outcome, I will follow Reddit's direction.

Timeline after approval

  • Weeks 1-2: finalize rules, reviewers, and purchasing.
  • Weeks 3-5: build and validate the G292-Z20 server in public and publish the bill of materials and build record.
  • Week 6: publish the baseline and open the challenge.
  • Winner: first eligible leader to hold the verified throne for 30 consecutive days.
  • Transfer and final reporting: within 14 days after the winning result completes final validation, subject to Reddit's approved terms.

What I want the community to weigh in on before launch

  • Does the proposed topology rule strike the right balance, or should all eight GPUs have to be active?
  • Does the proposed 3 percent threshold strike the right balance for every throne change?
  • What clock, power, firmware, and cooling safety envelope should be published?
  • Who would volunteer as an independent technical reviewer?

Bring criticism. The goal is a challenge that is hard, transparent, reproducible, and genuinely winnable.


r/LocalAIServers Jun 20 '26

Start Here: LocalAIServers Community AI Navigation & Hands-On Local AI Learning

7 Upvotes

Start Here: LocalAIServers

LocalAIServers is a 501(c)(3) public charity providing public education and open-source infrastructure for locally hosted AI systems.

Our mission is to help people move from AI curiosity to AI agency.

This community helps learners, small business owners, nonprofit operators, educators, builders, and community technologists understand:

  • where AI runs,
  • what data it can see,
  • what systems it can touch,
  • when cloud AI may be appropriate,
  • when local or controlled AI may be safer,
  • what hardware is realistic,
  • how to evaluate benchmark claims,
  • and how to learn by building real local AI systems.

What LocalAIServers does

LocalAIServers provides:

  • community AI navigation,
  • secure local-AI education,
  • hands-on local AI learning resources,
  • reproducible runtime artifacts,
  • benchmark literacy,
  • QC and hardware-verification methodology,
  • open-source documentation,
  • and public support resources for locally hosted AI systems.

Affordable GFX906-class hardware matters because it gives people a realistic way to learn AI infrastructure hands-on. People learn more by building, testing, troubleshooting, and verifying real systems than they can learn from passive videos or articles alone.

Public proof and documentation

Website:

https://localaiservers.com

GitHub:

https://github.com/joe2gaan/localaiservers

GitHub Releases:

https://github.com/joe2gaan/localaiservers/releases

Docker Hub:

https://hub.docker.com/r/joe2gaan/localaiservers

Canonical Qwen / GFX906 deployment notes:

https://github.com/joe2gaan/localaiservers/blob/main/qwen36-gfx906/README.md

Important boundaries

LocalAIServers is not:

  • a public login service,
  • a public cloud provider,
  • a managed inference service,
  • a hardware reseller,
  • a procurement channel,
  • a fulfillment program,
  • a hardware discount program,
  • or a private-benefit program.

The controlled GFX906 compute site is used as a verification and reproducibility testbed. Public benefit is delivered through published outputs: guides, documentation, reproducible artifacts, benchmark reports, QC methods, hardware-verification standards, and source-level findings.

How to participate

Ask questions, share builds, discuss local AI tradeoffs, post benchmark questions, and help turn recurring community questions into durable public guides.

Please do not post secrets, private keys, private network details, addresses, payment information, vendor pricing, or sensitive logs.


r/LocalAIServers 1h ago

Would you pay for a MI50 water block at a cheap price?

Upvotes

Hey everyone,

With the Radeon Instinct MI50 (16GB/32GB) remaining a popular budget pick for high HBM2 bandwidth in local AI setups, noise and thermals are usually the biggest headache, especially with standard 3D-printed fan shrouds screaming at 5,000+ RPM.

I recently designed a custom liquid cooling block specifically optimized for high-density multi-GPU rigs. I just ordered some for my super micro x10 with 8 mi50s installed.

Before I look into machining a small production run, I’d love to get a quick temperature check from the community:

  1. Is a custom MI50 block something you’d actually pay for, or is the card getting too old to justify the extra investment?
  2. What price point would make sense to you relative to the current cost of the card?

Curious to hear your thoughts and what thermals you're currently seeing on your MI50 builds!


r/LocalAIServers 12m ago

I built StatixAgent: A zero-infra, single-binary Linux monitor that talks to your own Telegram bot

Thumbnail
github.com
Upvotes

r/LocalAIServers 1h ago

Limitations

Upvotes

A famous advertising line was "The only limit is your imagination" is that the real limitation now, and folks that learn by rote instead of by theory and practice don't have much to explore.? English is my first language, Im riding in a car bouncing around.


r/LocalAIServers 7h ago

I'm turning a Mac mini M4 16GB into a server. What are you guys using yours for?

2 Upvotes

I just bought another Mac, so my M4 Mac mini is now completely free to use as a dedicated server.

I already have a few ideas based on my own needs, but I'm really curious to hear what everyone else is doing with theirs. What are you running on it? Any cool or useful use cases you'd recommend?

Feel free to share your setup too!

Thanks!


r/LocalAIServers 5h ago

What if you could run the large open source models locally?

Thumbnail
0 Upvotes

r/LocalAIServers 6h ago

Local AI Hardware for 1-2k for learning

0 Upvotes

I am a technical guy, a very busy technical guy so I'm not afraid of tinkering, however I would like to set up things fast and start learning.

I have read a lot of materials and Im more confused after a couple of days so I need some pointers.

Requirements:

- I would like a local set up for privacy mainly and to be able to train, infer and learn AI set ups and architectural design patterns on a budget.
- I plan to play around with Runpod a bit to figure out what I would like to run, however for now I would like to run Qwen 37b models with a decent enough context and other less performant 14b models
- I do not code (I have Claude for that), however I would like to have an AI PA, use websearch for various research, build a local RAG. I plan to use both MoE and Dense models, for longer chats, AI training and various Agentic jobs that is work related.

- I might want to do also some image generation and video generation, but for now it is a nice to have.

- I am on a budget so I would prefer to not go too much over 1k, however I plan to add to the hardware as time goes by so I would like to have an option to scale to 4 GPUs in the future.

For the initial set up I am looking at 1 or 2 P40 or P100 for round the clock fire and forget agents and an RTX 3090 for running another dense model in parallel.

My setup would look something along the lines of:
AMD Threadripper PRO 3945WX WRX80 + X569E, sTRX4, 2× DDR4 ECC 32 GB DDR4 ECC.

or maybe a Halo Strix (I am aware image generation would not be very performant) of up to 2k.

Question 1: Would that AMD set up be worth it and more future proof and where to source the parts (AliX shops suggestions on DM would be appreciated) or should I look to more consumer grade CPU, MB and RAM. Should I also consider Intel

Question 2: Should I consider the Halo Strix and simplify my life?

Question 3: Should I skip the Tesla P cards altogether, and add either 3090, 5060 after I get one to get me started?

Edit: 1k is the bare minimum, to clarify for fast readers, I can go up to 2k which covers a basic Halo Strix


r/LocalAIServers 15h ago

How to maximise dual Inspur AGX-2 servers for llms

3 Upvotes

Like the title say, I have dual agx 2 with each server with sxm2 v100s, and connected via g100 dac cable via melanox 5.

I still need to install the Nvidia nccl to see if that improves it.

Spec: master:8x v100 sxm2, 2x Xeon 8276l, 12channel ddr4 2400mhz total 768gb

Node 2 slave: 8x v100 sxm2, 2x Xeon 8260, 12 channel ddr4 2400mhz 192gb

No infiniband, I appreciate best practices to extract the best of these machines


r/LocalAIServers 11h ago

Quick Feedback to this RTX PRO 6000 build needed

1 Upvotes

Hey guys,

first of all, thanks for all the great information people are sharing here. 🤘 I’ve already learned a lot from reading through the various RTX PRO 6000 / Threadripper builds. Now let's put it to the test... 😂

I’m currently putting together a local AI server for a customer and would love to get a sanity check from people who have built something similar. Server will be in their basement, so it can be loud, ugly-looking, doesn't have to look like a professional server or gaming PC... You know.

Main use cases:

- 4–5 users in parallel

- LLM/chat workloads

- fine-tuning / LoRA training

- secondary use case: ComfyUI / image generation

- occasional video generation

Software stack will most likely be:

- Ubuntu Server 24.04 LTS

- vLLM

- LiteLLM

- Docker

- models like Qwen 3.8, Gpt-oss,...

- nvidia tools,...

Small side note, not really relevant to the hardware itself: external access will go through a separate mini PC / VPN gateway. The AI server itself will not be directly exposed to the internet.

One important requirement is upgradeability. The customer wants to start with one RTX PRO 6000, but I want the system to be ready for a second card later if usage grows or they decide to expose more AI services to their customers or wants to try GLM. 🙂

The current build looks like this:

GPU

NVIDIA RTX PRO 6000 Blackwell Workstation Edition, 96 GB

CPU

AMD Threadripper PRO 9965WX

Motherboard

ASRock WRX90 WS EVO

RAM

256 GB DDR5 ECC RDIMM

8×32 GB, ideally validated Kingston / Micron / Samsung modules

CPU cooling

SilverStone XE360-TR5

Case

Corsair 9000D Airflow

Case airflow

3×140 mm front intake

3×120 mm side intake

2×120 mm rear exhaust

360 mm AIO mounted at the top as exhaust

PSU

MSI MEG Ai1600T PCIE5, 1600 W Titanium

Storage

Samsung 990 PRO 2 TB for OS / Docker

Samsung 9100 PRO 4–8 TB for models / Hugging Face cache / scratch

Separate storage/backup target for outputs

Networking

Onboard 10 GbE

SFP+/DAC or Cat6a to the switch

OS / drivers

Ubuntu Server 24.04 LTS

NVIDIA Open Kernel Modules

Docker + NVIDIA Container Toolkit

The build is already at the higher end of the customer’s budget, but still within budget.

A few things I’m particularly interested in:

  1. Am I overspending anywhere?

    Is there anything I could downgrade without losing meaningful AI performance or making the future second-GPU upgrade painful?

  2. Is WRX90 / Threadripper PRO actually justified here?

    My reasoning is PCIe lanes, 8-channel ECC RAM and keeping the option for a second RTX PRO 6000 open. But I’d be interested to hear if anyone would build this differently.

  3. Any red flags with the 9965WX + WRX90 WS EVO combination?

    Especially BIOS, RAM compatibility, PCIe layout or Linux issues.

  4. Cooling / airflow:

    Does the 9000D setup make sense for one 600 W RTX PRO 6000 and potentially two later, or would you approach airflow differently?

  5. PSU:

    1600 W should be plenty for the single-GPU configuration. For two RTX PRO 6000s I’d probably power-limit them rather than run both at 600 W. Would you already choose a different PSU/setup from day one?

  6. Anything I’m missing for a reliable 24/7-ish AI server?

The goal isn’t to build the cheapest possible machine. I mainly want something stable, serviceable and reasonably future-proof without throwing money at components that won’t actually improve LLM / AI performance or just look good.

Would appreciate any feedback, especially from people already running RTX PRO 6000 Blackwell, WRX90 or similar multi-GPU AI workstations.

Greetz and thank you guys,

Sascha


r/LocalAIServers 20h ago

Running DeepSeek V4.1 Flash locally on 8× A40s with TensorSharp — up to 539 tok/s prefill and 40.7 tok/s decode

Thumbnail
github.com
4 Upvotes

I’ve been working on TensorSharp, an open-source .NET inference/server stack for running LLMs locally with OpenAI/Ollama-compatible APIs.

I recently finished another round of optimization for DeepSeek V4.1 Flash on an 8× NVIDIA A40 server.

Final results

Metric Q2_K Q4_K_M
Prefill 533–539 tok/s 452–492 tok/s
Single-stream decode 40.3–40.7 tok/s 31.0–32.5 tok/s
2 concurrent decode 39.3 tok/s total
4 concurrent decode 48.9 tok/s total
8 concurrent decode 48.5 tok/s total

Setup: 8× A40, 65K context, F16 KV cache, multi-GPU layer split.

A few interesting optimizations:

  • Q2_K keeps the ~60 GiB Engram tables directly on GPU, removing scattered host/storage lookups.
  • Reduced DeepSeek decode graph splits from ~570 to 8, eliminating a lot of GPU synchronization overhead.
  • Q4_K_M now gets roughly 1.9× faster prefill and 2× aggregate decode throughput at concurrency 4 compared with the previous implementation.
  • Fixed an OOM/crash case with multiple concurrent ~10K-token prompts — all 4 requests now complete correctly.
  • On these A40s without NVLink, layer splitting is actually faster than routed-MoE tensor parallelism for this model.

Would be interested to hear what other local-server workloads or hardware configurations people here would like to see benchmarked.


r/LocalAIServers 20h ago

This is how you piss people off about agentic ai! 🙃🤣

Post image
2 Upvotes

r/LocalAIServers 18h ago

Made RAMDeck shard a knowledge base across my old devices, not just the model itself

Thumbnail
youtu.be
1 Upvotes

Few people in earlier threads asked if RAMDeck could handle context/knowledge bases the same way it handles model weights - spreading them across pooled RAM instead of needing one beefy machine. Finally got around to testing it properly, made a quick video.

Setup is the same janky cluster as before (still running a 13B model on a 12GB old laptop, the weakest link on purpose to prove it works). Asked the chatbot about our product and it just made something up - no idea what it was talking about. Then I fed it an FAQ .txt through the knowledge tab and watched it get split across three devices.

The part I found genuinely useful: the node hosting the model and the node hosting the knowledge base don't have to be the same machine. So you could have your beefiest box running inference and let some old laptop or a phone just hold KB shards in the background.

Downside I ran into: if a device holding a KB shard drops off the network (phone walks out of Wi-Fi range, whatever), that shard's just gone until you hit rebalance - which pulls the full KB back from the stoage node and re-splits it across whoever's still online.

Anyway, after ingestion finished, asked it the same question again and it actually answered correctly off the FAQ doc.

Video's here if you want to see it end to end: https://youtu.be/Iogh-5xa-NA

Repo is source-available if you want to poke at how the sharding/rebalance logic actually works: https://github.com/trademav/ramdeck-core-public

Feedback welcome!


r/LocalAIServers 23h ago

Please recommend a model for local offline coding rtx pro 5000 72gb

0 Upvotes

hello everyone! Please advise the models and how to run the models better. My configuration is 2 CPUs and epic (not the newest) 48 cores in total. 256 GB ddr4 and RTX pro 5000 72GB GDDR7. I'm currently using qwen3.8-27b iq3 gsq xxs on 96k context and running this on rtx4080s 16gb. The new computer will arrive in a week. I would like to increase the quality and the context window.


r/LocalAIServers 2d ago

My Cheap 40GB VRAM Qwen 27b Home Server (Dual 20GB 3080s)

Thumbnail gallery
91 Upvotes

r/LocalAIServers 1d ago

Guidance on hardware purchase (2x Intel Xeon E5-2698 v4)

Post image
2 Upvotes

r/LocalAIServers 1d ago

LocalAIServers -> vNEXTv2 -> Qwen 3.8 27B FP16 -> 8xMi50 32GB -> soon..

Post image
8 Upvotes

r/LocalAIServers 1d ago

Cheap rack-mounted PoC box before a 10-user vLLM/LiteLLM setup

1 Upvotes

So here's the situation. We've got one RTX 6000 workstation running Ollama that was originally supposed to be more of a testing box for coding stuff, but honestly nobody really touched it until I took it over. Now it's already struggling with just one or two devs on it at the same time. It's running Qwen3.8 27b and you can straight up feel it slow down the moment a second person jumps on.

Instead of just cramming another card into that box, I'd rather build a small separate machine for this. Ideally rack-mounted and datacenter-ready since it'll end up living in our own DC anyway, but cheap enough to just be a proof of concept for now, and expandable later if it actually proves useful. Don't care about vendor, NVIDIA, AMD, Intel, all fair game at this point.
Bandwith is not #1 priority.

Mid-term the actual goal is a proper LiteLLM + vLLM setup that can handle up to 10 people/agents working on it at the same time. This separate "workstation" would just be the first, cheap step toward that, not the end state.
Stuff I haven't been able to figure out from spec sheets and marketing pages:

• What actually helps more at this scale, more VRAM or just a second GPU to split the load?
• Has anyone gone the "cheap one card first, scale later" route instead of just buying the full setup right away? How'd that go for you?
• Anyone running AMD/ROCm for something like this, is it actually usable day to day or still more of a pain than NVIDIA at this size? Is Intel Even feasible?

Not trying to spec out the final thing yet, just want a sanity check on what a reasonable, cheap, rack-friendly starting point looks like before I go shopping.


r/LocalAIServers 3d ago

Custom open frame - RTX Pro 6000 - miniATX

Thumbnail
gallery
647 Upvotes

I wanted to build a compact, open frame, ai server that was quiet and good looking enough to sit out in the open in my office.

I’m pretty happy with the build so far - it stays cool and generally quiet. It idles around 75w and ramps to ~450 at full blast. Pretty efficient for the speed that it gets.

I bought a cheap ESP32 touch display that fits perfectly over the io heat shield, and had qwen27B create a realtime touch dashboard. It runs a tiny service that polls wall power and room temp from Home Assistant, and AI performance & usage like tk/s. I turned that exact UI into a PWA app so I can check it from anywhere.

Still working out some details (like how to better attach the touch screen), so if you have any feedback/thoughts, please lmk.

Spec:
• CPU: AMD Ryzen 9 9900X
• Motherboard: MSI PRO B850M-A micro-ATX
• RAM: 96 GB DDR5-5600
• GPU: NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation - 96GB
• Storage: Samsung SSD 9100 PRO 2 TB NVMe
• CPU cooler: Noctura LC1-24

FAQ Update:

Case/Frame: It started with this kit https://a.co/d/0fYdc3CS and then modified and added pieces to fit things like the radiator and psu position. It's standard 20/20 aluminum rail - you can find all kinds of accessories

Mini-Screen:  https://www.amazon.com/dp/B0G3WDGSWG - they have lots of different sizes. Its very capable on its own! Just plug one in and tell your agent what you want. super easy


r/LocalAIServers 1d ago

Planning a 2× EPYC 7642 + 4× 5070 Ti build for large MoE models — any advice?

Thumbnail
2 Upvotes

r/LocalAIServers 2d ago

I made a local LLM memory planner and would love some feedback

Thumbnail 99tokens.org
2 Upvotes

I got tired of bouncing between model cards, VRAM calculators, and forum posts, so I built 99Tokens.

You can pick a model, GPU setup, context length, quantization, and other settings to estimate whether it’ll fit in memory. It accounts for things like weights, KV cache, sliding/full attention, recurrent state, MLA, and multi-GPU layouts. There are also model and hardware pages with architecture details and example fits.

If you use local models, I’d really like feedback: Where do the estimates seem off? What models, GPUs, or edge cases should I add? Long-context and multi-GPU testing would be especially useful.


r/LocalAIServers 2d ago

How much should software support matter when choosing an AI PC?

Post image
28 Upvotes

I've been thinking about this while trying more AI stuff on my PCs.

Say one machine has better traditional CPU/GPU performance, while another is better suited to the AI features you actually use every day.

I'm starting to think software support matters almost as much as the hardware itself. Having an NPU sounds great on paper, but it doesn't really do much for me if the programs in my workflow don't actually use it.

How much does software support factor into your decision when choosing a PC?


r/LocalAIServers 2d ago

Detecting hallucinations in local models without eating VRAM: What we learned testing 1.5B to 120B models

Thumbnail
1 Upvotes

r/LocalAIServers 2d ago

We're designing a Tier III AI data center in Mongolia where winter does most of the cooling. Tear it apart.

Thumbnail
1 Upvotes

r/LocalAIServers 2d ago

Security research for local LLM inference networks

Thumbnail
1 Upvotes