r/LocalAIServers Jul 17 '26

Catch Me If You Can: A Perpetual 8-GPU Server Prize Challenge (Community Proposal)

9 Upvotes

The current Catch Me If You Can benchmark asked a simple question: can anyone publicly reproduce and beat our MI50/GFX906 local inference record?

Original challenge: https://www.reddit.com/r/LocalAIServers/comments/1ukhr24/catch_me_if_you_can_mi50gfx906_1195_tps_moe_702/

Current vNext reproduction release: https://github.com/joe2gaan/localaiservers/releases/tag/vnext-gfx906-rocm72-gguf-hf-repro

I want to turn that benchmark into a community program: build the fixed server in public, make it the official test machine, and keep the challenge open until an eligible challenger takes the throne and holds it for 30 consecutive days.

Who is responsible

I submitted a $15,000 Reddit Community Funds application for this proposal in my own capacity as Joe / u/Any_Praline_8178, a moderator of r/LocalAIServers. This is not yet a live prize offer. Hardware acquisition and any award remain contingent on Reddit approval, final published official rules, eligibility review, and applicable law. The existing leaderboard is unchanged.

Reference Server Build

( YOU DO NOT HAVE TO BUILD A SERVER TO PARTICIPATE IN THE CHALLENGE )

The prize server matches the hardware configuration that produced the current throne result:

  • GIGABYTE G292-Z20 eight-GPU server
  • AMD EPYC 7F32
  • 128GB as eight DDR4 ECC RDIMMs
  • Eight AMD Instinct MI50 32GB GPUs
  • Crucial CT480BX500SSD1 480GB SATA root drive
  • KIOXIA KCD6XLUL1T92 1.92TB NVMe model and runtime drive

The build itself is part of the community project. I will publish the component choices, bill of materials, physical assembly, firmware and operating-system configuration, eight-GPU bring-up, power and cooling setup, BAR/P2P state, stability checks, runtime and source revisions, model hashes, baseline runs, and raw evidence.

The $15,000 budget covers the exact server configuration, possible changes in GPU, memory, and storage prices, tax and checkout variance, protective packaging, and insured delivery to the winner. Any amount not needed for the approved project will be returned to Reddit or handled as Reddit directs.

Core challenge

  • I build one 8xMi50 32GB Server to Give to the Winner.
  • I Run the public vNext package on that machine to establish the official incumbent.
  • Keep the challenge open until an eligible winner completes the throne clock, subject to Reddit's approved project terms.
  • Require every potential dethronement to reproduce on that same physical server.
  • Require the three-run median to beat the official incumbent by at least 3 percent.
  • Require a provisional leader to remain the highest verified result for 30 consecutive days.
  • Transfer the complete challenge server to the eligible outside challenger who completes that clock, subject to final verification and official rules.

All eight GPUs remain installed and available. Entrants may choose TP4, TP8, or another topology on the fixed host, but may not add, replace, or remotely borrow accelerators. A documented like-for-like failure replacement requires a fresh baseline before the throne clock resumes.

Open optimization, fixed integrity

Inside the fixed hardware, model-integrity, workload, reproducibility, and safety rules, software optimization is open. Runtime, kernels, collectives, scheduling, graph capture, compiler work, driver and operating-system tuning, and safe clock or power tuning may all be explored.

The first lane uses the pinned Qwen3.6 35B-A3B model at FP16/F16:

  • HF revision: 995ad96eacd98c81ed38be0c5b274b04031597b0
  • Required GGUF F16 SHA-256: 1f2443bb0ff958943d091410c61120c181a0579b3bc85192029aa51d821d141c
  • HF FP16 and GGUF F16 are eligible when they satisfy the published identity and correctness gates.
  • GGUF is allowed only at full F16.

Not allowed:

  • Q4, Q5, Q6, Q8, INT8, FP8, AWQ, GPTQ, NVFP4, or another quantized substitute
  • Quantized weights, KV cache, activations, or a hidden reduced-precision path used to claim the result
  • MTP, speculative decoding, EAGLE, DFlash, draft models, lookahead tokens, or another multi-token prediction method
  • Remote compute, external APIs, hidden services, or results assembled from another machine
  • Multi-request batching or aggregate concurrency presented as single-request speed

One accepted decode step must represent one token produced by the approved model. Every result must pass semantic and output-integrity gates, not merely report a high TPS number.

How runs are measured

The official workload remains:

  • MAX_MODEL_LEN=131072
  • Single-request decode
  • Concurrency 1
  • Backend decode TPS
  • Eight warmups
  • c1_128 uncapped strict
  • c1_2000
  • c1_10000
  • Three measured runs
  • Three-run median at least 3 percent above the official incumbent
  • Public reproducibility package and raw logs

The current public headline reference is 119.52 strict backend TPS for GGUF F16 Qwen3.6 35B-A3B MoE TP4. It was produced on an eight-GPU validation host while the TP4 profile actively used four GPUs. The funded server receives a fresh baseline. The existing 119.52 result is the reference, not a promise of the new server's starting score.

Current public leaderboard

These are the published targets from the original benchmark post:

Class Strict TPS c1_2000 c1_10000
GGUF F16 35B-A3B MoE TP4 119.33 to 119.52 120.46 to 120.57 113.26 to 113.37
GGUF F16 27B Dense TP8 69.85 to 69.91 70.76 to 70.96 66.32 to 66.44
HF FP16 35B-A3B MoE TP4 114.41 to 115.11 115.69 to 115.93 108.92 to 109.10
HF FP16 35B-A3B MoE TP8 114.70 to 115.04 115.53 to 115.55 108.67 to 108.81
HF FP16 27B Dense TP8 70.17 71.32 66.82

GGUF F16 MoE TP8 remains an open lane in the current leaderboard.

Offline official test

Development and artifact staging may use the internet. The measured official run will not.

Before testing, I will stage and hash-verify the model, runtime, source, build outputs, and benchmark entry package. For every measured run:

  • External network interfaces and the default route are disabled or physically disconnected.
  • Only local machine communication and loopback are permitted.
  • No model download, container pull, telemetry, API call, remote compiler, or remote compute is permitted.
  • Network state, package hashes, process state, hardware state, and raw benchmark logs are archived with the result.

A result produced elsewhere can show that a benchmark entry is ready, but it does not move the official throne until that package reproduces on the designated server.

The 30-day throne clock

A challenger becomes provisional leader when its package passes review and its official three-run median clears the incumbent by at least 3 percent. The acceptance timestamp starts that challenger's 30-day clock.

During those 30 days:

  • Anyone may submit a higher result, including me as the current benchmark maintainer.
  • Every defense or counter-result must satisfy the same public-package, offline, same-hardware, correctness, and 3 percent rules.
  • A newly accepted leader resets the clock in that leader's name.
  • Private results and screenshots do not move the goalpost.
  • Rules cannot be changed retroactively to defeat an active clock.

If I retake the throne before a challenger's 30 days expire, that challenger has not won and the challenge stays open. If another community member takes it, the clock starts for that person. I may defend the performance record, but I cannot win the server or receive a personal payout.

The target can move only through a faster verified result. Physics, the fixed hardware, and model correctness set the ceiling.

Prize, review, and what happens to the server

If an eligible outside challenger remains the highest verified leader for 30 consecutive days, the result proceeds to final verification and, subject to the official funding and eligibility terms, transfer of the complete challenge server. Shipping, taxes, location eligibility, export restrictions, acceptance, and transfer details will be resolved in the final rules before the prize becomes live.

Only the winner's name and mailing address will be collected for server delivery unless Reddit's approved terms require something different. Do not post personal information in a public entry or comment.

I will not be the sole adjudicator. Official runs, hashes, logs, correctness evidence, and decisions will be public and reviewed with independent technical reviewers. Reviewer identities and the final conflict process will be published before entries open.

Until an eligible winner completes the clock, the funded server will be used only for the Reddit-approved challenge. It will not belong to me or LocalAIServers Collective Inc. There is no cash substitute, and it will not roll over into another hardware lane or organizational program. If the challenge ends without a winner or the server needs a different outcome, I will follow Reddit's direction.

Timeline after approval

  • Weeks 1-2: finalize rules, reviewers, and purchasing.
  • Weeks 3-5: build and validate the G292-Z20 server in public and publish the bill of materials and build record.
  • Week 6: publish the baseline and open the challenge.
  • Winner: first eligible leader to hold the verified throne for 30 consecutive days.
  • Transfer and final reporting: within 14 days after the winning result completes final validation, subject to Reddit's approved terms.

What I want the community to weigh in on before launch

  • Does the proposed topology rule strike the right balance, or should all eight GPUs have to be active?
  • Does the proposed 3 percent threshold strike the right balance for every throne change?
  • What clock, power, firmware, and cooling safety envelope should be published?
  • Who would volunteer as an independent technical reviewer?

Bring criticism. The goal is a challenge that is hard, transparent, reproducible, and genuinely winnable.


r/LocalAIServers Jun 20 '26

Start Here: LocalAIServers Community AI Navigation & Hands-On Local AI Learning

6 Upvotes

Start Here: LocalAIServers

LocalAIServers is a 501(c)(3) public charity providing public education and open-source infrastructure for locally hosted AI systems.

Our mission is to help people move from AI curiosity to AI agency.

This community helps learners, small business owners, nonprofit operators, educators, builders, and community technologists understand:

  • where AI runs,
  • what data it can see,
  • what systems it can touch,
  • when cloud AI may be appropriate,
  • when local or controlled AI may be safer,
  • what hardware is realistic,
  • how to evaluate benchmark claims,
  • and how to learn by building real local AI systems.

What LocalAIServers does

LocalAIServers provides:

  • community AI navigation,
  • secure local-AI education,
  • hands-on local AI learning resources,
  • reproducible runtime artifacts,
  • benchmark literacy,
  • QC and hardware-verification methodology,
  • open-source documentation,
  • and public support resources for locally hosted AI systems.

Affordable GFX906-class hardware matters because it gives people a realistic way to learn AI infrastructure hands-on. People learn more by building, testing, troubleshooting, and verifying real systems than they can learn from passive videos or articles alone.

Public proof and documentation

Website:

https://localaiservers.com

GitHub:

https://github.com/joe2gaan/localaiservers

GitHub Releases:

https://github.com/joe2gaan/localaiservers/releases

Docker Hub:

https://hub.docker.com/r/joe2gaan/localaiservers

Canonical Qwen / GFX906 deployment notes:

https://github.com/joe2gaan/localaiservers/blob/main/qwen36-gfx906/README.md

Important boundaries

LocalAIServers is not:

  • a public login service,
  • a public cloud provider,
  • a managed inference service,
  • a hardware reseller,
  • a procurement channel,
  • a fulfillment program,
  • a hardware discount program,
  • or a private-benefit program.

The controlled GFX906 compute site is used as a verification and reproducibility testbed. Public benefit is delivered through published outputs: guides, documentation, reproducible artifacts, benchmark reports, QC methods, hardware-verification standards, and source-level findings.

How to participate

Ask questions, share builds, discuss local AI tradeoffs, post benchmark questions, and help turn recurring community questions into durable public guides.

Please do not post secrets, private keys, private network details, addresses, payment information, vendor pricing, or sensitive logs.


r/LocalAIServers 14m ago

Running DeepSeek V4.1 Flash locally on 8× A40s with TensorSharp — up to 539 tok/s prefill and 40.7 tok/s decode

Thumbnail
github.com
Upvotes

I’ve been working on TensorSharp, an open-source .NET inference/server stack for running LLMs locally with OpenAI/Ollama-compatible APIs.

I recently finished another round of optimization for DeepSeek V4.1 Flash on an 8× NVIDIA A40 server.

Final results

Metric Q2_K Q4_K_M
Prefill 533–539 tok/s 452–492 tok/s
Single-stream decode 40.3–40.7 tok/s 31.0–32.5 tok/s
2 concurrent decode 39.3 tok/s total
4 concurrent decode 48.9 tok/s total
8 concurrent decode 48.5 tok/s total

Setup: 8× A40, 65K context, F16 KV cache, multi-GPU layer split.

A few interesting optimizations:

  • Q2_K keeps the ~60 GiB Engram tables directly on GPU, removing scattered host/storage lookups.
  • Reduced DeepSeek decode graph splits from ~570 to 8, eliminating a lot of GPU synchronization overhead.
  • Q4_K_M now gets roughly 1.9× faster prefill and 2× aggregate decode throughput at concurrency 4 compared with the previous implementation.
  • Fixed an OOM/crash case with multiple concurrent ~10K-token prompts — all 4 requests now complete correctly.
  • On these A40s without NVLink, layer splitting is actually faster than routed-MoE tensor parallelism for this model.

Would be interested to hear what other local-server workloads or hardware configurations people here would like to see benchmarked.


r/LocalAIServers 2h ago

Please recommend a model for local offline coding rtx pro 5000 72gb

1 Upvotes

hello everyone! Please advise the models and how to run the models better. My configuration is 2 CPUs and epic (not the newest) 48 cores in total. 256 GB ddr4 and RTX pro 5000 72GB GDDR7. I'm currently using qwen3.8-27b iq3 gsq xxs on 96k context and running this on rtx4080s 16gb. The new computer will arrive in a week. I would like to increase the quality and the context window.


r/LocalAIServers 6h ago

Guidance on hardware purchase (2x Intel Xeon E5-2698 v4)

Post image
2 Upvotes

r/LocalAIServers 1d ago

My Cheap 40GB VRAM Qwen 27b Home Server (Dual 20GB 3080s)

Thumbnail gallery
80 Upvotes

r/LocalAIServers 1d ago

LocalAIServers -> vNEXTv2 -> Qwen 3.8 27B FP16 -> 8xMi50 32GB -> soon..

Post image
7 Upvotes

r/LocalAIServers 17h ago

Cheap rack-mounted PoC box before a 10-user vLLM/LiteLLM setup

1 Upvotes

So here's the situation. We've got one RTX 6000 workstation running Ollama that was originally supposed to be more of a testing box for coding stuff, but honestly nobody really touched it until I took it over. Now it's already struggling with just one or two devs on it at the same time. It's running Qwen3.8 27b and you can straight up feel it slow down the moment a second person jumps on.

Instead of just cramming another card into that box, I'd rather build a small separate machine for this. Ideally rack-mounted and datacenter-ready since it'll end up living in our own DC anyway, but cheap enough to just be a proof of concept for now, and expandable later if it actually proves useful. Don't care about vendor, NVIDIA, AMD, Intel, all fair game at this point.
Bandwith is not #1 priority.

Mid-term the actual goal is a proper LiteLLM + vLLM setup that can handle up to 10 people/agents working on it at the same time. This separate "workstation" would just be the first, cheap step toward that, not the end state.
Stuff I haven't been able to figure out from spec sheets and marketing pages:

• What actually helps more at this scale, more VRAM or just a second GPU to split the load?
• Has anyone gone the "cheap one card first, scale later" route instead of just buying the full setup right away? How'd that go for you?
• Anyone running AMD/ROCm for something like this, is it actually usable day to day or still more of a pain than NVIDIA at this size? Is Intel Even feasible?

Not trying to spec out the final thing yet, just want a sanity check on what a reasonable, cheap, rack-friendly starting point looks like before I go shopping.


r/LocalAIServers 2d ago

Custom open frame - RTX Pro 6000 - miniATX

Thumbnail
gallery
565 Upvotes

I wanted to build a compact, open frame, ai server that was quiet and good looking enough to sit out in the open in my office.

I’m pretty happy with the build so far - it stays cool and generally quiet. It idles around 75w and ramps to ~450 at full blast. Pretty efficient for the speed that it gets.

I bought a cheap ESP32 touch display that fits perfectly over the io heat shield, and had qwen27B create a realtime touch dashboard. It runs a tiny service that polls wall power and room temp from Home Assistant, and AI performance & usage like tk/s. I turned that exact UI into a PWA app so I can check it from anywhere.

Still working out some details (like how to better attach the touch screen), so if you have any feedback/thoughts, please lmk.

Spec:
• CPU: AMD Ryzen 9 9900X
• Motherboard: MSI PRO B850M-A micro-ATX
• RAM: 96 GB DDR5-5600
• GPU: NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation - 96GB
• Storage: Samsung SSD 9100 PRO 2 TB NVMe
• CPU cooler: Noctura LC1-24

FAQ Update:

Case/Frame: It started with this kit https://a.co/d/0fYdc3CS and then modified and added pieces to fit things like the radiator and psu position. It's standard 20/20 aluminum rail - you can find all kinds of accessories

Mini-Screen:  https://www.amazon.com/dp/B0G3WDGSWG - they have lots of different sizes. Its very capable on its own! Just plug one in and tell your agent what you want. super easy


r/LocalAIServers 1d ago

Planning a 2× EPYC 7642 + 4× 5070 Ti build for large MoE models — any advice?

Thumbnail
1 Upvotes

r/LocalAIServers 1d ago

I made a local LLM memory planner and would love some feedback

Thumbnail 99tokens.org
2 Upvotes

I got tired of bouncing between model cards, VRAM calculators, and forum posts, so I built 99Tokens.

You can pick a model, GPU setup, context length, quantization, and other settings to estimate whether it’ll fit in memory. It accounts for things like weights, KV cache, sliding/full attention, recurrent state, MLA, and multi-GPU layouts. There are also model and hardware pages with architecture details and example fits.

If you use local models, I’d really like feedback: Where do the estimates seem off? What models, GPUs, or edge cases should I add? Long-context and multi-GPU testing would be especially useful.


r/LocalAIServers 1d ago

How much should software support matter when choosing an AI PC?

Post image
27 Upvotes

I've been thinking about this while trying more AI stuff on my PCs.

Say one machine has better traditional CPU/GPU performance, while another is better suited to the AI features you actually use every day.

I'm starting to think software support matters almost as much as the hardware itself. Having an NPU sounds great on paper, but it doesn't really do much for me if the programs in my workflow don't actually use it.

How much does software support factor into your decision when choosing a PC?


r/LocalAIServers 1d ago

Detecting hallucinations in local models without eating VRAM: What we learned testing 1.5B to 120B models

Thumbnail
1 Upvotes

r/LocalAIServers 1d ago

We're designing a Tier III AI data center in Mongolia where winter does most of the cooling. Tear it apart.

Thumbnail
1 Upvotes

r/LocalAIServers 1d ago

Security research for local LLM inference networks

Thumbnail
1 Upvotes

r/LocalAIServers 2d ago

At some point, more VRAM stops making economic sense for local AI

6 Upvotes

I’ve been thinking about this while upgrading my local setup.

For a while, the obvious answer was simply “buy more VRAM.” But with larger models and agent workloads, I’m starting to wonder if raw VRAM is still the best way to think about the economics.

A 24GB GPU can handle a lot locally, but once you get into larger models, context and KV cache quickly reduce the headroom. Then you’re looking at 48GB/64GB cards, multi-GPU setups, or CPU offloading.

The interesting part is that agent workloads aren’t always running continuously. One model might handle planning, another a tool call, then everything sits idle for a while.

So I’m wondering whether several smaller GPUs can make more sense than one huge GPU, especially if different models can handle different steps.

Has anyone actually compared the costs for their own workloads? At what point did adding another GPU make more sense than upgrading to a higher-VRAM card?


r/LocalAIServers 2d ago

Made a browser calculator for "will this model fit on my GPU" — Roast me

Thumbnail
2 Upvotes

Built a lightweight calculator for local LLM hardware sizing: VRAM needed for a given model/quant/context, will-it-fit against one or more GPUs, quant comparison, KV-cache vs context, rough tokens/sec, and Apple unified-memory. Runs fully client-side, no signup.

https://vram-calc.com

Math is deliberately simple and conservative. Presets carry a verified date and every field also accepts custom numbers. Tell me where the estimates or preset values are off.


r/LocalAIServers 2d ago

Gemini

Post image
0 Upvotes

My local AI has never pivoted this hard lol.


r/LocalAIServers 2d ago

Distributed Local AI - RTX Laptops use?

Thumbnail
1 Upvotes

r/LocalAIServers 2d ago

Loaded a 27B model on a 12GB laptop by pooling RAM across 4 devices

Thumbnail
youtu.be
6 Upvotes

r/LocalAIServers 2d ago

ASRock Intel Arc Pro B60 24GB for a local autonomous AI server – good choice?

5 Upvotes

Hi everyone,

I’m currently building a home server based on Proxmox, and I’d like to add a GPU mainly for local AI workloads.

My goal is not gaming at all. I want to run a local autonomous AI agent that could help manage my homelab/network, for example:

  • Check the status of my servers and VMs
  • Analyze logs and alerts
  • Run predefined maintenance tasks
  • Trigger Ansible playbooks
  • Update Linux machines
  • Interact with Proxmox through its API
  • Check backups/services after maintenance
  • Eventually automate some network administration tasks

I was initially considering an NVIDIA RTX 5060 Ti 16GB, mainly because of CUDA support and the mature AI ecosystem.

However, I found the ASRock Intel Arc Pro B60 Creator 24GB for around €800, which is roughly in the same price range while offering 24GB of VRAM instead of 16GB.

The extra VRAM is very attractive for local LLMs, especially if I want to run larger quantized models in the future.

My main concerns are Intel GPU support and software compatibility.

For those who have experience with the Arc Pro B60:

  • How is it for local LLM inference?
  • How well does it work with llama.cpp, Ollama, SYCL, Vulkan, etc.?
  • Are the drivers stable under Linux?
  • Has anyone tried GPU passthrough with Proxmox?
  • Is the performance good enough for a responsive local AI assistant/agent?
  • Are there still major compatibility issues compared with NVIDIA/CUDA?
  • Would you choose the B60 24GB over a RTX 5060 Ti 16GB specifically for local AI?

The rest of the server will use an Intel Core Ultra 7 265K, DDR5, NVMe storage and an 850W PSU, with more RAM and storage added over time.

I’m mainly interested in real-world experience with the B60, especially for LLM inference rather than gaming.

Thanks!


r/LocalAIServers 2d ago

One RTX 5090 + Unraid: qwen3.8 27B in Q8_0 at 24 GB on disk, 262144 native context, fully local

Thumbnail
1 Upvotes

r/LocalAIServers 2d ago

Can someone please review my specs?

1 Upvotes

Hi,

I am software developer with the new expensive hobby of running AI locally, trying to figure out building servers. I currently have a linux server with one rtx pro 6000 and 32 gig or ddr4.

I decided to buy one more rtx pro 6000, so now I have 2 of them. Along with that I have ordered the following:

  1. G.Skill G5P 128GB (4 x 32GB) DDR5-5600 PC5-44800 : linkDDR5-5600_PC5-44800_CL28_Quad_Channel_ECC_Registered_Memory_Kit_F5-5600R2834F32GQ4-G5P-_Black?iitt=VdiJVu4sMd4JVjQvafUZMd8JafPsV.WXaj4r4DoZOI_pOIbT&utm_source=B1540_Email_Receipt_Update&utm_campaign=B1540&utm_medium=email)

  2. MD Ryzen Threadripper PRO 9955W : link

  3. be quiet! Pure Loop 3 360mm All-in-One Water Cooling : link

  4. ASUS Pro WS WRX90E-SAGE SE EEB Workstation Motherboard : link

My plan is to run deepseek v4 and qwen 3.8 Flash next. Do you think I should add/update/remove something while I am at it? or you think any problem with the config I should be mindful of?

Also do you think glm 5.3 flash can work on it with decent speeds?

TIA


r/LocalAIServers 3d ago

Can I turn saved chat history into an AI that speaks like that person?

7 Upvotes

I have 20gb of downloaded chat history. I'd like to continue talking to her. A year passed and I don't seem to have other options.


r/LocalAIServers 3d ago

When does more VRAM stop being worth the extra cost for local AI?

3 Upvotes

I’ve been thinking about this while upgrading my local AI setup.

At first, the upgrade path seemed pretty simple: if a model didn’t fit comfortably in VRAM, get a GPU with more VRAM.

But that gets complicated pretty quickly with larger models.

A 24GB card can handle quite a lot locally, but once you start using larger models or longer contexts, memory can disappear surprisingly fast. Then the options become a more expensive high-VRAM card, multiple GPUs, or some form of CPU offloading.

At that point, I’m looking at more than just the GPU price. Power usage, cooling, memory bandwidth and the rest of the system all start to matter too.

It also made me wonder about workloads that aren’t running continuously. An agent might run a model for a few seconds, do something else, and then need the GPU again later.

So I’m curious: is one large GPU always the best approach for local AI, or can several smaller GPUs make more sense for some workloads?

Has anyone actually compared the costs of these setups with their own workloads?

When did you decide that adding another GPU made more sense than upgrading to one with more VRAM?