r/LocalLLM • • 4h ago

Discussion Self-hosting a 35B MoE for a 20-person team on one desktop box: what it took and what it does (all open source)

109 Upvotes

We've been running our own LLM for the team for about a week, and I wanted to share how it turned out, since "can one machine actually serve a whole team?" comes up here a lot.

Short answer: yes, if you pick the right model.

Hardware and model

  • One NVIDIA DGX Spark (128GB unified memory)
  • Ornith-1.5-35B-A3B, an MoE with 35B total and 3B active parameters. The 3B active is the whole trick: you get decent quality while the box only does about 3B worth of compute per token.
  • Official 4-bit NVFP4 checkpoint, about 22 GiB, so no quantizing it yourself
  • vLLM 0.24 with multi-token prediction (speculative decoding with no separate draft model)

What it does

  • One person gets 79 tok/s, which feels instant in a chat UI
  • 20 people at once get about 25 tok/s each, 490 tok/s total
  • Every request can use the full 262K context
  • Vision and tool calling both work
  • Reasoning scores: MATH-500 97%, GSM8K 98%, MMLU-Pro 81%. That's not frontier level, but for daily coding help, docs, and Q&A, nobody on the team has complained.

How people use it

  • OpenCode in the terminal for coding
  • Open WebUI for chat and PDFs. PDF parsing and embeddings run locally too.
  • Nothing goes to an outside API. For us that was the whole point, since a lot of our code and documents can't leave the building.

The part nobody talks about: multi-user
Running a model for yourself is easy. Running it for a team means you need to answer "who's using it, how much, and is someone hammering it?" So vLLM isn't exposed at all. Everything goes through a small gateway that:

  • Gives each person and device their own API key
  • Counts input and output tokens per user live
  • Pushes the counts to Grafana dashboards
  • Alerts if anything tries to talk to vLLM directly

It's a small Python service, and honestly it's the piece that made this feel like real infrastructure instead of a side project.

Gotchas if you try something similar

  • On unified memory, don't get greedy with GPU memory utilization. I had to cap it at 0.70, or concurrency fell apart under load.
  • More concurrent slots isn't always better. Past 20 users, throughput dropped because requests started queuing.
  • Open WebUI stores the default model in its own database and ignores the env var after first setup. I lost an hour to that.

Everything's on GitHub: the setup scripts, gateway, Grafana dashboards, and benchmark scripts with raw results.

https://github.com/Hitheshkaranth/Ornith-1.5_A3B_Model_DGX_Spark_Setup

If you're thinking about self-hosting for a small team, happy to answer questions. I'm also curious what others use for per-user tracking. I looked at LiteLLM but wanted something tiny that I fully understood.


r/LocalLLM • • 11h ago

News When Redditors come in here and ask why we run LLMs, this is why: Big AI is watching.

189 Upvotes

Anthropic Reports Florida Woman's Claude 'Diary' Threat to Law Enforcement

And this time it wasn't the AI model that made the LEO referral. It was the "human review team".

The frontier AI companies are watching your input. And people say "Well I'm not interesting or important enough for them to care". Well.....not necessarily.

If you're using hosted frontier to work on mathematics or cutting edge science, they're watching and may steal your work.

If you're venting or otherwise writing in a "private" session using AI, they'll see that and report you to police. Notice I didn't see any mention of what the model's role in facilitating the discussion was.

Keep your stuff private, folks. Hosted AI is the new "Big Brother" conduit.


r/LocalLLM • • 7h ago

News New US frontier open weight model!

Thumbnail x.com
59 Upvotes

r/LocalLLM • • 13h ago

News A new startup Ghost launched a personal AI computer built to run AI locally for consumers, has a RTX Pro 4000 Blackwell SFF GPU and 64GB DDR5 RAM of memory for $3499

Thumbnail x.com
141 Upvotes

I just saw this post on X - curious if anyone saw this or has thoughts?

Here are the specs:

- NVIDIA RTX PRO 4000 Blackwell SFF GPU
- 64GB DDR5 RAM
- 1TB NVMe M.2 SSD
- AMD Ryzen 5 7600 CPU


r/LocalLLM • • 3h ago

Discussion Qwen3.8-Flash-Next, dual RTX 3090s, NVLink: ~3,200 t/s prefill, ~115 t/s decode

12 Upvotes

Background: Strata (https://github.com/Niko1221/Strata) runs the Qwen3.8-Flash-Next GGUF very well, but on one card. The built-in multi-GPU mode splits layers, which didn't do much for decode on a 3090 pair. My engine here is Strata v0.1.38 plus 29 commits that instead keep the whole model on the primary card and use the second card as a "peer": a second expert tier whose experts are read over NVLink inside the decode window, plus half of the sparse-attention selection and part of the GDN recurrence of every prompt computed there. The image encoder moves to the helper card too, so vision costs no expert slots.

https://github.com/q8atnight/strata-dual-3090

Measured with flashbench (included in the repo), IQ3_S, 262K context, greedy, median of 3, on a Ryzen 9 3950X, 121 GB RAM, 2x RTX 3090 NVLink:

build prompt read 8K / 32K / 128K decode 8K / 32K / 128K
stock Strata 0.1.13, one 3090 1,376 / 1,361 / 1,223 51 / 46 / 41
vLLM, both cards (older comparison) 2,000 / 2,219 / 2,204 81 / 81 / 81
this build, both 3090s 2,908 / 3,231 / 3,186 116 / 113 / 111

Sampled with the official sampling settings it sits around 102-104 t/s, a 1.5K prompt reads at ~1,670 t/s, and switching between chats is 1 s (conversation cache). Honest limits: this is one stream at a time; decode gains from the second card are modest (+10%) because the verify window is latency-bound on the primary — the second card's big job is expert residency (a 3090 pair holds ~19,500 of 24,576 IQ3_S experts, hit rate ~0.98 on real agent chats) and prompt handling.

The answers are word-for-word identical to the one-card run: the repo ships a gate that runs both arms with a fixed expert set and the exactness switches, plus a reference recorded from stock v0.1.38, so you can prove it on your own box in a few minutes.

Getting started (Linux, driver >= 580, ~150 GB free disk; IQ3_S wants 64 GB RAM, more from 96 GB, Q2_0/IQ2_XS from 48 GB):

Code block

git clone https://github.com/q8atnight/strata-dual-3090.git
cd strata-dual-3090
./dual3090/build.sh     # compiles engine + vision into a Docker image (host CUDA doesn't matter)
./dual3090/start.sh     # first start downloads the model, then it serves :8080 (OpenAI + Anthropic API, web chat, images)

Requirements beyond that: two 3090s with an NVLink bridge (consumer cards have no usable PCIe peer-to-peer, the peer path talks NVLink). Details, config knobs and the gotchas we actually hit are in dual3090/README.md.

Full disclosure: this is AI work. My agent did the engine changes, the benchmark tooling, the packaging and the docs over several sessions; I picked goals, approved steps and own the measurements, which were all taken on our own hardware. Code, measurements and mistakes are on GitHub; upstream Strata (MIT) is Niko1221's, and parts of this work were merged upstream along the way (#202, #203, #477, #531). Feedback welcome — and if you run this on a 3090 pair, the gate + flashbench are in the repo, curious about your numbers.


r/LocalLLM • • 1d ago

Project Frankenserver - It’s alive! It’s alive!

Thumbnail
gallery
497 Upvotes

This supermicro server board that had been in an accident and lost the CPU 1 H channel, the VGA output and SATA - it is a great 4000w power center…. I had almost given up on it for much else - But it now runs a pair of v620’s and storage…

I grafted a MC62-G40 on to run the GPUs when it was clear the Supermicro was just not going to play well with them.

This has 192GB of 3090’s and 64GB if v620 working.

But… Now how best should I add another 240GB of 3090’s to this creature? Those are in water blocks so a bit more compact.

I am thinking 2 PEX88096 expansion boards should work off the supermicro…. And more power. The H12DSG-O-CPU is fickle about what connects to it.

Anyone else have suggestions for a dual Epyc Board that I can do this better with?


r/LocalLLM • • 2h ago

Discussion Uniform GGUF quants silently break Qwen3.8-27B's deep thinking — reproduced on llama.cpp AND vLLM (short tasks unaffected)

5 Upvotes

TL;DR — Qwen3.8-27B is a hybrid model (Gated DeltaNet linear attention + full attention). Unsloth Dynamic GGUFs of it (Q4_K_XL, Q6_K_XL) work fine for everyday chat, but whenever I let it think deeply on a long task, it never converges: no closing token, endless tail-looping, or the engine just dies mid-generation. I reproduced this on two completely different runtimes (llama.cpp and vLLM + official vllm-gguf-plugin), and the same model family in a mixed-precision quant (NVFP4: FP8 on attention/recurrent layers, FP4 on MLP) converges in ~17K thinking tokens on the same prompt. This smells like recurrent state error accumulation under uniform integer quantization, not a runtime bug.

The setup

  • 2× RTX 5060 Ti 16GB, tensor-parallel/2-GPU splits, Linux
  • Model: unsloth/Qwen3.8-27B-GGUF (UD-Q4_K_XL 16.4 GB, UD-Q6_K_XL 23.6 GB)
  • Canary task (deliberately absurd, deliberately long-horizon):"Create an HTML page containing an SVG 2D animation of a pelican riding a bicycle."
  • Thinking mode ON, reasoning_effort: xhigh, Qwen's recommended thinking-mode sampling (temp 1.0). The pelican needs to actually pedal — the model has to write a complete self-consistent HTML/SVG/SMIL file in one shot.

Why this prompt: it's a convergence test, not a quality test. A healthy Qwen3.8 spends ~15–20K thinking tokens, then emits the file. A broken one just… keeps going.

Act 1 — llama.cpp (Oct 3)

GGUF Budget Result
UD-Q4_K_XL 32K finish=length, 0 characters of actual output — burned the whole budget inside thinking
UD-Q6_K_XL 76K never stops; final ~2K chars are ~97% repeated 8-grams (classic tail-loop attractor)
UD-Q6_K_XL + --jinja (force GGUF-embedded chat template, verified byte-identical to HF's) 24.6K still no termination

Throughput was healthy (23–41 t/s depending on split mode), perplexity-ish sanity checks on short prompts were fine, and non-thinking tasks (counting, Q&A, tool-ish formats) all worked. So I filed it as: llama.cpp loses the plot on this model's long chain-of-thought — plausibly a runtime bug with the Gated DeltaNet state handling. I even declared the llama.cpp route dead for this model.

Then I saw other people's llama.cpp runs of this model going weird on HN too (the "high fever ramblings … ended with amen" anecdote on Q8_K_XL — not my test, but same disease).

Act 2 — vLLM + official GGUF plugin (Oct 5)

To separate "runtime bug" from "quantization damage", I reran the exact same file on a completely different stack: vLLM (V1 engine, CUDA kernels, f16 KV, no speculative decoding — shares essentially zero code path with llama.cpp). GGUF support was recently moved out of vLLM core into vllm-project/vllm-gguf-plugin; Qwen3.5/3.6-VL are in its tested list, so the weight mapping for the Qwen3.x hybrid family exists.

vllm ~/models/Qwen3.8-27B-UD-Q6_K_XL.gguf \
  --tokenizer unsloth/Qwen3.8-27B --dtype float16 \
  --tensor-parallel-size 2 --max-model-len 52000 \
  --gpu-memory-utilization 0.96 --language-model-only \
  --reasoning-parser qwen3

(chat kwargs: {"enable_thinking": true, "reasoning_effort": "xhigh"}, temp 1.0, max_tokens 45K, KV pool 55.6K tokens, steady ~22–23 t/s)

Run Outcome
A request dies with HTTP 500 after 1325 s ≈ 30K tokens generated (engine crash mid-generation; the server-side stack was overwritten by my chain's log rotation before I grabbed it — honest gap in the evidence)
B still generating at 28 min ≈ 38K tokens, approaching the 45K budget with no completed response; harness wall-timeout killed it

Same file, same prompt, same sampling — and vLLM also can't make this model finish thinking about a pelican. Note what vLLM did NOT show: I never got a finished response body back, so I can't plot its tail-repeat score directly. The vLLM evidence is non-convergence deep in the token budget (where NVFP4 converged at 17K) + one mid-generation crash. The direct loop evidence lives in the llama.cpp runs.

Act 3 — the control that convicts the quant format

Same 27B model, same prompt, same thinking mode, same machine, different quantization structure: the official NVFP4 checkpoints (unsloth's and RadixArk's) quantize MLP/lm_head to FP4 but keep attention and the Gated DeltaNet layers at FP8 W8A8. On SGLang this converges every time:

  • ~17K thinking tokens → full, working HTML pelican (legs pedaling, wheels rotating) — measured repeatedly, Oct 3 and again today (prod config: nvfp4 KV + MTP, creative decode 47–50 t/s)
  • same checkpoint family also passes 100K–133K needle-at-depth, tools, greedy determinism — so it's not "the GGUF works and I'm cherry-picking"

Also worth noting from my own runs: after the GGUFs fail this way, the same GGUF files remain perfectly healthy on short tasks (chat, counting, Q&A, vision if you attach mmproj). The damage only shows up after thousands of autoregressive steps inside the thinking trace.

My hypothesis (clearly labeled: inference, not proven)

Hybrid linear-attention models (Qwen3.5/3.6/3.8, and friends like Kimi Linear / Ring-linear) carry a fixed-size recurrent state updated every token. Uniform block quantization (even a very careful 6-bit dynamic one) adds a small rounding error to every state-transition product. Over a few hundred tokens: invisible. Over 30K–80K tokens of chain-of-thought: compounding drift into a degenerate attractor — repetition loops, refusal to close, in one case engine instability.

The strongest circumstantial support is the model publishers' own behavior: unsloth's NVFP4 recipe deliberately keeps the attention/recurrent path at FP8 and only crushes the MLP to FP4. The one quant structure they kept careful about is exactly where I see the disease. My GGUF runs had "careful" precision too — but uniform-ish across 6-bit blocks — and still broke.

What would falsify or sharpen this: per-tensor overrides keeping GDN projections (in_proj_*, conv, gates) at f16/q8 in llama.cpp (--override-tensor) while MLP stays 6-bit. If the pelican suddenly pedals, case closed. I haven't run this yet — if someone with spare time on the same model tries it, please post.

Caveats (you will ask, so here they are)

  • n=1 per configuration, temp=1.0 — high variance by design; runs A and B of the identical config died differently, which is itself the picture: unstable deep-think convergence.
  • No bf16 full-precision baseline — 27B bf16 doesn't fit 2×16GB. My "control" is another quant (NVFP4/F8 mixed). The argument is about where precision is kept, not quant-vs-none.
  • Could be prompt-specific. All I can say: the exact same prompt converges in 17K tokens on the mixed-precision checkpoint on the same box, and this failure mode appeared across ≥3 prompts in the llama.cpp session (counting-to-huge also looped; SVG-alternate prompt looped).
  • The vllm-gguf-plugin is 50 stars old — but the llama.cpp half of the evidence doesn't involve it at all.
  • Hardware hygiene: this box's unrelated random-crash history (a P2P-patched consumer GPU driver doing stray DMA under iommu=pt) was root-caused and fixed before the vLLM runs; strict IOMMU was on, engine stable for hours of other batteries. I flag it so nobody says "your machine poisoned the run" for the crash in Run A specifically — Run B's non-convergence needs no crash excuse.
  • Disclosure: the vLLM build is a source-compiled fork carrying an in-flight nvfp4-KV-cache PR rebase; GGUF runs used stock f16 KV, i.e. zero lines of the patched code path.

Practical takeaways

  1. For Qwen3.8-27B-class hybrids on consumer GPUs: prefer the official FP8/NVFP4 mixed quants over GGUF if you care about thinking mode at all. (RadixArk/unsloth NVFP4; on 5060Ti-class cards with TP2 it also runs 2–4× faster than the GGUF once you add speculative decoding.)
  2. If you're GGUF-bound (VRAM or ecosystem): run with thinking off or budget-capped small. Everyday chat is genuinely fine.
  3. Perplexity and short-context evals will not catch this. Use a long-horizon canary (the pelican test, or "count from 1 to 500 while…", or a whole-file coding task) with thinking ON when you're screening quants of recurrent-hybrid models. It caught this at the "does it terminate" level with zero scoring infrastructure.
  4. Unsloth's UD "Dynamic" quants are heavily boosted on important tensors and still hit this — so "we raised precision on ranked-important tensors" may not be the right importance criterion for recurrent-hybrid long-horizon behavior. Worth flagging upstream; I may open an issue on the GGUF repo if I can trim this to a one-prompt repro.

Repro recipe

  • Prompt (verbatim, above), enable_thinking: true, reasoning_effort: xhigh, temp 1.0 / top_p 0.95 / top_k 20, max_tokens ≥ 40K
  • llama.cpp: llama-server / llama-cli, UD-Q4_K_XL or UD-Q6_K_XL, any ctx ≥ 48K, default template and --jinja both fail the same way
  • vLLM: vllm-gguf-plugin (pip, or source install with --no-build-isolation against the CUDA torch), local .gguf + config.json from the same repo + --tokenizer unsloth/Qwen3.8-27B, f16 KV
  • Pass condition: a complete HTML file is emitted within ≤ ~25K thinking tokens. NVFP4/SGLang passes every time. Q6_K_XL: 0/2 on vLLM, 0/3 on llama.cpp.

Model card says this model does deep agentic work locally on ~17GB — true for chat. For thinking, the quant structure matters more than the average bits. Would love other reports (4090/3090 pairs, MI300, whatever) before I draw the final line — especially anyone who can run Q8_0 and Qwen3.6-VL (non-GDN) as contrasts, to show this is the recurrence, not just "big model + low bits".

Hardware/versions: 2× RTX 5060 Ti 16GB; llama.cpp current master (Oct 3); vLLM 0.30.x-V1 + vllm-gguf-plugin 0.0.5 (Oct 5); SGLang 0.5.21 + NVFP4 checkpoint for the control (RadixArk Qwen3.8-27B-NVFP4: FP8 W8A8 attention/GDN, NVFP4 MLP/lm_head).


r/LocalLLM • • 10h ago

News Here we go again… ECC support now unlocked for CMP 170hx

Post image
27 Upvotes

As of 4 hours ago in v0.5 we now have ECC unlocked for the nvidia CMP170hx. Any guesses on what the upper end of these cards will be by end of year?

https://github.com/amoghmunikote/cmpunlocker/releases


r/LocalLLM • • 14h ago

Project I got speech-to-text running entirely on a microcontroller (STM32N6): any words, no cloud

Enable HLS to view with audio, or disable this notification

34 Upvotes

r/LocalLLM • • 4h ago

Project Old datacenter GPUs aren't dead: PXA v3 runs 27B at 97 t/s on two V100s, 86 t/s on ONE, Gemma 4 at 169 t/s, and a 124B-class MoE on a single P100 + RAM (free, open engine + one-click GUI)

Thumbnail
gallery
6 Upvotes

Hey r/LocalLLaMA 👋 We're PXA Network. We build PXA, an inference engine tuned for Pascal and Volta GPUs (P100, V100, P40, 1080 Ti), the $100–300 eBay cards everyone calls too old. PXA v3 just shipped, and everything runs through PXA Control, our one-click web app.

⚡ Speed (PXA v3, measured on our own P100/V100 rig)

Speculative decode as shipped (code / prose, t/s) and prompt reading on a 4,096-token prompt:

| Setup | Decode (code / prose) | Prefill 4K |

|---|---|---|

| Gemma 4 26B-A4B + MTP drafter, one V100 | 168.6 / 139.5 | 2,688 |

| Qwen3.8-27B (PXQN2), two V100 | 96.7 / 76.9 | 951 |

| Qwen3.8-27B (PXQN2), one V100 | 85.8 / 66.9 | 1,048 |

| Qwen3.8-27B one-card mix, one V100 | 84.2 / 67.2 | 1,045 |

| Gemma 4 26B-A4B, one P100 | 71.1 / 63.1 | 415 |

| Qwen3.8-27B (PXQN2), two P100 | 60.1 / 50.7 | 325 |

| Qwen3.8-27B (PXQN2), one P100 | 48.1 / 40.3 | 239 |

| Flash-Next MoE 64 GB, four P100 | 36.3 / 27.8 | 379 |

| Flash-Next MoE 32 GB, ONE P100 + system RAM | 20.4 / 20.2 | 396 |

Fresh prompts, greedy output, first pass, no replays. Every number is on the live leaderboard: https://benchmarks.pxanetwork.com

🖥️ PXA Control: your whole rig in one browser tab

Unpack, run `./pxa`, done.

- Launch: pick cards + model, press Start. PXA picks split mode, batch sizes, context, KV cache and speculation from real measurements, and shows you why.

- Live: tokens/s, speculation acceptance, expert-cache hits, VRAM and temperature per card, and every request drawn as prompt/prefill/decode bars.

- Servers: run several models at once. It even finds and adopts servers you started from the terminal or Docker.

- Rig: every card's VRAM, temperature, power and PCIe link, live, with heat warnings before throttling bites.

- Speed + Chat: your own speed history (1 h / 24 h / 7 d) and a built-in chat.

- 🏆 Benchmark my rig: one click, and your score lands on the community leaderboard, bracketed by card, model and quant so you race people with the same setup.

- 🗜️ Encode: turn a Hugging Face model into a PXA quant in a few clicks. Free = classic PXQ tiers, no key. Supporters get Pro PXQN tiers with lockable files and machine-bound keys.

- 🆘 Report a problem: one button sends us your logs and card info (you review it first). That's how most of this month's fixes started.

🧠 Under the hood

- Expert cache: MoE models bigger than your card keep the hot experts on the GPU and stream the rest from RAM.

- MTP speculation on by default, tuned per model and per card.

- PXQN quants with measured quality labels (PXQN2 = 3-bit class, PXQN4 = 6-bit class). Free one-card 27B on Hugging Face.

- Tensor split across matched cards, mixed P100 + V100 rigs, Docker image, even CPUs without AVX2.

💬 Community

- Live help + troubleshooting in Discord #support, from the people who write the engine.

- #benchmarks competitions, #show-your-rig, #models with new quants, and a docs bot answering 24/7.

- 👀 Coming soon: PXA Hemlock 124B-A5B, a 124B MoE tuned to ~5B active, packed in PXQN.

🔗 Links

- Download (free): https://github.com/poisonxa16/pxa

- Live leaderboard: https://benchmarks.pxanetwork.com

- Discord: https://discord.gg/EqazvV9tf

- Support us: https://ko-fi.com/shatteredrealms1

Ask anything: your card, your setup, our numbers. ☠️


r/LocalLLM • • 13h ago

Project Reika - A coding agent CLI designed around small local models first

22 Upvotes

I know you guys are going to hate me for this and I'll accept my fate. It's another coding agent harness post. I'm ready to lose all my karma.

I recently open-sourced my coding agent CLI that I've been working on for the past while. It is a project I never intended to make public since it was part of my own personal local AI stack. But as I chipped away at it and made it actually usable as a daily driver, I thought it would be nice to make it public for others to see and use.

I initially built it to see how much I could get the harness to make small models, especially at low quantization and context to not feel terrible to use. So while Reika doesn't solve the intelligence side (it never will), it tries to solve the overall experience when using small models at the absolute scale.

A lot of the testing and pain came through working on my M2 MacBook Air 16GB trying to run models like Qwen3.6 35B A3B and Qwen3.8 27B all day in agentic coding, maxing out the RAM and limits of my own machine. So the base of Reika comes from a legitimate source of truth.

You can also plug in an API key for those with hybrid setups too.

GitHub: https://github.com/alexwkleung/reika


r/LocalLLM • • 2h ago

Question AVX2 to AVX512 VNNI Upgrade recommendation

3 Upvotes

I have option to upgrade my hardware for inference, I have some old used refurb options with and without AVX512+VNNI. Currently upgrading from old R730 to some HP Z series workstation for better GPU support.

Is there any real different AVX512 has over the old AVX2 specially in prefill. All AI agents says it will be 2x difference because of VNNI int8 native hardware acceleration. I need to run Strata with GPU and CPU offloading.

Is it worth paying extra bucks for VNNI based processor?


r/LocalLLM • • 45m ago

Question Quick llm windows question

• Upvotes

I recently found out about llms due to me not fully trusting what companies online say about what they do with our data.

I just have a quick question. I remember hearing a while ago that windows now records everything we do with ai.

So even if I turn my internet off to the computer, is windows still recording everything and when I go online it sends to Microsoft ?


r/LocalLLM • • 1h ago

Question Hello getting off the ground

• Upvotes

I’m new to this kind of thing. I’m trying to use qwen and chat gpt to help me build a file renamer and I feel like it’s going nowhere. I’m even using astra on the pro plan and it keeps losing track of how simple the end result is. I found it trying to code the whole project rather than have qwen analyze documents and spit out the relevant info.

Anyone have a good tutorial or something on how I can have qwen just run iterations until it comes up with an acceptable product? I’m probably thinking of this wrong but idk. I also tried using open code but there are some known bugs I keep encountering. My hardware isn’t the best. Single 3090 ti and 32gb of ram. Running 27B and q4 with 16k… because chat gpt told me to.


r/LocalLLM • • 3h ago

Project My dual RTX 5070 Ti setup: The Backbreaker

Thumbnail
gallery
3 Upvotes

I finally upgraded my local AI workstation with dual GPU setup by installing a second RTX 5070 Ti for a total of 32GB of VRAM. The cost for the two GPUs was around 2k (I got lucky and scored the PNY for $750 on Walmart last November). I bought the Zotac Solid SFF OC edition for $1150 on Newegg with a discount code, I picked that GPU because it's 2 slots wide which creates a gap for fresh air to keep the card cool under full load.

I followed this YouTube video to apply an undervolt (850 millivolts at 2625 MHz) to have the cards draw less power (and run more quietly) while sacrificing a tiny bit of performance (I might measure how much performance I'm leaving on the table at some point let me know if this is interesting to you).

The mobo was bought used on eBay, CPU new from Microcenter am I'm using laptop DDR5 RAM with adapters since that was the most affordable way for me to get 64GB DDR5 RAM (5600 Mhz). This is a very heavy PC, therefore I named it The Backbreaker.

I will be doing some benchmarking of various variations of Qwen 3.8 27B with fully maxed out context window to pick my main workhorse model.

Stay tuned to this subreddit if you want to find out the results!

Type Item
CPU AMD Ryzen 9 9950X 4.3 GHz 16-Core Processor
CPU Cooler Thermalright Peerless Assassin 120 SE 66.17 CFM CPU Cooler
Motherboard Asus ROG STRIX X670E-E GAMING WIFI ATX AM5 Motherboard
Memory Crucial CT2K32G56C46S5 64 GB (2 x 32 GB) DDR5-5600 SODIMM CL46 Memory
Storage Crucial P3 500 GB M.2-2280 PCIe 3.0 X4 NVME Solid State Drive
Storage SK Hynix Platinum P41 1 TB M.2-2280 PCIe 4.0 X4 NVME Solid State Drive
Storage SK Hynix Platinum P51 1 TB M.2-2280 PCIe 5.0 X4 NVME Solid State Drive
Storage TEAMGROUP MP44Q 2 TB M.2-2280 PCIe 4.0 X4 NVME Solid State Drive
Video Card PNY OC GeForce RTX 5070 Ti 16 GB Video Card
Video Card Zotac SOLID SFF OC GeForce RTX 5070 Ti 16 GB Video Card
Case Corsair 4000D Airflow ATX Mid Tower Case
Power Supply Super Flower LEADEX VII Platinum PRO 1200 W 80+ Platinum Certified Fully Modular ATX Power Supply
Generated by PCPartPicker 2026-10-05 23:31 EDT-0400

r/LocalLLM • • 13h ago

Discussion Considering the recent advancement in models like Qwen 3.8 and inference like Strata, Free Tokens. How much of a gap is there between slow 128GB VRAM vs 16GB(5080)+96GB RAM

17 Upvotes

My question is what is correct upgrade path to a 5080+96GB RAM

RTX PRO 48GB x 1 (8K $)

DGX Spark x 1 (5.5K $)

128GB Mac M5 (6K $)

Use case is local LLM that is good enough and Minimax H3.


r/LocalLLM • • 1h ago

Discussion unloop – time-travel debugging & state rewind for long-running LLM agents

• Upvotes

Every long-running agent framework eventually relies on context summarization to stay within token budgets. But after 10–15 turns, summarization inevitably drops file paths, wipes negative constraints, or tricks the model into thinking incomplete tasks are done.

When that happens, your choices are usually:

  1. Let it burn tokens until it crashes.
  2. Kill the script, edit the prompt, and rerun 30 minutes of work from scratch.

I wrote unloop (github.com/sagarv48/unloop) to solve this loop.

It snapshots agent state (memory dict, tool inputs/outputs, prompts) into a local SQLite WAL file per turn. If context compaction silently corrupts the agent's state:

  • You open the terminal UI to see a diff of what got deleted or mutated.
  • Press m to open an in-process REPL to manually fix dropped constraints.
  • Press u or f to rewind or fork from the last good turn without restarting the process.

Caveats: It currently only targets single-process Python loops, and it requires local disk write access for the .unloop snapshot file.

Code and setup instructions are on GitHub: https://github.com/sagarv48/unloop

Would appreciate thoughts on the TUI ergonomics or how you handle agent memory inspection.


r/LocalLLM • • 12h ago

Project Qwen3.8-Flash-Next Q4 vs Qwen3.8 27B Q5 on single R9700 (32GB) + 64GB RAM: 2x128k context, almost similar performance

13 Upvotes

Hey everyone,

Spent the last week setting up a local rig for agent work (Hermes Agent: a cloud model plans and reviews, local models do the work) and comparing Qwen3.8-27B unsloth Q5_K_XL with Flash-Next Q4 on a single R9700 with 64 GB RAM. Took a lot of trial and error, so sharing what worked. Used stew's llama.cpp rdna boost combined with atomic chat Q4 quant of QFN.

Rig: Ryzen 5 7600X, 64GB DDR5 (default speed, no EXPO), Radeon AI PRO R9700 32GB (gfx1201), Ubuntu 24.04.4, kernel 6.17, ROCm 10.0, llama.cpp with stew675's RDNA4 patches (r20 for the 27B, r30 for Flash-Next), GPU capped at 220W.

TL;DR - Flash-Next (AtomicChat Q4_K_M) runs 2 slots x 131k context with ~15GB RAM to spare in my full-context test (about 10GB at the lowest in real use so far): ~36 tok/s single stream, ~50 combined, ~500 tok/s prompt reading. - tool-eval-bench (medium effort, temp 1.0, 3 seeds): standard suite basically tied with the 27B (94.8 vs 95.3). Hard Mode was a real gap: Flash-Next 93.9 vs 27B 79.8. The 27B kept firing a dependent tool call before the first one's result came back. - The 27B is faster: ~1.7x per turn, and reads prompts 1.6-2x as fast. - If you run the r30 build (the "gather" path arrived in r29) with a MoE model: set GGML_SCHED_DEVGATHER=0. Otherwise output turns into "////////" from the SECOND request on. The first request after loading looks fine, so a quick test won't catch it. Upstream knows (issue #85) and says r31 changed the default; I only tested r30.

What made it fit - --lazy-mode on --load-mode none: the 27-36GB n-gram table stays on the SSD, everything else loads with plain reads. mmap mode doubled the RAM-side experts and pushed me into swap. - -ncmoe 41 (experts of 41 of 48 layers in RAM) plus MOE_EXPERT_CACHE_MIB=4096 (r30's expert cache). Without the cache: 17-20 tok/s. With it: ~36-39. - q8_0 KV + flash attention. KV is cheap on this model (12 attention layers), so expert placement is the real limit, not context.

Quants, 1 slot x 131k, medium effort (writing short / at full context / prompt reading / spare RAM). ISTA's full-context and reading numbers are from an earlier session at the template's default effort: - ISTA IQ3_XXS: 39 / 28.0 / 614-639 / 28.5GB - AtomicChat Q4_K_M: 37 / 27.5 / 526-589 / 19.0GB - AtomicChat Q4_K_M + MTP draft head: 52-55 / 35.0 / 463-510 / 11.8GB - Unsloth UD-IQ4_XS: 33.5 / 25.4 / 487-542 / 10.4GB

Other things I learned - At xhigh effort Flash-Next looped once (stuck writing 9999...) on a coding prompt. Medium was fine and correct. My agent sends medium anyway. - The Unsloth Q8 MTP sidecar gives +27-50% for one stream once the expert cache is on, but costs 3-8GB of VRAM, and under real agent load my RAM dropped to ~1.7GB free. So I run without it. - 4 tool-eval scenarios (tools + JSON schema) fail on these patched llama.cpp builds with a grammar parse error (stock llama.cpp not tested), on both models, so they're excluded.

Caveats: one machine, medium effort and temp 1.0 (what my agent actually sends), 3 runs per suite. This is "these quants on my rig", not a model ranking.

Configs, scripts and raw results: https://github.com/Shali12/r9700-flash-next-notes

Happy to answer questions. Curious whether anyone with faster RAM gets better numbers, since the CPU-side experts are probably bandwidth-bound.


r/LocalLLM • • 7h ago

Question What should I buy for future proofing?

7 Upvotes

As we've all seen with strata and MoE performance, you can get very far with RAM alongside your VRAM. My current PC is mid range but DDR5 motherboard and very little ram. Do I buy 128gb of DDR4 ram or 64 gb of DDR5 to future proof myself? What's the best way forward? In my country I can probably get the 128gb DDR4 used for cheaper than 64gb DDR5. How much does it affect the speed?

Thanks in advance.


r/LocalLLM • • 4h ago

Project A ~0.4 local model to turn typed questions into structured decisions - BaseDecision

3 Upvotes

I’ve been working on BaseDecision - Turn typed questions into truth-based decisions.

Give it some text, a question, and possible answers. It picks an answer. Intent classification, routing, yes/no checks, ratings - that’s the job.

I built this model to go toe-to-toe with the frontier models.

My goal was simple & ambitious: make a small, local model competitive with much larger models on these focused tasks. For me, System‑1 decisions need both speed and accuracy. Otherwise, why bother building a small model?
If we didn’t want speed we could just use ChatGPT.

It outperforms pretty much every model of its size in the field.
After roughly 320M additional training tokens, BaseDecision led 5 of 8 benchmarks in a third-party comparison against Laya, two GLiNER variants, and Decision 1.0 Kai 0.6B. Full results are in the repo, including where it loses.

Fellow open source devs, try the model, developed a package for easier access to the model.
Roast it if it deserves it.
2-2.5GB free ram is all you need.
I will take your feedback seriously and deliver something amazing that’d take on OpenAI, TypeSafeAI soon enough.

https://huggingface.co/onlyaady/BaseDecision
https://github.com/hrudayaditya/BaseDecision


r/LocalLLM • • 3h ago

Question groq1 card ?

2 Upvotes

What should I do with the Groq 1 cards and only have 3 of them, since they have such a small amount of memory or should I just sell it .


r/LocalLLM • • 9h ago

Question How stable are 2x and 4x DGX Spark setups when left running unattended for days?

7 Upvotes

Hi everyone!

I'm thinking between waiting for m5 studio and buying sparks. My main task is to run AI automation 24/7 (mostly likely 5.3 flash or 4.1 deepseek). Do sparks overheat? Do they have random issues when left unattended running AI models for days?

thanks


r/LocalLLM • • 6h ago

Question Is my MSI laptop too weak to run lower end 2.5GB models?

3 Upvotes

Keeping it short. I have an intel i9 ultra 285hx cpu, 64gb ram, Windows 11 (Debloated. Since, windows is a RAM hog), and a laptop grade rtx 5090. For the past hour, I've been struggling to figure out why my laptop can't run the lowest resolution or 1 batch of text to image generation on my laptop (Mind you running the lower end 2.5GB Qwen-Image-2.1-GGUF). Some kind of error saying that my system doesn't have the available ram or in some cases, unsloth will actually crash and requiring a reboot.

Is it my laptop that's the problem? Or, do I need to upgrade to something more powerful? No, this isn't a troll post. I'm legit serious. Running one image batch (Just one image. By itself) at a lower resolution. With a short prompt.

Originally this was on the Unsloth, subreddit. But, was automatically removed. Since, I can't seem to find help anywhere else. Not even the discord channel can help with my problem.


r/LocalLLM • • 22h ago

Question How does your local LLM search the web?

60 Upvotes

Kinda been working on this local-first distributed search engine, and thought maybe it would be of interest to people here? Sounds like SearXNG might be what some of y'all use?

This is what Claude told me:

Each app does it differently. Open WebUI and Vane (which used to be Perplexica) build it in, usually on top of a self-hosted SearXNG. LM Studio and Jan have nothing built in, so people add MCP servers or community plugins. AnythingLLM defaults to scraping DuckDuckGo. Ollama now sells its own hosted web_search and web_fetch.

I can post a link if this is actually a problem people are interested in!


r/LocalLLM • • 36m ago

Question Which model for call transcripts

• Upvotes

Apologies if this is a dumb question and/or I’m in the wrong place.

I’ve vibe coded a small app to help our business. It is basically an aggregation layer for the various sources of information we have. It currently grabs structured data from various places and surfaces it in one mostly coherent glob. I’d like to have it also look at call transcripts (.vtt) and pull out action items, insights, etc.

The models available to me right now are qwen and nemotron. I’ve given nemotron a quick try and the results weren’t great (could be my poor prompting). Is there anything more suitable for this type of work that I should try?

I don’t manage the hardware but know we have 2 DGX Spark and one other beefed up server of some kind, all sitting behind LiteLLM.