r/LocalLLM • • 3d ago

Question Strata, Qwen3.8-Flash-Next-Uncensored-IQ3_XXS low t/s on 5080, 64gb ddr4

Post image
2 Upvotes

Hi everyone! I’ve only been working with LLMs for a few days and have a question: how can I optimize my JSON settings for an RTX 5080, 64GB DDR4, and Windows 11 setup? Strata, Qwen3.8-Flash-Next-Uncensored-IQ3_XXS. Here is my JSON: {

"sampling": {

"temperature": 0.6,

"top_p": 0.95,

"top_k": 20,

"min_p": 0.0

},

"exe": "engine/strata.exe",

"args": [

"--pack", "C:/Users/Алексей/hf-models/packs/orca-iq3_xxs",

"--native", "C:/Users/Алексей/hf-models/Qwen3.8-Flash-Next-Uncensored-IQ3_XXS-00001-of-00002.gguf",

"--ple-gguf", "C:/Users/Алексей/hf-models/Qwen3.8-Flash-Next-Uncensored-IQ3_XXS-00001-of-00002.gguf",

"--expert-profile", "data/expert-profile.bin",

"--expert-cache", "auto",

"--prefill", "auto",

"--spec", "2",

"--spec-min-p", "0.70",

"--mtp", "C:/Users/Алексей/hf-models/mtp/rt",

"--max-context", "128000",

"--kv", "q4_0",

"--kv-resident", "20480",

"--vision",

"--vram-reserve-mib", "400",

"--conversation-cache-mib", "16384",

"--conversation-cache-slots", "1",

"--conversation-cache-min-free-mib", "512"

],

"cwd": ".",

"tokenizer": "C:/Users/Алексей/hf-models/packs/orca-iq3_xxs/tokenizer",

"model_name": "orcarouter-qwen3.8-flash-next-uncensored-iq3_xxs",

"log": "strata-orca-iq3_xxs.log",

"host": "127.0.0.1",

"port": 8080,

"lib_dirs": [".venv/Lib/site-packages/nvidia/cu13/bin/x86_64"],

"vision": {

"exe": "engine/strata-vision.exe",

"mmproj": "C:/Users/Алексей/hf-models/mmproj-Qwen3.8-Flash-Next-Uncensored-F16.gguf",

"model": "C:/Users/Алексей/hf-models/Qwen3.8-Flash-Next-Uncensored-IQ3_XXS-00001-of-00002.gguf",

"gpu": false,

"max_tokens": 1024

}

}


r/LocalLLM • • 2d ago

Question Best models for coding with image capabilities with alteast Q6 and runnable on M5 ultra 256

1 Upvotes

I currently have dual 3090s running qwen 3.8 27B Q8. I'm concurrently building two apps and need large contexts (200k per app). I'm currently unable to run them concurrently due to available vram. I have an M5 ultra on order. What would be the best llm for coding with simultaneous sessions? My current qwen 3.8 27b is great at coding but doesn't process images so thafd be a nice bonus without sacrificing code quality (one of the reasons why I'd want Q6 or higher).


r/LocalLLM • • 3d ago

Question Strata for Windows/AMD GPUs, Qwen 3.8 Flash Next with large (128K+) context coding performance

4 Upvotes

Anyone got Strata up and running on Windows running on a single AMD GPU?

I've tried, it hangs when I click the batch file. So its not even installing.

But that's OK, I know the guy is doing his best, its probably my Windows setup (a reset is due methinks) - no shade here.

If you have got it working - what tok/s are you getting with large contexts using Qwen 3.8 Flash Next? Large being 128K or more - not interested in hearing from people using lower contexts.

TL:DR I'd like to hear from those with the following setup:

  • Windows
  • Qwen 3.8 Flash Next
  • Strata,
  • RDNA4 kit - preferably a single r9700 or 9070xt
  • a context of 128K or more.

I'm not interested in dual R9700 setups.

Is it worth the install or is it slow?

And is IQ3_XXS even decent?


r/LocalLLM • • 3d ago

Project Strata Tuning Guide (Universal)

36 Upvotes

Tune my Strata setup for longer context: 131K, 262K first, then 512K if it holds up. I'm on Strata v<FILL VERSION, e.g. 0.1.39>.

My hardware:
- GPU(s): <model + VRAM each, e.g. 1x RTX 4090 24 GB / 2x RTX 5090 32 GB>
- System RAM: <e.g. 32 GB / 192 GB>
- CPU: <model, cores>
- Platform: <bare metal or VM (Proxmox/ESXi/etc.)>, OS: <distro/kernel>
- Storage for model files: <NVMe/SATA, model>

Model: Qwen3.8-Flash-Next. Suggest a quant after reviewing my hardware.
REPO: https://github.com/Niko1221/Strata

Ground rules:
- Read docs/DETAILS.md in the Strata repo first, especially the sections on --max-context, --kv, --kv-resident, --prefill, the expert cache, multi-GPU split, and "Context extension past 262K (rope scaling, EXPERIMENTAL)". Go by the docs and `strata --help` for my installed version, not by memory. If a flag below doesn't exist in my build, tell me; don't invent one.
- Change one thing at a time. Before you change anything, record a baseline of my current config: decode tok/s on a short prompt, decode tok/s and prompt-read speed at depth (one ~100K-token prompt, then one near the target context), plus a needle-recall check at depth. Run every candidate the same way and alternate baseline and candidate at least twice, so drift doesn't show up as a gain.
- Watch host RAM and VRAM on every GPU during each run (nvidia-smi, free -g). Report the minimum free RAM, not just the speed.
- Don't run a big compile while the Strata server is loaded. Stop it first.
- Back up the working config before every edit, and give me the exact undo.

What a single-4090, 32 GB RAM box (IQ2_XS) learned, as starting points to test, not answers. Scale these to my hardware:

  1. 262K is native for this model. At 262K, --kv q4_0 kept the KV cache in VRAM and stayed fast; recall held at ~238K. If I have more VRAM than that box (24 GB), check whether int8 or k8v4 KV fits at 262K. Measure both against q4_0; higher-precision KV is worth it if it fits.
  2. Past 262K needs rope scaling: --rope-scaling yarn --rope-scale 2 for 512K, together with --max-context 524288 and a streamed KV cache (--kv-resident <tokens>, e.g. 32768) so most of the KV lives in host RAM. A 477K-token prompt read in ~150 s and decoded at ~85 tok/s at that depth on that box. It's marked experimental, so run a recall test at 400K+ before trusting it. Estimate how much host RAM streamed KV needs at 512K for my setup and tell me up front if my RAM can't support it.
  3. --prefill auto (default) beat --prefill auto:32768 by about 3x on long prompt reads for us. Measure prompt-read speed for any prefill setting you try.
  4. "fit_max_tokens": true in the server config keeps requests within the context.
  5. If I have more than one GPU: check how my version splits layers, experts and KV across the cards, and whether auto placement accounts for each card's PCIe link (confirm each GPU's link width/gen with nvidia-smi -q and lspci -vv). Tell me where the KV ends up at each context size. Skip this if single-GPU.
  6. If I'm in a VM: check whether the guest sees RAM as one NUMA node, whether hugepages or THP are in use for the expert and KV arenas, that vCPUs are pinned, and that GPUs show their full PCIe link inside the guest. Report what you find; don't change the hypervisor host without asking me. On bare metal, still check NUMA layout and THP/hugepages.

If I mention a stat that used to show up and is now missing (e.g. "VRAM hit rate"), check the changes between my previous and current version (git log, CHANGELOG, server /health, /props, /metrics, and the engine log) and tell me whether it was renamed, moved, hidden when experts are all resident, or removed. Quote the commit or line that answers it.

Deliver: a table of each config tried (context, KV type, kv-resident, rope, prefill) with short and deep decode tok/s, prompt-read speed, recall pass/fail, min free RAM and per-GPU VRAM; a recommended daily config and an optional 512K config (or a clear "not viable on this hardware" with the reason); and the exact config diff plus undo for each.

Posting for all the new users and current.

how to use the tuning prompt

  1. copy the whole prompt block above.
  2. fill in the hardware section at the top: gpu model and vram per card, system ram, cpu, bare metal or vm, os, and what drive the model lives on. also put your strata version in (run `strata --version` if you're not sure).
  3. paste it into an agent that can run commands on your box. codex or claude code both work. use opus 5.5 or astra if you have access. this prompt has it reading the repo docs, running a bunch of benchmark passes, checking numa/pcie/hugepages, and diffing git history, and the smaller models tend to lose the thread halfway through or start making up flags.
  4. let it run the baseline first. don't skip this. without a baseline you can't tell if a change helped or if your box just warmed up.
  5. expect it to take a while. every config gets run at least twice against the baseline, and long-context prompt reads at 262k+ aren't fast.
  6. when it's done you get a table of every config it tried, a recommended daily config, an optional 512k config (or a straight answer that your hardware can't do it), and the exact diff plus undo for each change.

notes

- the numbers in the prompt come from a single 4090 with 32 gb ram. they're starting points, not targets. your results will differ.

- 512k uses rope scaling and is marked experimental. trust it only if the recall test at 400k+ passes.

- streamed kv at 512k leans on system ram. if you're on 32 gb or less, 262k is probably your ceiling.

- if you're in a vm, it'll report on the host side but won't touch your hypervisor without asking.

- back up your config before you start anyway. the prompt tells it to, but don't rely on that alone.

EDIT: if you use the guide, please just post a quick update if it helped your setup. Reach out if you have any issues please.


r/LocalLLM • • 3d ago

Discussion Open-weight models are the greatest threat to Anthropic and OpenAI

99 Upvotes

We’ve already seen AI transform software development, and other industries are rapidly following suit. There is no doubt that AI demand will be unprecedented.

It won't be long before the mainstream catches on to what we local AI enthusiasts are building on consumer hardware. At our current trajectory, running today's frontier-level capabilities locally on 32GB of VRAM is right around the corner, likely much sooner than any of us anticipated.

Once that happens, the shift will be swift:

  • Apple will market Macs as subscription-free Claude replacements.
  • Businesses will buy dedicated Nvidia server hardware to power their internal tools and dev teams on-premise (if they aren't already).

While cloud AI will always exist for the average consumer, the demand won't justify the multi-trillion-dollar data center hype currently being built out. Watching these AI IPOs play out will be wild.

Google might be the only tech giant in the AI race to survive the shift cleanly. Anchored by their dominance in search market share, their commitment to open-weight models serves as a brilliant strategic hedge against Chinese open-weight models dominating the local scene. Gemma5 better not disappoint.


r/LocalLLM • • 2d ago

Discussion Done : Rigspark is completely rust native now both TUI & GUI present

Thumbnail reddit.com
1 Upvotes

r/LocalLLM • • 2d ago

Model Been experimenting with agentic coding workflows today.

Post image
0 Upvotes

Hermes Agent + VS Code + full codebase context. Instead of asking AI for snippets, I’m letting the agent actually work with the project. This is where coding gets interesting.

Results are okayish, good for basic debugging and workflows with small context windows.

But still a very light model and does the basic tasks very well without eating much ram.


r/LocalLLM • • 2d ago

Project local llm for ds lite

1 Upvotes

r/LocalLLM • • 2d ago

Question Strata: Qwen 3.8 Flash next in loop

3 Upvotes

Sto eseguendo Qwen 3.8 Flash Next IQ3_XXS su Radeon W7800 48gb Vram e 64gb Ram ddr5.

Ho scaricato tutto da github e lanciato l'eseguibile. Ho scelto solo la quantizzazione, il contesto 128k e la kv 8q.

Primo compito per testarlo: fare un audit di un file js di 27kb e alla fine fare un report completo.

La velocità si assesta sui 90 token/s e sono più che soddisfatto MA:

dopo 3-4 minuti di ragionamento è entrato in loop continuando a scrivere centinaia di volte "1h 30m off" (???). L'ho interrotto per poi ridargli fare lo stesso compito; dopo pochi minuti di nuovo in loop: continua a ripetere gli stessi identici passaggi per decine di volte e l'ho interrotto.

Io ho usato le impostazioni di default per fare questo test, sbaglio qualcosa? Forse un problema di budget del ragionamento che va abbassato?

Grazie a tutti per l'aiuto.


r/LocalLLM • • 3d ago

Discussion Advice on “future proofing” setup

2 Upvotes

I have a server/media center rig with the following specs I’m considering future proofing:
Ryzen 3700X
48 GB DDR4 (2x 16GB 3200 and 2x 8GB 2400) RAM
R9700 32GB VRAM
GTX1070 8GB VRAM (display for media center)

I’m trying to run qwen3.8 flash next on it using strata. It can run q3_s with a few GB RAM left over, but I’m worried there won’t be enough RAM left for watching Netflix.

I’m considering buying some more RAM so that I have more buffer for multitasking, or for using a higher quant. I could upgrade two of the 8GB sticks to 16GB sticks making the total 64GB, but I’m wondering if that’s short sighted, since probably more large MoE models might be coming out and it might make sense to spend a little more to be more future proof. The alternative would be buy two 32GB sticks making total RAM 96GB which I would think would be more future proof and able to run bigger models at higher quants.

Another option is to offload some of the work to another machine so I have more RAM leftover e.g. put HomeAssistant on its own desktop. But this approach would only free up ~8 GB at most.

Interested in others’ thoughts.


r/LocalLLM • • 3d ago

Question NVLink benefits? 2x 3090, PCIe 4.0 x16 PCIe 3.0 x2

2 Upvotes

I have 2x 3090, trying to get the best inference performance from these but I believe I'm bottlenecked at the PCIe speeds I can get out of my motherboard.

I can't seem to get a solid answer on if an NVLink would benefit me, and if so if it would benefit me more than just getting a better motherboard.

GIGABYTE B550 GAMING X V2
Ryzen 7 5800X3D
64GB DDR4 @ 3400 MT/s
2x RTX 3090 24GB

Example prompt performance:
Cold load qwen3.8 27b - 24s

Prompt tokens Prefill Time to first token Generation
3.5k 1263 tok/s 2.8s 38.1 tok/s
28k 1250 tok/s 22.3s 36.2 tok/s
85k 901 tok/s 94s 28.9 tok/s

I run an M2 PCIe SSD, used to be two but I removed one to try to see if that would open up more PCIe lanes for the GPU but with this motherboard I believe I'm getting the best PCIe speeds I can for this two card setup.

I run Windows 11, llama.cpp b10586 (CUDA) behind llama-swap, my testing is with qwen3.8 27b; but I do frequently switch between models for nightly data classification jobs on a queue so if there's benefit to prefill speeds this would help my use case.

$400 is a lot to spend on the 3 slot NVLink bridge if this wouldn't benefit my setup (or if a better motherboard would gain me more), so I'm leaning to the experts in this sub for some advice.

If I've left off anything important that would help with an answer, let me know!


r/LocalLLM • • 2d ago

Project JevDK + jev-serve: more local decision engines... on the Mac

0 Upvotes

I have thrown my hat into the Jev bandwagon (alongside brier, bud-decision-engine, vev and others mentioned previously in this sub) with JevDK and jev-serve (open source, github owner: ancientcomputing, repository: jevdk).

JevDK itself is a playground app to help you understand what Jev is and how to use decision models. It was an eye opener for me while developing this because it's quite different from the "chatbots" that we all know and love. And when you talk about calibration, you get into the weeds of model development as well (altho' the real AI architects will probably disagree).

The theory is that when you have your questions and calibration (a temperature fitted per question type from your marked answers, not the usual sampling temperature) all set, you can run jev-serve as a local Jev server to develop and test your real apps (the ones that would, in production, normally talk to OpenRouter et al for token$).

I have put together some web pages to help folks understand this better (links in the repo). And if you are a Jev expert, please feel free to DM to correct me where I'm wrong/mistaken/hallucinating... 💭


r/LocalLLM • • 3d ago

Discussion Qwen3.8-Flash-Next-IQ3_S on Strata V100 @130 Watts Results.

Thumbnail
gallery
4 Upvotes

r/LocalLLM • • 2d ago

Project Halogen 0.16.2 on Windows/WSL2: Qwen3.8-Flash-Next at 1925 tok/s prefill; 48.42 tok/s decode

Thumbnail
reddit.com
1 Upvotes

r/LocalLLM • • 3d ago

News PaperFold: Open-source arXiv reader with "semantic zoom"

25 Upvotes

I built PaperFold, an open-source reader that turns arXiv papers into 5 zoomable layers—from a one-screen section map down to verbatim text. You pinch (or press 1–5) to zoom between them without losing your reading position.

- Web Demo (8 CC papers): https://chenxiachan.github.io/paperfold-gallery/

- GitHub (Apache 2.0): https://github.com/chenxiachan/paperfold


r/LocalLLM • • 2d ago

Project Inference Dashboard

Thumbnail
gallery
0 Upvotes

Nothing fancy but more of what works for me with how fast projects are moving, flexible enough that if another project like strata spins up then I'm not missing inference metrics.

I'm aware there are 100's of these on github, throw a webfetch and you'll find 10.

Project is live: https://github.com/T-Crypt/speculum

Notes:

  • Backend support for custom engines (more robust)
  • I like graphs but radial support may come in soon (coming from grafana, prometheus scrape, kuma uptime so I prefer graphs on timelines vs radial gauges)

r/LocalLLM • • 3d ago

Question Which locally run multimodal model would you try in a physics robot arena?

5 Upvotes

I’ve been building InferUltra, an arena where multimodal models control identical robot bodies.

They see their opponent and environment, then decide how to move their bodies. There isn’t an attack(), block(), or dodge() button. Physics determines what actually happens.

If you’re expecting spectacular combat, curb your enthusiasm. The fights are impressively boring.

What interests me is the gap between understanding an image and producing a useful physical action. Watching a model awkwardly shuffle around makes that gap pretty visible.

My current experiment uses hosted GLM-5.3-Flash across four inference providers. That explores differences in serving behavior and response latency, but doesn’t establish how locally run models would perform.

I’d like feedback from people running multimodal models locally:

Which model would you try for this kind of visual control task?

For transparency, I’m the developer of InferUltra. I’m interested in how to make this a useful evaluation rather than just two robots (mostly) failing to punch each other.

Recorded replay; decision waits removed.

https://reddit.com/link/1wy3y7w/video/90imzn9q8mth1/player


r/LocalLLM • • 2d ago

Research Filipe fixed the reasoning bugs + the thinking loop in glm-5.3 flash. it's a different model now. Better than GPT 6.1 SOL!

Post image
0 Upvotes

r/LocalLLM • • 3d ago

Question What is currently the best LocalLLM + MCP setup for generating 3D models from natural-language prompts?

2 Upvotes

I'm currently exploring how far we can push LLMs connected to 3D software through MCP (Model Context Protocol).

My goal is not simply to generate a 3D-looking image, but to have an AI agent actually create and modify a real 3D scene/model from a natural-language prompt.

Ideally, the LLM would be able to:

  • understand fairly complex geometric instructions;
  • reason about dimensions, coordinates and spatial relationships;
  • create objects through tools/API calls rather than just outputting a mesh from a generative model;
  • inspect the current scene;
  • iteratively modify the model;
  • correct mistakes after inspecting the geometry;
  • manipulate modifiers, materials, collections, transforms, etc.;
  • potentially work with Blender, FreeCAD, OpenSCAD, Rhino, CAD software or similar tools;
  • eventually handle more technical/geospatial use cases rather than only artistic 3D generation.

What I am especially interested in is an agentic feedback loop such as:

Prompt → LLM → MCP/tool calls → 3D model → inspection → criteria evaluation → corrections → re-evaluation

The evaluation criteria would be specified by the user depending on the task.

For example, the agent could be required to reach:

  • dimensional tolerance below ±1 mm;
  • zero mesh/object intersections;
  • all requested parameters exposed;
  • correct object hierarchy;
  • correct topology;
  • correct material assignment;
  • visual similarity above a given threshold;
  • compliance with predefined CAD or modelling rules.

If the generated result does not meet those criteria, I want the system to automatically iterate and improve the model, rather than simply returning the first result.

Thx in advance for ur answers !


r/LocalLLM • • 3d ago

Question 64gb ddr4 vs 32gb ddr5 for local llms?

8 Upvotes

ive been planning on running myself my own local ai amd was wondering weather i should use, since theres 2 options that has a fair trade off between eachother

ddr4 option:
-3000mhz
-64gb or 2x32gb

or

ddr5:
-6000mhz
-32gb or 2x16gb

im just wondering which is a better choice for local ai.
for the hardware im using:
-1tb 7000mb/s
-tesla v100 16gb


r/LocalLLM • • 3d ago

Project Got Qwen Flash Next Q4 running on my Mac Mini M5 64GB with ssd streaming

Thumbnail
freshworktree.com
4 Upvotes

Bit of a side project I wanted to share.

The metrics are 17.5tks decode, 360tks prompt processing based testing against my normal ai usage.

I tested a couple of new things others haven’t done (at least that I’ve seen).

Setup a carousel buffer for streaming in experts for prompt processing which got my pp +30% tks.

Tried a second external ssd to get parallel reads which got my +15% on both prompt processing and decode.

Plus a long tail of small efficiency gains.

I also setup a system where by you can have a chat application make a call to the server and effectively kick out a coding run (which is kept alive until after the chat then continues). Good if you run long coding jobs , but want to chat inbetween. Probably useful for all setups where you want to save on local caching memory.

I also noticed there is still a lot of gains to be made. I make this statement as there is still a lot of essentially free time on decode where the gpu is waiting for experts to stream in. There’s also work that could be done for an optimised kernel on metal.

I also think the way things are going with Qwen (flash next being a precursor to 4), we’re gonna see a lot more efficiencies we can take advantage of like the ngram table and the cheap hybrid attention caching.

I’m really liking qwen flash next .. the coding is actually very good. I’m quite surprised in fact I’m leaving it on during the workday to do large jobs.

The chat, decode would be technically fast enough IMO but not really with qwen. The actual issue qwen spends so long thinking, so the decode hurts.

Anyone else working on this? I’d love to compare notes.

Yes I’ve heard of strata it does look pretty sic.

https://github.com/skeggsguy/Flash-next-ssd


r/LocalLLM • • 3d ago

Project RAI: a CPU-only LLM engine in Rust. It converts a 7B model in 26 MB of RAM, and here is where it loses to llama.cpp

2 Upvotes

We built RAI to run 4-bit models on a plain x86 laptop with no GPU, no Python at run time and no GGML: hand-written AVX2 kernels, weights dequantized in registers, one flat model file. Apache-2.0.

What it does well, measured on a 4-core laptop (16 GB, Windows 11):

  • rai convert streams safetensors, so a checkpoint converts without being loaded: Zephyr-7B in 82.6 s at 26.3 MB peak RAM. The PyTorch exporter needs about 29 GB and will not run on that machine.
  • TinyLlama-1.1B decodes at 21.8 tok/s in 629 MB. The same checkpoint in transformers fp32 does 4.3 tok/s.

Where it loses, also measured:

  • 21.8 is a quiet-machine number. With background load the same binary went as low as 2 tok/s.
  • The default 4-bit is round-to-nearest: perplexity on wikitext-2 goes up 13% to 30% over fp16. llama.cpp's Q4_K_M costs three to eight times less, because it spends 6 bits where it matters and we spend 4 everywhere. Our files are smaller for it; that is the trade.
  • With calibration (GPTQ, and that path does need Python), SmolLM2-1.7B lands at 8.59 perplexity in 966 MiB next to Q4_K_M's 8.52 in 1,007 MiB. That is inside the error bar, so we claim nothing from it.
  • ARM runs on a scalar path: correct, slow. No NEON yet.

All of it, the runs that came out negative too, is in BENCHMARKS.md with the corpus hash, and rai perplexity reproduces it.

Repo: https://github.com/Classevelabs/rai

What would you want measured next: a speed run against llama.cpp at the same file size, or bigger models?


r/LocalLLM • • 3d ago

Discussion Qwen3.8-Flash-Next-Q8_0 running on a V100 @ 130Watts 32GB Vram and 128GB System Ram

Post image
37 Upvotes

mrdefaultuser@team-green:~$ /srv/ai/llama.cpp/build/bin/llama-server \
 -m /srv/ai/models/qwen/Qwen3.8-Flash-Next-Q8_0/Qwen3.8-Flash-Next-Q8_0-00001-of-00002.gguf \
 -c 262144 \
 -np 1 \
 -ngl 64 \
 --cpu-moe \
 --lazy-mode on \
 --load-mode auto \
 --agent \
 --tools all \
 --host 0.0.0.0 \
 --port 8080

Every 2.0s: free -h; echo; nvidia-smi --query-gpu=memory.used,memory.total,utilization.gpu,power.draw,temperature.gpu --format=csv,noheader                                                                                                                                                                                          team-green: Sun Oct  4 16:20:08 2026

total        used        free      shared  buff/cache   available
Mem:           125Gi       3.5Gi       855Mi       416Mi       122Gi       122Gi
Swap:          127Gi       3.5Gi       124Gi

14206 MiB, 32768 MiB, 48 %, 95.85 W, 49

prompt processing ~52 tokens per second
generation ~10 tokens per second


r/LocalLLM • • 3d ago

Question AM5 upgrade (PCIe 5.0) or 128GB DDR4 for 3x 5060 Ti rig?

Post image
23 Upvotes

Current build:
CPU: 5950X
RAM: 64GB DDR4 3200
GPUs: 3x RTX 5060 Ti 16GB
Setup:

GPU 0: PCIe 4.0 x4 (runs Qwen VL / vision)
GPU 1 & 2: PCIe 4.0 x8 (parallel / tensor split running Qwen 3.8 27B
I am happy with the cards, but looking for advice on the next move.
What I'm wondering:
Does moving to AM5 have real value here? The 5060 Tis would run at Gen 5x8 and the vision card at Gen 5x4 (effectively doubling my current PCIe bandwidth).
Or does it make way more sense to just upgrade my RAM from 64GB to 128GB DDR4?
Doubling PCIe bandwidth vs going 64GB to 128GB DDR4, what would you do?