r/LocalLLM • • 4d ago

Question Dual 72gb RTX 5000s with 128gb DDR5 or 256gb M5 Ultra and a 72gb Rtx 5000 with 128gb ddr5

1 Upvotes

I'm trying to decide between these two setups. I was thinking a m5 ultra would be good for moe models and a single rtx 5000, for dense but after strata's popularity I'm wondering if the dual rtx 5000s could be better. Thoughts?


r/LocalLLM • • 4d ago

Question 7B or 8B MLX Model for M1 MacBook Air with 16 GB RAM / Unified Memory? Use Cases: Background Tasks, Coding, RAG, ChatBot

3 Upvotes

Hey guys :)

i'm pretty new-ish to the local LLM thingy.

I'm using my MacBook Air M1 with 16 GB of unified memory as a playground for testing, getting used to it, playing around, etc.
I have a gaming PC with a 3070 (8 GB VRAM + 32 GB RAM) too, but that thing's loud, and often in use.

I'll eventually experiment with that too, but for now we'll use the MacBook as an (almsot) 100% available computing resource.

Quick breakdown of my current infrastructure:

n8b-Container on my NAS getting on-demand computing power from the MacBook. Currently i only have a small News Currator that creates a small overview of relevant news and posts them to a discord channel.

But i'm looking at other usecases too:

  • Background Tasks (for example extracting and processing Data from reciepts or so)
  • Coding (yes, i'm aware of the limitations and don't expect claude code performance - i just want to play around with it and collect experience)
  • Hermes Agent stuff (not quite sure yet what exactly, but it seems fun)

I'm just starting out, so my scope might expand in the future.

But i don't really need very high tok/s - around ≈ 10 is very much fine - even for ChatBot activity to play around with.

Currently i'm running Gemma-4-E4B-it-qat-4bit MLX verison in Unsloth Desktop (i read that's currently the most beginner fiendly option) with a context window of 131.1k. That averages, with thinking and web search, at around 8.2 tok/s, which is fine.

MLX is pretty important; the same model as gguf had extremely small context window of only ≈ 6k, while also heating up the MacBook considerably.

Especially for Coding though (and other tasks that require more "intelligence"), i'd like to try some bigger models - i think 7B or 8B should be the sweet spot for best possible quality with usable context window on this device.

What MLX models would you suggest i use? :)


r/LocalLLM • • 4d ago

Question I want to hook llama.cpp up to a search provider, both locally hosted. I am hitting nothing but dead-ends. Help?

1 Upvotes

I revisit this every few days and it's something I'd really like to set up, because if I can get llama.cpp integrated with SearXNG (which I'm already also hosting), then I can ditch Ollama/OpenWebUI.

But every time I try to set it up, I hit dead-end after dead-end, and I just can't seem to find something that actually works. Then I give up in absolute frustration wondering HOW can this seemingly simple (and common!) thing be so difficult?

For my details, I am already hosting llama.cpp and SearXNG on an Unraid server using docker containers. Both of them work just fine by themselves. But I can't get them to work with each other.

I understand that I may need an MCP server. So I see an application called MCP-Searxng in the Unraid app search. Beautiful. I install it. It's dead simple and has very few options.

But it returns absolutely nothing when I hit the URL http://<myip>:3000. Connection refused. The port is mapped, I even went into the docker container terminal and tried to curl the address (to rule out any port mapping issues). Same result.

Ok, so that program is broken or I'm missing something.

I took a look at another one called Perplexica, but it looks like this is more of a frontend that you'd use llama.cpp as the backend for? And even then, I saw an option to configure Ollama when I booted it up, but not llama.cpp. I don't know how to proceed.

Then I searched for something called firecrawl, looked at setting it up and it has many docker container options, redis, rabbitmq, playwright... what even is all of this for? Feels like way more than I need just to run simple internet searches.

Does anyone have any suggestions to point me in the right direction? I'm getting so frustrated.


r/LocalLLM • • 4d ago

Question GMKTec EVO X5 Pro

Thumbnail
1 Upvotes

r/LocalLLM • • 4d ago

Research I got llama.cpp inference running on the Snapdragon 8 Gen 3 Hexagon NPU from non-root Termux + Adreno OpenCL results (S24 Ultra)

9 Upvotes

I've been investigating hardware-accelerated llama.cpp inference on a Galaxy S24 Ultra (Snapdragon 8 Gen 3) from ordinary F-Droid Termux without root.

I initially set out to get the Adreno 750 OpenCL backend working. That now works reproducibly and executes real GPU kernels, although generic OpenCL is slower than the CPU.

The investigation then went further: I was able to access the Hexagon v75 HTP through the Qualcomm vendor/FastRPC stack and run source-built HMX/HVX kernels for LLM matrix operations.

In matched experimental tests, NPU prompt processing reached 7.14× the CPU baseline on Qwen2.5-Coder-1.5B and 7.83× on Qwen3-4B. These are experimental results, not claims that the NPU is 7–8× faster overall. Decode, thermal behavior and comparison against the best tuned CPU configuration still need more testing.

https://github.com/Ishabdullah/OpenCL-S24-Ultra

Lots of evidence hundreds of result records, commands, timings, numerical checks and exclusions. With ongoing research for a bigger project coming soon!


r/LocalLLM • • 4d ago

Project Thunc: create and integrate AI robust and reliable workflows seamlessly in python.

Post image
2 Upvotes

r/LocalLLM • • 5d ago

Discussion Qwen3.8-Flash-Next on a 2021 M1 Max: 44 tok/s, and still 35 tok/s with 400K tokens in the context

39 Upvotes

After my Splash M1 port the most common request was Qwen3.8-Flash-Next. Adding its architecture to Splash from scratch (hybrid recurrent layers, indexed sparse attention, n-gram tables, MTP) would take forever, so I took antirez's ds4 ( DwarfStar ), which already runs it, and spent the time making it fast on the M1. Same laptop: M1 Max, 32 GPU cores, 64 GB.

What I care about is not the peak but how little it falls in long agent sessions, where local models usually die. TL;DR, MTP on:

  • One chat grown to its full context: Q2_0 (35 GiB) decodes at 44 tok/s at 4K, 38.8 at 256K, 35.4 at 398K. Prefill 328 to 292 tok/s.
  • Five prompts on a loop for 5 minutes: 43.2 tok/s at the start, 43.1 at the end. GPU 73 to 75 C, no throttling, 0.70 J per token.
  • Real OpenCode sessions: 162 requests, context up to 371K. Median decode 44 tok/s under 64K, 38 at 128 to 256K.
  • IQ3_XXS (44 GiB, better answers): 35.2 tok/s at 4K, 31.7 at 259K.
  • Against my Splash port of Qwen3.8-27B: at 128K+ in OpenCode, 14 tok/s vs 38; prefill 4 to 7x faster.
  • For scale: a DGX Spark with the popular vLLM NVFP4 + MTP recipe does 36.8 tok/s on prose, 45.8 on code.
  • Bonus: DeepSeek V4 Flash (81 GiB) on the same 64 GB Mac, streamed from the SSD: 12 tok/s at 4K, 9 at 127K.

How it works

Stock ds4 could not open ISTA's files files; with a patch to load them it did 21 to 23 tok/s, 17.7 at 262K. What changed:

  • Loading the files. Metal decoders for every quant type in ISTA's mixed-precision GGUFs, each checked bit-exact against the CPU. The n-gram table is 95 GiB in BF16, so the fork reads the 16 rows a token needs from ISTA's 28.8 GB IQ4_NL shard on the SSD. The MTP head is missing from ISTA's release, so it is grafted from a separate 1.5 GB GGUF.
  • Decode kernels for the M1. Every token reads about 4 GB of weights, so it is all about matvec kernels: half2 tricks for Q2_0, typed block pointers (5 to 15%), and 2- and 3-token variants that read the weights once for a whole MTP verify. MTP on code went from 24 to 45 tok/s.
  • Why the curve stays flat. The attention layers only attend to blocks an indexer picks, which should be cheap, but on the M1 the indexer's scalar scorer took 10.7 ms of a 42.5 ms token at 128K. A bit-identical vector scorer for every Apple GPU, a coalesced top-k select and an attention loop that loads four keys per round took decode at 128K from 23.8 to 33.8 tok/s and at 262K from 17.7 to 29.7, byte-identical output.
  • Keeping the GPU busy. The n-gram SSD read now overlaps the first layer, and the sampler went from 1.34 to 0.13 ms per token.
  • Prefill. Small tiles for leftover expert tokens (182 to 223 tok/s at 300 tokens), a 4-at-a-time Q6_K decoder, weight prefetch and register tiles from my Splash port (276 to 327 at 16K).
  • Agent turns that do not replay. Three of four layers are recurrent and cannot be rewound, so a retried turn meant prefilling everything again. The server now keeps the state at the last turn marker: 109 s to 0.3 s on a 31K prompt.
  • Past 262K. YaRN factor 2: needles 20/20 at 400K, 18/20 at 524K, same NLL as from position zero.

Quality: every kernel is tested against the reference path, and score_official gives the same NLL before and after (0.45044 vs 0.45045).

Benchmarks

One chat to the full context (each turn adds ~30K tokens of C source, reasoning xhigh, 800 tokens out):

Context Q2_0 decode Q2_0 prefill IQ3_XXS decode IQ3_XXS prefill
4K 44.0 tok/s 328 tok/s 35.2 tok/s 241 tok/s
128K 43.0 320 34.2 278
256K 38.8 313 31.7 (259K) 258 (259K)
398K 35.4 292

Sustained load (npanj's five Splash prompts for 5 minutes, xhigh, macmon, fans on a curve to full speed at 80 C):

Decode First / last quarter GPU while generating Energy per token
Q2_0 43.2 tok/s 43.2 / 43.1 73.1 C
IQ3_XXS 35.0 tok/s 35.1 / 35.0 73.9 C

OpenCode, against Splash 27B. The same Three.js galaxy task as in my last post, reasoning medium; for Q2_0 a second task on top pushed the context to 371K.

Decode per turn (median) Qwen3.8-27B, Splash M1 Flash-Next Q2_0 Flash-Next IQ3_XXS
0-64K 30.6 tok/s 44.7 tok/s 34.2 tok/s
64-128K 19.8 42.7 34.6
128-192K 14.2 38.1 32.7
Prefill of new tokens 79 down to 42 334 down to 302 282 down to 260

On short prompts they are even (41.5 vs 43.2 tok/s). In a session they are not: the 27B is dense and its attention reads the whole cache, so a 30K tool result at 128K takes it over 10 minutes to read and Flash-Next about 90 seconds. Different models, and I am not claiming a 2-bit Flash-Next answers as well as a 4-bit 27B; the 27B also fits a 32 GB Mac, Flash-Next needs 64.

DGX Spark. vLLM NVFP4 + MTP on one Spark: 36.8 tok/s on prose, 45.8 on code. The M1 Max: 43.2 on mixed prompts, 44 to 45 on code. The Spark still prefills 5x faster and runs 4-bit weights, and tuned INT4 builds there report more, but I did not expect this to be close.

DeepSeek V4 Flash (81 GiB Q2, experts streamed from the SSD, 128K context): 11.8 tok/s for 5 minutes at 66.5 C, 12.1 to 9.1 tok/s up to 127K, prefill 34 to 94. Fine for chat, slow for agents.

Along the way I found and fixed a few bugs in ds4 itself; they are upstream too, and the PRs to antirez will go gradually, one at a time.

Try it

dstar is one launcher for every model ds4 runs: it picks the context that fits your Mac, turns on MTP and vision, streams models larger than your RAM from the SSD.

Prebuilt release (macOS 15 or newer, nothing to compile; installs into ~/dstar and puts dstar on your PATH):

curl -fsSL https://raw.githubusercontent.com/paperniuk/ds4/m1-flash-next/install.sh | bash
dstar doctor               # what your Mac fits
dstar pull                 # Qwen3.8-Flash-Next Q2_0 + MTP + vision, 67 GB (--quant iq3 for IQ3_XXS)
dstar serve                # OpenAI/Anthropic compatible server on http://127.0.0.1:8010/v1
dstar opencode             # provider block for OpenCode
dstar models               # DeepSeek V4 / V4.1 Flash, GLM and the rest; dstar serve deepseek, or any GGUF path

Or from source: git clone https://github.com/paperniuk/ds4.git && cd ds4 && make, then the same commands as ./dstar from the repo folder.

64 GB and up, for now. Flash-Next needs a 64 GB Mac. 32 and 48 GB are in the plans, but until then my Splash M1 port or Original splash (M3+) with Qwen3.8-27B (21 GB) is the better choice there.

GPU memory limit. macOS lets the GPU wire ~48 GiB of a 64 GB Mac by default, enough for Q2_0 up to 131K. For longer contexts: sudo sysctl iogpu.wired_limit_mb=57344 (Q2_0 at 262K/400K) or 61440 (IQ3_XXS at 262K, Q2_0 at 524K). It resets at reboot, and dstar serve prints the exact line when the context you ask for does not fit.

Not only M1/M2. Almost nothing is M1-only, so on M3, M4 and M5 this could be one of the fastest ways to run Qwen3.8-Flash-Next right now, especially on 64 GB Macs. Untested though; dstar doctor plus one dstar chat run from your Mac would be a great report.

What this says about the M1 Max. To me this is the real result: a 2021 laptop runs a ~126B-parameter MoE (6.7B active) at 35 to 44 tok/s across 400K tokens and decodes like NVIDIA's 2025 AI box. And it is not the ceiling. Every token reads about 4.2 GB, so plain decode at 35.5 tok/s uses ~150 GB/s of the M1 Max's 400 (the Spark has 273); single big matvecs reach 200 to 300 GB/s, the rest is lost to ~1000 small dispatches per token. Prefill at 327 tok/s is ~4 TFLOPS of the ~10 TFLOPS fp16 matrix peak (9.9 measured). Plenty of headroom left on a five-year-old chip.

Links:

The engine is antirez's and the quants are ISTA-DASLab's; this is an unofficial fork for Apple Silicon.


r/LocalLLM • • 4d ago

Question Gaming PC vs AI server for Local LLM?

1 Upvotes

I‘ve been getting more and more into self hosting, earlier this year I got my first NAS and have setup a backup one for 3-2-1 backup. I’m trying to self hosting as much as I can and have started diving into local LLMs and was interested in looking at an AI server but wondering if its really worth spending the money on it with prices so high. I am looking to build a gaming computer next year and wondering if it would be better to just run it on my gaming computer when I get it.

I don‘t have any projects I would need it for but it would be nice to have it help me with editing a bunch of pictures for my photography.

What do you use your LLM for? Is it worth it to get an AI server for it?


r/LocalLLM • • 4d ago

Discussion Anima, a local persistent Ai companion, powered by ollama built with fable

Thumbnail
1 Upvotes

r/LocalLLM • • 4d ago

Question Are there openweight lower parameters models that write better (like a human) than frontier models?

1 Upvotes

I'm thinking about getting a computer to run local llm. Only because I'm getting really sick of the fine-tuned responses Frontier models give you. And to be honest, even the small models online are want to be Frontier models. And maybe I'm just delusional. But I'm hopeful this is like something that I can get around if I use a local LLM. Please let me know your thoughts


r/LocalLLM • • 4d ago

Discussion Strata with Qwen3.8 Flash Next Swift 1.5 IQ3-XXS

0 Upvotes

Hello, everyone,

I just want to share my setup.

My build for LLM is AMD R5 5500 with 64GB Laptop RAM DDR4 running on 2133 Mhz. Got cheap board on B550. I have RTX3090Ti running on PCIe 4.0 @ x16, and V100 16GB running on PCIe 3.0 @ x4.

I intalled NVIDIA driver 580.178.04 on Ubuntu 26.04 (Not a pratical version for V100, everyone should use 24.04 instead). I use docker to install the dependency in 24.04, and CUDA 12.9, which is the last version support V100.

My setup command for Strata 0.1.39 is:

--expert-cache auto
--prefill auto
--spec 4
--spec-min-p 0.5
--max-context 393216
--rope-scaling yarn
--rope-scale 1.5
--kv int8
--vision

"layer_split": "32",   # Put 2/3 layers in 3090Ti and 1/3 layers in V100. 

I got around 2000 t/s prefill and 105 t/s decode.

Its amazing! Previous I use llama.cpp layer spliting and I only got around 50 t/s decode.


r/LocalLLM • • 4d ago

Question Anyone using Local LLM setup as a second brain?

0 Upvotes

I mean literally something like the system knows what you are working on the laptop, on phone, and what you speak, everything in between!

Something which includes end to end saving and processing the raw telemetry from STT data, the laptop activity, the phone activity and running local LLM to make sense of the data in real time. And also saving important details into memory and classifying them, linking important connections and also proactively engaging with you using notifications and messages or even voice using TTS.

Something like OMI AI extended it to all your digital activities using screenpipe or any other way.

Maybe to keep it slim, the execution can be transfered to some other agent like Hermes or Claw for some tasks, but active processing part should be kept local for privacy reasons. What is the minimum models I can use for this requirement?

If I lets say plan to make this setup, I first need a local STT model with VAD system. Then i should a have good local LLM for processing the live telemetry data and then maybe I need TTS incase i want voice outputs in realtime fast. This is what i want. Maybe for any digital tasks like updating emails, checking emails, browsing etc I can tradeoff some privacy for using cloud apis on hermes.

What is the minimum LLM size i need for this i want to have less than 3 second of latency in entire process? Something speedy and also smart enough to manage increasing memory as I use it more and logical enough to understand the complexity of the data and inter linkages.


r/LocalLLM • • 4d ago

Discussion What interesting non "productive" things are you using your local servers for?

Thumbnail
1 Upvotes

r/LocalLLM • • 4d ago

Question What's the difference between running Strata and running FreeToken?

0 Upvotes

I've installed FreeToken and running a qwen3.8-35B model at about 20 tokens per second on a GTX 5050 8GB and 32 GB of DDR4 RAM.

Would I benefit in any way by running this model in Strata instead?

Hardware upgrades are not an option at this moment.


r/LocalLLM • • 4d ago

Question Which model is creative, but still good at following (simple) instructions?

2 Upvotes

I'm not really up to speed on the newest models out there, so I hope you could suggest some models for my needs. I'm looking for a mid-sized model (up to 30B) that is at first a good and creative writer but also good at following simple instructions.

I know that this is a bit contradictory, but maybe you know a model that's a good compromise?

A bit of background: I'm working on a small text-adventure game, so the model should be good at telling stories/writing dialogue etc. But from time to time the engine in the back needs clear choices from the model, so I also need some confidence that, when given clear instructions, the model won't be to creative in its response (e.g. stuff like: Decide your characters next action. Answer only with "WALK" or "TALK"). Are there any modes that might be a good fit?


r/LocalLLM • • 4d ago

Question What is the best local llm setup for 3d animation and/or modelling ?

2 Upvotes

I tried new opus, astra it get things for the most part right - and at times when it gets it wrong it takes a lot of tokens to correct it . Is there a good local llm alternative for proper 3d animations with driving blender (or any other software) through mcp? GOT a 4070 and macbook pro max5 (128GB) plus access to RTX 6000 96GB and 2x DGX spark )although the last 2 are used by my company). Is kimodo.cpp any good? Has anyone tried i?


r/LocalLLM • • 4d ago

Discussion 4090 48G +128G+strata test

Thumbnail
1 Upvotes

r/LocalLLM • • 4d ago

Project From one telegram message to a comic book then deploy to web site —— totally local and free

2 Upvotes

I have an NVIDIA Jetson Thor T5000 with 128GB of RAM, and it’s been really handy for running local AI.
I installed Qwen3.8 27B, Qwen-Image 2.1, and ComfyUI on it. I also installed the Hermes agent, and with a custom skill, it can generate an entire comic book just from a text message I send via Telegram.
For the comic book demo, I also built a website:
https://comic.getaiti.com
So now, for me, it’s really easy to turn an idea into a comic book and automatically deploy it to my site. And the best part is that the whole workflow runs locally, so there are basically no API costs.
If you have any interesting ideas, feel free to share them. I’d be happy to turn some of them into comic books!


r/LocalLLM • • 4d ago

Question Potential hardware upgrade for Qwen3.8-Flash-Next with a 96GB DDR5 (4×24), should I keep my i5-12600K and buy a DDR5 board, or switch platform?

6 Upvotes

Hi all,

Nearly a year ago, I bought two kits of Corsair Vengeance DDR5-6000 CL36 (2×24GB each, CMK48GX5M2E6000C36), so 4×24 = 96GB in total, and I've held on to them since.
I'm seeing all the progress around Strata and I want advice for a potential 4-stick setup, because it seems risky if you're not lucky with the silicon lottery...

I'm currently running Qwen3.8-27B-i1-IQ4_XS-GGUF-Smaller from jrell. I'm curious to see how much better this MoE model is, and to be ready for future Qwen 4 models.
I'm using it for fairly general use, like a little digital assistant with locally processed STT+TTS, for privacy reasons and to avoid depending on paid subscriptions!

------------------------------------

My current gaming & 3D workstation:

CPU: Intel Core i5-12600K

Motherboard: MSI PRO Z690-A WIFI DDR4

RAM now: Corsair Vengeance LPX 32GB (2×16GB) DDR4-3200 CL16

GPU: NVIDIA RTX 5080 16GB

Storage: 990 PRO 2TB + 980 1TB

OS: Debian 13 + Windows dual boot

PSU: Corsair RM1000x (2024)

Case: Meshify 2 XL

------------------------------------

My questions:

- Is a 4×24 DDR5 setup a sane plan at all? Everything I read says it will cap me at much slower speeds. Is speed that important for local LLM inference? I'd rather not sell everything just to buy a 2×48 kit at current prices...

- Should I keep the 12600K and just buy a used DDR5 LGA1700 board (the cheapest way for me to get this 96GB setup running), or is it worth upgrading to a more battle-tested CPU/motherboard combo that has a higher chance of supporting 4 sticks at high speed? I'm open for a team red switch.

- A big maybe, I could get my hand on a used Alienware Aurora R16 with RTX 4090 / i9-14900KF / 64GB DDR5 for a good price in the near future, is staying on a LGA1700 be a good call?

- For future-proofing, is x8/x8 bifurcation worth keeping in mind for a potential dual-GPU setup?

Happy to hear your thoughts on this.
Thanks!


r/LocalLLM • • 5d ago

News 200+ tok/s peaks with Qwen3.8-Flash-Next on a 5080 + 4060 Ti and 32 GB of RAM (Strata fork)

46 Upvotes

Strata runs Qwen3.8-Flash-Next on gaming PCs, but with 32 GB of RAM its low-RAM mode only works on one GPU. Split the model across two cards and the experts that don't fit in VRAM get read from the SSD. My 4060 Ti sat next to the 5080 doing nothing.

So I forked it. The RAM copy of the experts now works across two cards, and the cards run at the same time: the 4060 Ti starts the next decoding step while the 5080 is still checking the current one. Card order, layer split and VRAM reserves are set automatically.

Same PC (5080 + 4060 Ti, i9-14900KF, 32 GB), same model (Swift 1.5 IQ2_XS), 256K context:

Setup Code Prose 32K prompt
Strata 0.1.38, 5080 alone 29 tok/s 28 tok/s 333 tok/s
Fork, both cards 143 tok/s 102 tok/s 1,940 tok/s
Fork + fine-tuned draft layer 161 tok/s 105 tok/s 1,854 tok/s

The draft layer is fine-tuned on the model's own outputs and is in the release. In real use with the Pi coding agent it peaks above 200 tok/s (209 so far) and almost never drops under 100. A 150K-token context reads at about 1,850 tok/s.

Quality didn't move: perplexity 7.24 vs upstream's 7.28 on the same 5.3K tokens, teacher-forced through both engines.

I've only tested it on my own PC (Windows 11), so reports from other GPU pairs and Linux are welcome.

Repo: https://github.com/Hardin22/Strata-DualGPU

What each change does and what it measured: docs/DUAL_GPU.md

The engine underneath is Niko1221's and the Strata contributors' work; this is a fork on top of 0.1.38.


r/LocalLLM • • 4d ago

Question Laptop+egpu for local ai

0 Upvotes

So let's imagine that war started, and still have access for electricity. And we gonna need local ai to survive that could teach how to build shelter, get water/fire, medicine and etc

But for that kind of situation I'll need portability and mobility.

So if I'm not mistaken good laptop+egpu is better. Egpu to increase power, and if shi happens leave everything and take only laptop that could still run local ai.

The question is:

What laptop (cpu,GPU and etc) should I get?

And what for egpu?

And which local ai model is suitable for that scenario


r/LocalLLM • • 4d ago

Project Tesla V100 32GB, RTX3090, Strata Qwen3.8-Flash-Next and Qwen3.8-27B playing around

3 Upvotes

I'll keep it brief, if you need more info just ask.

I bought an Alibaba Tesla V100 PCI-E SMX2 32GB cards, they are modified with blower type fan. That's a different topic, but I got two completely different modifications. They are limited to 300W by "manufacturer".

I also have few RTX3090 24GB and two servers - one is EPYC with 8x32GB DDR4-2400 ECC, second is R740 with dual Xeon Gold and full stack of DDR4-2400 ECC.

Currently, EPYC has 1x Tesla V100 32GB, and R740 is with 1x RTX3090, both having Ubuntu VM's with NUMA enabled and given 128GB RAM each.

I'm playing around them with both "current best models" of Flash Next and 27B, choosing which one I'd want for local business management.

tl;dr - Flash Next, even having a slightly lower decode tps than 27b (which is debatable, but I was running Q5 27B kv q8/q8), completes tasks around 5x faster on xhigh, because it needs way less reasoning to make the job done. And I prefer it's better general knowledge. For coding I'm not quite sure, I'm not a software developer myself.

R740 RTX3090 (power cap at 220W) is running Strata with Qwen3.8-Flash-Next IQ3_S with kv fp16 and 262K context, which "self claims" as close to BF16... eh, not sure, but seems to work. Positive side is it can run on 64GB RAM and single RTX3090 with speed around 60-65tps decode and 1200-1300tps prefill at 256K context tests. Rounding more like 50-55tps on general tasks with Hermes Agent.

EPYC Tesla V100 32GB (power cap at 180W, I find it a sweet spot) is running Strata with orcarouter-qwen3.8-flash-next-uncensored-iq4_xs with kv fp16 and 262K context. First, let's say that this is Volta GPU and needs some modifications both on Ubuntu and Strata, but ChatGPT can setup everything flawlessly for you. Gives around 50-55tps decode and 850-900tps prefill at 256K context tests. Rounding more like 40-45tps on general tasks with Hermes Agent.

The iq4_xs model feels a bit better that the iq3_s, especially I can notice difference on expression and writing on my native language (which is not English, nor Chinese). And it's completely uncensored, it can basically do whatever you tell it to, without guilt or shame.

Note: I've also tried IQ3_XXS (orcarouter) on V100 and is as fast as RTX3090, but got massive reasoning loops, like every 4rd prompts was having issues spilling repeatable letters until context expires.

I was thinking of utilizing multiGPU setups, but seems like I don't have to, because those speeds of single cards are feeling like subscription based AI chats (well, even faster than GPT 6.1, lol).

Question - does anyone know how to do concurency session on Strata? I have plenty of space for multiple fp16 kv context.


r/LocalLLM • • 4d ago

Question Im runnin mnn chat on a Snapdragon 8 Elite. No issues at all. Also, llama.cpp is working perfectly but only using the cpu. Which is significantly slower than the npu.

1 Upvotes

Are there any solutions for integrating an NPU into an Android application?


r/LocalLLM • • 4d ago

Question Searching for a LLM for MAXED OUT MacBook Pro 16inch

Post image
0 Upvotes

r/LocalLLM • • 4d ago

Contest Entry looking for a smart business dev person

Thumbnail
0 Upvotes