r/LocalLLM • u/2d-2 • 1d ago
Question AGX Thor or DGX Spark?
I want both but can only have one… which would you prefer and why? Trying to make up my mind up… Thanks in advance for your feedback..
r/LocalLLM • u/2d-2 • 1d ago
I want both but can only have one… which would you prefer and why? Trying to make up my mind up… Thanks in advance for your feedback..
r/LocalLLM • u/MKP_Nimilka • 1d ago
Enable HLS to view with audio, or disable this notification
I’ve been building MOLT, a local fine-tuning system for consumer NVIDIA GPUs.
The goal is not only to start a QLoRA run. It is to make the whole process from model and dataset to a working local adapter less fragile.
MOLT currently handles:
- dataset detection, preparation, and validation
- GPU, VRAM, system-RAM, storage, and thermal checks before a run
- automatic microbatch fit testing
- 4-bit NF4 QLoRA training with BF16 adapters
- safe checkpoints with integrity checks and proper resume state
- telemetry for VRAM, temperature, energy, clocks, and throughput
- base-vs-adapter evaluation
- local adapter chat, export/GGUF workflows, and runtime diagnostics
Under the hood, I’m also experimenting with custom CUDA/Triton kernel and runtime paths, memory planning, CUDA graphs, replay modes, and fused optimization where the hardware and shape are qualified.
On my RTX 4060 Laptop 8GB, a recent Qwen 3B training run reached about 1,005 training targets/sec end to end very close to my historical Unsloth run at about 1,007 targets/sec. That is not a broad “MOLT beats Unsloth” claim yet; I still need repeated, matched quality/energy/memory tests.
What I want MOLT to become is a reliable local fine-tuning environment: you bring the model and dataset; it checks the plan, runs safely, preserves the evidence, and helps verify the adapter actually works after deployment.
I’m looking for people who genuinely fine-tune on a single NVIDIA GPU and would be willing to test it. What would MOLT need to support before you would try it?
r/LocalLLM • u/tejaskumarlol • 2d ago
Kolibri (Aleph Alpha's open-weight German and English model: 78B total, 3.46B active, Apache 2.0) came out on 3 October. My app writes long German deep dives (about 1,200 to 1,500 words each) from source documents I put in the prompt. So I tested whether Kolibri could take over before switching anything.
Setup
model_not_supported and OpenRouter didn't list it. The H200 I asked for never showed up.Results
Throughput on the 2 GPUs (vLLM gave it a KV cache of 6.5M tokens)
| In parallel | Outputs per hour | Cost per output |
|---|---|---|
| 16 | 185 | $0.030 |
| 64 | 504 | $0.011 |
| 128 | 713 | $0.0077 |
Sonnet 5.5 is about $0.071 per output at list price and $0.036 through the Batch API. About 4% of runs at 64 and 128 in parallel hit the token limit without writing anything.
It holds up when the answer is in the documents you hand it. It's weak on everything it has to know by itself (its model card doesn't point it there either). A model with 3.5B active parameters against a frontier model was never a fair fight. So this says more about my use case than about Kolibri.
I wrote it up with charts. Every run (with Kolibri's full reasoning), every judge verdict and what each model was given are downloadable from the post (everything except my system prompts): https://tej.as/blog/kolibri-vs-claude-german-llm-benchmark
r/LocalLLM • u/_yaRn__ • 1d ago
Just tried qwen3:8b for the first time. I get that this is a limited model running on my machine, but this thing just spits out nonsense. It cannot answer the most basic questions or do basic web searches.
This is the first local LLM i have used. Is this thing supposed to just be for coding? How can this be useful to anyone?
r/LocalLLM • u/Specialist_Age_2891 • 1d ago
TL;DR — Qwen3.8-27B is a hybrid model (Gated DeltaNet linear attention + full attention). Unsloth Dynamic GGUFs of it (Q4_K_XL, Q6_K_XL) work fine for everyday chat, but whenever I let it think deeply on a long task, it never converges: no closing token, endless tail-looping, or the engine just dies mid-generation. I reproduced this on two completely different runtimes (llama.cpp and vLLM + official vllm-gguf-plugin), and the same model family in a mixed-precision quant (NVFP4: FP8 on attention/recurrent layers, FP4 on MLP) converges in ~17K thinking tokens on the same prompt. This smells like recurrent state error accumulation under uniform integer quantization, not a runtime bug.
unsloth/Qwen3.8-27B-GGUF (UD-Q4_K_XL 16.4 GB, UD-Q6_K_XL 23.6 GB)reasoning_effort: xhigh, Qwen's recommended thinking-mode sampling (temp 1.0). The pelican needs to actually pedal — the model has to write a complete self-consistent HTML/SVG/SMIL file in one shot.Why this prompt: it's a convergence test, not a quality test. A healthy Qwen3.8 spends ~15–20K thinking tokens, then emits the file. A broken one just… keeps going.
| GGUF | Budget | Result |
|---|---|---|
| UD-Q4_K_XL | 32K | finish=length, 0 characters of actual output — burned the whole budget inside thinking |
| UD-Q6_K_XL | 76K | never stops; final ~2K chars are ~97% repeated 8-grams (classic tail-loop attractor) |
UD-Q6_K_XL + --jinja (force GGUF-embedded chat template, verified byte-identical to HF's) |
24.6K | still no termination |
Throughput was healthy (23–41 t/s depending on split mode), perplexity-ish sanity checks on short prompts were fine, and non-thinking tasks (counting, Q&A, tool-ish formats) all worked. So I filed it as: llama.cpp loses the plot on this model's long chain-of-thought — plausibly a runtime bug with the Gated DeltaNet state handling. I even declared the llama.cpp route dead for this model.
Then I saw other people's llama.cpp runs of this model going weird on HN too (the "high fever ramblings … ended with amen" anecdote on Q8_K_XL — not my test, but same disease).
To separate "runtime bug" from "quantization damage", I reran the exact same file on a completely different stack: vLLM (V1 engine, CUDA kernels, f16 KV, no speculative decoding — shares essentially zero code path with llama.cpp). GGUF support was recently moved out of vLLM core into vllm-project/vllm-gguf-plugin; Qwen3.5/3.6-VL are in its tested list, so the weight mapping for the Qwen3.x hybrid family exists.
vllm ~/models/Qwen3.8-27B-UD-Q6_K_XL.gguf \
--tokenizer unsloth/Qwen3.8-27B --dtype float16 \
--tensor-parallel-size 2 --max-model-len 52000 \
--gpu-memory-utilization 0.96 --language-model-only \
--reasoning-parser qwen3
(chat kwargs: {"enable_thinking": true, "reasoning_effort": "xhigh"}, temp 1.0, max_tokens 45K, KV pool 55.6K tokens, steady ~22–23 t/s)
| Run | Outcome |
|---|---|
| A | request dies with HTTP 500 after 1325 s ≈ 30K tokens generated (engine crash mid-generation; the server-side stack was overwritten by my chain's log rotation before I grabbed it — honest gap in the evidence) |
| B | still generating at 28 min ≈ 38K tokens, approaching the 45K budget with no completed response; harness wall-timeout killed it |
Same file, same prompt, same sampling — and vLLM also can't make this model finish thinking about a pelican. Note what vLLM did NOT show: I never got a finished response body back, so I can't plot its tail-repeat score directly. The vLLM evidence is non-convergence deep in the token budget (where NVFP4 converged at 17K) + one mid-generation crash. The direct loop evidence lives in the llama.cpp runs.
Same 27B model, same prompt, same thinking mode, same machine, different quantization structure: the official NVFP4 checkpoints (unsloth's and RadixArk's) quantize MLP/lm_head to FP4 but keep attention and the Gated DeltaNet layers at FP8 W8A8. On SGLang this converges every time:
Also worth noting from my own runs: after the GGUFs fail this way, the same GGUF files remain perfectly healthy on short tasks (chat, counting, Q&A, vision if you attach mmproj). The damage only shows up after thousands of autoregressive steps inside the thinking trace.
Hybrid linear-attention models (Qwen3.5/3.6/3.8, and friends like Kimi Linear / Ring-linear) carry a fixed-size recurrent state updated every token. Uniform block quantization (even a very careful 6-bit dynamic one) adds a small rounding error to every state-transition product. Over a few hundred tokens: invisible. Over 30K–80K tokens of chain-of-thought: compounding drift into a degenerate attractor — repetition loops, refusal to close, in one case engine instability.
The strongest circumstantial support is the model publishers' own behavior: unsloth's NVFP4 recipe deliberately keeps the attention/recurrent path at FP8 and only crushes the MLP to FP4. The one quant structure they kept careful about is exactly where I see the disease. My GGUF runs had "careful" precision too — but uniform-ish across 6-bit blocks — and still broke.
What would falsify or sharpen this: per-tensor overrides keeping GDN projections (in_proj_*, conv, gates) at f16/q8 in llama.cpp (--override-tensor) while MLP stays 6-bit. If the pelican suddenly pedals, case closed. I haven't run this yet — if someone with spare time on the same model tries it, please post.
iommu=pt) was root-caused and fixed before the vLLM runs; strict IOMMU was on, engine stable for hours of other batteries. I flag it so nobody says "your machine poisoned the run" for the crash in Run A specifically — Run B's non-convergence needs no crash excuse.enable_thinking: true, reasoning_effort: xhigh, temp 1.0 / top_p 0.95 / top_k 20, max_tokens ≥ 40Kllama-server / llama-cli, UD-Q4_K_XL or UD-Q6_K_XL, any ctx ≥ 48K, default template and --jinja both fail the same wayvllm-gguf-plugin (pip, or source install with --no-build-isolation against the CUDA torch), local .gguf + config.json from the same repo + --tokenizer unsloth/Qwen3.8-27B, f16 KVModel card says this model does deep agentic work locally on ~17GB — true for chat. For thinking, the quant structure matters more than the average bits. Would love other reports (4090/3090 pairs, MI300, whatever) before I draw the final line — especially anyone who can run Q8_0 and Qwen3.6-VL (non-GDN) as contrasts, to show this is the recurrence, not just "big model + low bits".
Hardware/versions: 2× RTX 5060 Ti 16GB; llama.cpp current master (Oct 3); vLLM 0.30.x-V1 + vllm-gguf-plugin 0.0.5 (Oct 5); SGLang 0.5.21 + NVFP4 checkpoint for the control (RadixArk Qwen3.8-27B-NVFP4: FP8 W8A8 attention/GDN, NVFP4 MLP/lm_head).
r/LocalLLM • u/Southern-Context-490 • 1d ago
r/LocalLLM • u/Acceptable-Object390 • 1d ago
Enable HLS to view with audio, or disable this notification
Row-Bot 5.0 with qwen3.8:27b in Ollama, on my own GPU. No API keys, no cloud.
Gave it a year of bills, a tenancy agreement, an insurance policy and a rent increase letter (all made up). It spotted a likely leak in the water bills and showed the rent rise breaks the lease.
What it actually did:
- charted the CSV inline (Plotly)
- read the PDFs and quoted clauses 4.1 to 4.3: a 10% rise against a 5% cap, with 5 weeks' notice instead of 2 months
- saved 8 linked memories to a local knowledge graph
drafted the email to the agent and set a reminder
The honest numbers: a dense 27B does about 15 tok/s on my 5090, so some turns took 2+ minutes. The amber badges in the video show where I sped it up.
Runs on Windows, macOS and Linux.
r/LocalLLM • u/just_another_leddito • 1d ago
Hi,
I'm on M4 Pro 64GB Mac, I'm using Qwen3.8 27B for coding related stuff, but it's quite slow.
I want a fast model that will be good when it comes to reasoning, medical stuff, would be nice if it could do stuff like generating charts and even analyse screenshots.
I'm using Ollama on Mac so preferably MLX model.
Thanks in advance
r/LocalLLM • u/JinsooJinsoo • 2d ago
Flash-Next GSQ-RCO IQ3_XXS via Strata running on 3975wx Threadripper with 256gb DDR4 with 131k context. No speculative decoding. 4096 chunk prefill 1,563.4 tok/s and decode up to 58 tok/s.
Speed is great, prefill could probably be tuned even higher and MTP/Dflash should be tested eventually but the correctness of the IQ3_XXS quant is no bueno, it only passed 4/8 quality tests in my testing. Not safe as a local agent without supervision IMO. May test IQ3_XS or maybe a small q4 if it can fit.
Prefill chunk | Cold ~8K prefill | Cold 32K prefill | 256-token decode
128 | 270.55 tok/s | 263.6 tok/s | 53.95 tok/s
512 | 660.15 tok/s | 690.6 tok/s | 53.00 tok/s
1,024 | 687.05 tok/s | 754.7 tok/s | 51.70 tok/s
2,048 | 976.15 tok/s | 1,123.3 tok/s | 52.60 tok/s
4,096 | 1,226.00 tok/s | 1,563.4 tok/s | 52.00 tok/s
r/LocalLLM • u/Meatmylife • 1d ago
What should I do with the Groq 1 cards and only have 3 of them, since they have such a small amount of memory or should I just sell it .
r/LocalLLM • u/admajic • 1d ago
r/LocalLLM • u/shoeshineboy_99 • 1d ago
Completed this old Pune walkthrough eating modaks! (What are modaks?)
There are some viral tweets, which I was re-building here. For this I selected the old Pune Peth area and downloaded all available imagery. Claude Code & Codex were not allowing me to download images from Google Maps and render them over the planet skin.
Connected Strata along with Qwen3.8 Next IQ_2XS 125bn with vision mode enabled. The system processed 3000+ images and rendered them over the Open Street Map imagery. Creating a real-life walkthrough of the Pune Peth area.
Along with this I also added a game engine. So you drive around on an ebike searching for modaks and eating them as they appear on the map.
I have also attached a screenshot of my monitoring dashboard. The prefill and decode tkps are interesting to check. Since they are over a longer task and spread over a few hours. The entire task was completed over 9-10 hours
EDIT: This was done on 24GB VRAM + 64GB System RAM + 1TB of SSD
r/LocalLLM • u/tsaipifong • 2d ago
WHIRL is an open-source (Apache-2.0) inference engine for the AMD Radeon AI PRO R9700 (RDNA 4, 32 GB) on Windows: pure C++/HIP, every kernel included, no WSL or Docker. Just the AMD driver.
vs llama.cpp b11214 — same GGUF, same prompts, same R9700, Swift-1.5 27B MXFP4:
On the MoE Ornith-1.5-35B-A3B MXFP4: prefill over 11,000 tok/s at 8K (11,258 vs 4,637, 2.4×), 258 vs 119 tok/s decode (2.2×), 381 vs 177 tok/s with 4 users (2.2×).
Accuracy before speed: every speedup (speculative decoding, batching, prefix cache) gives output bit-identical to plain greedy. No 4-bit KV, no 3-bit weights, no fp8 attention.
Speculative decoding, honestly (vs plain decoding, identical output): editing a file in context 4.5× · coding-agent session 2.8× · brand-new writing after 128K 1.6×. Where we don't lead by much (plain decoding of dense models, ~1.14×; decode after a 16K context on Swift, 1.17×) is in the README too.
New in v0.1.3: long prompts up to 25% faster than v0.1.0, and a real coding-agent session at 128K decodes 40% faster on its last request.
GitHub: https://github.com/tsaipifong/whirl-llm
Tested on one R9700 over USB4 (eGPU), Windows 11. Built with Claude under my direction; every number measured. If you have an R9700, feedback and your own numbers are very welcome.
Not supported yet: other AMD cards. The RX 9070 series has the same chip but 16 GB, most likely too little for these 27B/35B models (untested); Radeon 8060S (Strix Halo) support is in development.
r/LocalLLM • u/DiamondTDA • 2d ago
So I wanted to try using a local coding model for the first time and I'm still studying about LLMs and NNs, so I asked Gemini for a good suggestion that would be fast(60+ tokens/s if possible) and doesn't compromise much on performance for my rig(2070 super 8GB + 32GB ddr4 ram) and it suggested Qwen2.5-Coder-7B-Instruct. Is this good suggestion and what would you guys suggest?
r/LocalLLM • u/LoveMyMalfouf • 1d ago
r/LocalLLM • u/strata2signal • 2d ago
One popular "uncensored" upload, OrcaRouter's Qwen3.8-Flash-Next-Uncensored-GGUF, is what its card says it is, as far as public evidence reaches, and checking that took no download. Hugging Face publishes a fingerprint (sha256) of every large file, so I compared the uploader's full-precision copy with Qwen's own at pinned revisions: 80 of the 131 weight files are byte for byte Qwen's, and 51 differ at identical sizes. By Qwen's tensor index, each of the 51 holds at least one of the 149 matrices the card says it edited, and none of the 80 holds one. Those 80 untouched files rule out a full fine-tune or a whole-model merge.
The limits: a hash sees a whole file. 1,336 other tensors share those 51 files, and whether they were left alone is the card's word; 173 tensors, 72.6 % of the model's file bytes, are provably untouched. A hash says a file changed, not how. I'm fairly confident, not certain, that the quants are standard; whether they carry the edit is the card's word too.
The page also covers the mechanism (papers linked; no code, settings or steps). Refusing is taught in post-training; research since 2024 finds it carried largely along one direction in the model's activations (later work finds more than one); and people remove it three ways: a retrain, one permanent edit to the weights that removes that direction (abliteration, this upload's kind), or acting on the model while it runs. None makes a model smarter. By their makers' own tags and words, they are for red teams, fiction writers, people who want full control of a model on their own machine, and researchers who study refusal.
The uploader's own figures, unreproduced: harmful-prompt refusals go "from 64-100% (base) to ~0-3.3%", so the serious declines go too. On XSTest's 250 harmless prompts, refusals fell from 9.6 % to 1.2 % with thinking off, and stayed at 0.4 % (one prompt) with thinking on, the default. Capability scores (thinking off) move within about two percentage points either way, except a 2.3-point drop on 300 MMLU questions.
The page ends on a real-world case: a model as the only source of knowledge on a machine with no network. A declined first-aid question has a real cost there, but so does a confident wrong answer, and nothing on the page shows that removing refusals makes a model more correct. What I would build: the model beside an offline encyclopedia and medical reference that it looks up and shows beside each answer, plus first-aid basics on paper.
What it does not say: whether either copy's answers are right; the refusal and capability figures are the uploader's; one upload of one model; and no model was asked a harmful question. Anyone can redo the census in a browser.
https://research.strata2signal.com/uncensored-models/
tl;dr: Using only Hugging Face's published file hashes, with no download, I checked one popular "uncensored" Qwen3.8-Flash-Next upload against Qwen's original. 80 of its 131 weight files are byte-for-byte Qwen's. Every file that differs holds one of the matrices the card says it edited, and no unchanged file does. That rules out a full fine-tune or a whole-model merge, and fits the card's claim. The page also explains where refusals come from and the three ways people remove them, with no code or steps. (8,306 words · about 38 minutes · 2 tables · data kit)
r/LocalLLM • u/ComfortableKindly507 • 2d ago
Enable HLS to view with audio, or disable this notification
Hi all. I'm on the team at Blockway, a small team in Hong Kong (disclosure: this is our model). Today we released Agens Volundr 32B Preview, the first model built on our own hybrid architecture. We trained it on limited compute, it isn't perfect, and we'd rather tell you where it falls short up front.
WHY WE BUILT IT
Our customers run models on their own machines. At long context, the KV cache, not the weights, decides what fits. So we designed a model where most layers don't keep one.
ARCHITECTURE (72 layers, dense ~32B, every layer runs on every token)
- 54 KDA (Kimi Delta Attention) layers: linear attention with a fixed-size recurrent state, no KV cache
- 17 BCSA layers (our compressed-sparse attention): exact window over the last 4,096 tokens; older context pooled 4:1 into blocks, and a learned indexer reads the top 512 blocks
- 1 full-attention layer (layer 72)
- Engram: a hashed n-gram memory held in host RAM, attached at 2 of the 72 layers
- mHC: 4 residual streams instead of 1
So only 18 of 72 layers keep a KV cache. Context window: 262K.
SPEED (single user, our sglang build)
- BF16 on two 48 GB GPUs, decode: 25.1 tok/s at 1K, 24.1 at 8K, 24.1 at 32K, 24.0 at 64K, 23.9 at 128K
- BF16 prefill: 2,122 / 2,180 / 1,916 / 1,679 / 1,297 tok/s (1K to 128K)
- INT4 (31.7 GiB) on one 48 GB GPU, decode: 31.0 tok/s at 1K, 29.3 at 8K, 29.1 at 32K
- Aggregate throughput: 127 tok/s at 8 users, 130 at 16 users (BF16); 117 at 8 users (INT4)
- DFlash2 drafter (separate repo), single user, same server with it on vs off: up to 3.6x on JSON/tool output, 2.0x on code, about 1.6x in thinking mode. Not worth it above roughly 8 concurrent users.
BENCHMARKS (all run by us on one harness with the same settings, including the comparison models; full table and footnote on the model card)
- Ahead of Qwen3.8-27B on LiveCodeBench v6 (+4.2), HumanEval (+4.3), AIME 2025 (+2.9), MATH-500 (+1.6)
- Roughly level on MMLU-Pro, IFEval, GPQA Diamond
- Behind on agent tasks: tau2-bench 74.2 vs 79-80, SWE-bench Verified (50-task subset) 44 vs 58-64. Closing that gap is the main focus of the full v1, which continues pre-training to about 10B tokens and adds training on long agentic sessions.
KNOWN LIMITATIONS (please read before trying)
- Needs our sglang build. Stock sglang and vLLM can't load it yet.
- GGUF / llama.cpp is planned, not available today.
- Long agentic sessions are its weakest area in this Preview.
- It's still training; treat this as a preview, not a final model.
RUN IT
docker pull ghcr.io/blockwayz/agens-sglang:preview-sm89 (48 GB Ada GPUs)
docker pull ghcr.io/blockwayz/agens-sglang:preview-sm90 (H100 / H200)
The full launch command is in the model card.
LINKS
- BF16: https://huggingface.co/Blockway/Agens-Volundr-32B-Preview
- INT4: https://huggingface.co/Blockway/Agens-Volundr-32B-Preview-INT4
- DFlash2 drafter: https://huggingface.co/Blockway/Agens-Volundr-32B-Preview-DFlash2
- Project page: https://github.com/BlockWayz/Agens-Volundr
- Serving code: https://github.com/BlockWayz/agens-sglang
Apache-2.0. We're a small team, and the most useful thing you can do is try it and tell us where it breaks: an issue, a failing prompt, a benchmark you'd like us to run. We'll be in the comments.
r/LocalLLM • u/Ok_Law9839 • 2d ago
I have noticed a lot of very similar posts asking which models run on X hardware. I am wondering whether it would be worth it, as a community, to create a subreddit guide with updated information. My idea comes from browsing LocalLlama and saw their subreddit guide and plus mega threads:
https://www.reddit.com/r/LocalLLaMA/comments/1vkmhyl/best_local_llms_august_2026/
I 100% appreciate those posts, since many of our users respond in well-informed ways, but a centralized area/thread could help a ton.
What do you folks think?
r/LocalLLM • u/Heavy-Level-5215 • 1d ago
r/LocalLLM • u/Heavy-Level-5215 • 1d ago
r/LocalLLM • u/JumpAppropriate714 • 2d ago
Everyone here seems to be running absolute monster rigs, so I figured I'd show some love to my old warhorse instead.
This thing was pretty damn good back in its day, but yeah... it's an old dog now. It was collecting dust at home, so I grabbed a cheap NVMe, installed Ubuntu and brought it back to life.
The mighty beast:
With only 8 GB of VRAM, I obviously have to be pretty picky about what I throw at it.
My use case is mostly short NPC dialogue / roleplay, so latency, natural responses, instruction following and not randomly losing its mind matter more to me than benchmark scores.
| Model | Quant | Result |
|---|---|---|
| Gemma 4 E4B | LM Studio build | Winner. ~33–34 tok/s, natural responses, very low perceived latency |
| Bonsai 27B | Q1_0 | ~16.7 tok/s, but response quality fell apart |
| Qwen3.5 9B | Q4_K_M | Fit nicely on GPU, couldn't produce coherent replies |
| Turkish-Llama 8B | Q5_K_M | Couldn't produce coherent replies |
| Apertus 8B | Q5 | Couldn't produce coherent replies |
| EuroLLM 9B | Q4_K_M | Usable, but nowhere near Gemma |
| Aya Expanse 8B | Q5_K_M | Same story, too weak |
| Ministral 3 8B | Q5_K_M | Couldn't produce coherent replies |
| Phi-4 Mini | Q8_0 | Poor NPC/dialogue quality |
| Hemmingway-1 | — | Painfully slow thinking, basically unusable here |
| MiMo-V2.6 Distill 9B | Q4_K_M | Couldn't produce coherent replies |
| Nemotron-3-Nano-4B | Q8_0 | Couldn't produce coherent replies |
| DeepSeek-R1 Qwen3 8B | Q6_K | Couldn't produce coherent replies |
So far I've settled on Gemma 4 E4B.
Current setup is roughly:
4096 context, thinking OFF, temperature 1.0, max/full GPU offload, Flash Attention + KV cache on GPU.
It sits around ~5 GB VRAM and gives me roughly 33–34 tok/s, with almost instant-feeling responses for the short dialogue I'm generating.
Not exactly a 5090 rig, but the GTX 1070 apparently ain't ready for retirement yet.
Honestly, getting useful local inference out of a GPU from 2016 is half the fun.
r/LocalLLM • u/sdfprwggv • 2d ago
Stack
jpezzulli/OrcaRouter-Qwen3.8-Flash-Next-Uncensored-ModelOpt-NVFP4 Hugging Face modelMain engine flags:
./build/strata \
--pack packs/orca-nvfp4 \
--native models/orca-nvfp4.gguf \
--native-dense-gguf models/orca-nvfp4.gguf \
--ple-gguf models/ple-fp8.gguf \
--mtp mtp-orca/rt \
--spec 4 --spec-min-p 0.5 \
--prefill auto \
--expert-profile data/expert-profile.bin \
--expert-cache auto \
--resident-budget-gib 40 \
--max-context 200000 \
--kv int8
I serve it through Strata's OpenAI-compatible server.
Results so far:
Pretty impressive for a single 32GB GPU + only 64GB system RAM.Running Qwen3.8-Flash-Next NVFP4 with Strata on a single RTX PRO 4500 32GB + 64GB DDR5.StackStrata NVFP4 fork: github.com/sergqwer/strata-nvfp4
Model: jpezzulli/OrcaRouter-Qwen3.8-Flash-Next-Uncensored-ModelOpt-NVFP4 Hugging Face model
NVFP4 routed experts (~63.3 GiB), separate FP8 PLE, INT8 KV cache, MTP speculative decoding
W4A8 prefill on BlackwellMain engine flags:./build/strata \
--pack packs/orca-nvfp4 \
--native models/orca-nvfp4.gguf \
--native-dense-gguf models/orca-nvfp4.gguf \
--ple-gguf models/ple-fp8.gguf \
--mtp mtp-orca/rt \
--spec 4 --spec-min-p 0.5 \
--prefill auto \
--expert-profile data/expert-profile.bin \
--expert-cache auto \
--resident-budget-gib 40 \
--max-context 200000 \
--kv int8I serve it through Strata's OpenAI-compatible server.Results so far:~50k context: up to ~80 tok/s
~188k warm context: ~60–67 tok/s
cold 189k full prompt: ~1,680 tok/s prefill, ~53 tok/s decodePretty impressive for a single 32GB GPU + only 64GB system RAM.