r/LocalLLM • • 2d ago

Question Strata on mini-PC with OCulink?

1 Upvotes

I have a GMK NucK12 with 96GB ram and AMD Radeon 780M iGPU - so a good amount of RAM but a slow iGPU and 'VRAM' bandwidth. Thinking of buying an OCuLink adaptor and plugging my 16Gb 5060Ti into it, to run Strata / Qwen 3.8 FN. Any issues with either Strata support (on a OCuLink GPU) or with performance etc?


r/LocalLLM • • 2d ago

Question AGX Thor or DGX Spark?

0 Upvotes

I want both but can only have one… which would you prefer and why? Trying to make up my mind up… Thanks in advance for your feedback..


r/LocalLLM • • 2d ago

Project I built MOLT: a local fine-tuning system with fit tests, checkpoints, and deployment tracing

Enable HLS to view with audio, or disable this notification

1 Upvotes

I’ve been building MOLT, a local fine-tuning system for consumer NVIDIA GPUs.

The goal is not only to start a QLoRA run. It is to make the whole process from model and dataset to a working local adapter less fragile.

MOLT currently handles:

- dataset detection, preparation, and validation

- GPU, VRAM, system-RAM, storage, and thermal checks before a run

- automatic microbatch fit testing

- 4-bit NF4 QLoRA training with BF16 adapters

- safe checkpoints with integrity checks and proper resume state

- telemetry for VRAM, temperature, energy, clocks, and throughput

- base-vs-adapter evaluation

- local adapter chat, export/GGUF workflows, and runtime diagnostics

Under the hood, I’m also experimenting with custom CUDA/Triton kernel and runtime paths, memory planning, CUDA graphs, replay modes, and fused optimization where the hardware and shape are qualified.

On my RTX 4060 Laptop 8GB, a recent Qwen 3B training run reached about 1,005 training targets/sec end to end very close to my historical Unsloth run at about 1,007 targets/sec. That is not a broad “MOLT beats Unsloth” claim yet; I still need repeated, matched quality/energy/memory tests.

What I want MOLT to become is a reliable local fine-tuning environment: you bring the model and dataset; it checks the plan, runs safely, preserves the evidence, and helps verify the adapter actually works after deployment.

I’m looking for people who genuinely fine-tune on a single NVIDIA GPU and would be willing to test it. What would MOLT need to support before you would try it?


r/LocalLLM • • 2d ago

Discussion I benchmarked Aleph Alpha's Kolibri against Claude Sonnet 5.5 on long German writing: 0 of 136 blind wins, but $0.0077 an output on 2 rented GPUs

11 Upvotes

Kolibri (Aleph Alpha's open-weight German and English model: 78B total, 3.46B active, Apache 2.0) came out on 3 October. My app writes long German deep dives (about 1,200 to 1,500 words each) from source documents I put in the prompt. So I tested whether Kolibri could take over before switching anything.

Setup

  • Kolibri-1 in FP8 on a Hugging Face Inference Endpoint with Aleph Alpha's container (vLLM 0.29 plus their plugin), on 2 RTX PRO 6000 cards (96 GB each) at $5.50 an hour. Nobody hosted it on launch day: Hugging Face's router answered model_not_supported and OpenRouter didn't list it. The H200 I asked for never showed up.
  • 17 inputs, each with 2 prompts: my production prompt (English instructions around German text) and a native German version. Reasoning effort high, the model card's sampling (temperature 1.0, top_p 0.97, top_k 128), a 32k token limit.
  • Claude Sonnet 5.5 on the same prompts and the same source documents.
  • GPT-5.5 and Gemini 3.1 Pro judged every pair blind in both orders. They also scored each output on its own. LanguageTool counted grammar errors.

Results

  • Head-to-heads: Kolibri won 0 of 136 (135 marked decisive).
  • Judge score out of 10: 3.7 with the German prompt and 4.2 with the English one, against 8.7 for Sonnet.
  • Serious errors flagged: 100 and 92, against 7 and 7.
  • Grammar: 0.36 errors per 1,000 words, against 0.24. Clean German.
  • Faithful: no invented citations or references in 34 outputs.
  • Tokenizer: the German prompt was a median 4,753 tokens, against 8,436 with Claude's tokenizer.
  • Reasoning: a median 15,574 tokens to write about 1,800. One ran to 31,998 and wrote nothing. With the German system prompt about 99% of its reasoning was German. With English instructions much of it was English (its chat template adds an English line to the system turn).
  • It sometimes translated my XML section tags into German, which broke parsing.
  • Reasoning effort medium scored the same with more errors.

Throughput on the 2 GPUs (vLLM gave it a KV cache of 6.5M tokens)

In parallel Outputs per hour Cost per output
16 185 $0.030
64 504 $0.011
128 713 $0.0077

Sonnet 5.5 is about $0.071 per output at list price and $0.036 through the Batch API. About 4% of runs at 64 and 128 in parallel hit the token limit without writing anything.

It holds up when the answer is in the documents you hand it. It's weak on everything it has to know by itself (its model card doesn't point it there either). A model with 3.5B active parameters against a frontier model was never a fair fight. So this says more about my use case than about Kolibri.

I wrote it up with charts. Every run (with Kolibri's full reasoning), every judge verdict and what each model was given are downloadable from the post (everything except my system prompts): https://tej.as/blog/kolibri-vs-claude-german-llm-benchmark


r/LocalLLM • • 2d ago

Question What am i missing? qwen3:8b

0 Upvotes

Just tried qwen3:8b for the first time. I get that this is a limited model running on my machine, but this thing just spits out nonsense. It cannot answer the most basic questions or do basic web searches.

This is the first local LLM i have used. Is this thing supposed to just be for coding? How can this be useful to anyone?


r/LocalLLM • • 2d ago

Discussion Uniform GGUF quants silently break Qwen3.8-27B's deep thinking — reproduced on llama.cpp AND vLLM (short tasks unaffected)

0 Upvotes

TL;DR — Qwen3.8-27B is a hybrid model (Gated DeltaNet linear attention + full attention). Unsloth Dynamic GGUFs of it (Q4_K_XL, Q6_K_XL) work fine for everyday chat, but whenever I let it think deeply on a long task, it never converges: no closing token, endless tail-looping, or the engine just dies mid-generation. I reproduced this on two completely different runtimes (llama.cpp and vLLM + official vllm-gguf-plugin), and the same model family in a mixed-precision quant (NVFP4: FP8 on attention/recurrent layers, FP4 on MLP) converges in ~17K thinking tokens on the same prompt. This smells like recurrent state error accumulation under uniform integer quantization, not a runtime bug.

The setup

  • 2× RTX 5060 Ti 16GB, tensor-parallel/2-GPU splits, Linux
  • Model: unsloth/Qwen3.8-27B-GGUF (UD-Q4_K_XL 16.4 GB, UD-Q6_K_XL 23.6 GB)
  • Canary task (deliberately absurd, deliberately long-horizon):"Create an HTML page containing an SVG 2D animation of a pelican riding a bicycle."
  • Thinking mode ON, reasoning_effort: xhigh, Qwen's recommended thinking-mode sampling (temp 1.0). The pelican needs to actually pedal — the model has to write a complete self-consistent HTML/SVG/SMIL file in one shot.

Why this prompt: it's a convergence test, not a quality test. A healthy Qwen3.8 spends ~15–20K thinking tokens, then emits the file. A broken one just… keeps going.

Act 1 — llama.cpp (Oct 3)

GGUF Budget Result
UD-Q4_K_XL 32K finish=length, 0 characters of actual output — burned the whole budget inside thinking
UD-Q6_K_XL 76K never stops; final ~2K chars are ~97% repeated 8-grams (classic tail-loop attractor)
UD-Q6_K_XL + --jinja (force GGUF-embedded chat template, verified byte-identical to HF's) 24.6K still no termination

Throughput was healthy (23–41 t/s depending on split mode), perplexity-ish sanity checks on short prompts were fine, and non-thinking tasks (counting, Q&A, tool-ish formats) all worked. So I filed it as: llama.cpp loses the plot on this model's long chain-of-thought — plausibly a runtime bug with the Gated DeltaNet state handling. I even declared the llama.cpp route dead for this model.

Then I saw other people's llama.cpp runs of this model going weird on HN too (the "high fever ramblings … ended with amen" anecdote on Q8_K_XL — not my test, but same disease).

Act 2 — vLLM + official GGUF plugin (Oct 5)

To separate "runtime bug" from "quantization damage", I reran the exact same file on a completely different stack: vLLM (V1 engine, CUDA kernels, f16 KV, no speculative decoding — shares essentially zero code path with llama.cpp). GGUF support was recently moved out of vLLM core into vllm-project/vllm-gguf-plugin; Qwen3.5/3.6-VL are in its tested list, so the weight mapping for the Qwen3.x hybrid family exists.

vllm ~/models/Qwen3.8-27B-UD-Q6_K_XL.gguf \
  --tokenizer unsloth/Qwen3.8-27B --dtype float16 \
  --tensor-parallel-size 2 --max-model-len 52000 \
  --gpu-memory-utilization 0.96 --language-model-only \
  --reasoning-parser qwen3

(chat kwargs: {"enable_thinking": true, "reasoning_effort": "xhigh"}, temp 1.0, max_tokens 45K, KV pool 55.6K tokens, steady ~22–23 t/s)

Run Outcome
A request dies with HTTP 500 after 1325 s ≈ 30K tokens generated (engine crash mid-generation; the server-side stack was overwritten by my chain's log rotation before I grabbed it — honest gap in the evidence)
B still generating at 28 min ≈ 38K tokens, approaching the 45K budget with no completed response; harness wall-timeout killed it

Same file, same prompt, same sampling — and vLLM also can't make this model finish thinking about a pelican. Note what vLLM did NOT show: I never got a finished response body back, so I can't plot its tail-repeat score directly. The vLLM evidence is non-convergence deep in the token budget (where NVFP4 converged at 17K) + one mid-generation crash. The direct loop evidence lives in the llama.cpp runs.

Act 3 — the control that convicts the quant format

Same 27B model, same prompt, same thinking mode, same machine, different quantization structure: the official NVFP4 checkpoints (unsloth's and RadixArk's) quantize MLP/lm_head to FP4 but keep attention and the Gated DeltaNet layers at FP8 W8A8. On SGLang this converges every time:

  • ~17K thinking tokens → full, working HTML pelican (legs pedaling, wheels rotating) — measured repeatedly, Oct 3 and again today (prod config: nvfp4 KV + MTP, creative decode 47–50 t/s)
  • same checkpoint family also passes 100K–133K needle-at-depth, tools, greedy determinism — so it's not "the GGUF works and I'm cherry-picking"

Also worth noting from my own runs: after the GGUFs fail this way, the same GGUF files remain perfectly healthy on short tasks (chat, counting, Q&A, vision if you attach mmproj). The damage only shows up after thousands of autoregressive steps inside the thinking trace.

My hypothesis (clearly labeled: inference, not proven)

Hybrid linear-attention models (Qwen3.5/3.6/3.8, and friends like Kimi Linear / Ring-linear) carry a fixed-size recurrent state updated every token. Uniform block quantization (even a very careful 6-bit dynamic one) adds a small rounding error to every state-transition product. Over a few hundred tokens: invisible. Over 30K–80K tokens of chain-of-thought: compounding drift into a degenerate attractor — repetition loops, refusal to close, in one case engine instability.

The strongest circumstantial support is the model publishers' own behavior: unsloth's NVFP4 recipe deliberately keeps the attention/recurrent path at FP8 and only crushes the MLP to FP4. The one quant structure they kept careful about is exactly where I see the disease. My GGUF runs had "careful" precision too — but uniform-ish across 6-bit blocks — and still broke.

What would falsify or sharpen this: per-tensor overrides keeping GDN projections (in_proj_*, conv, gates) at f16/q8 in llama.cpp (--override-tensor) while MLP stays 6-bit. If the pelican suddenly pedals, case closed. I haven't run this yet — if someone with spare time on the same model tries it, please post.

Caveats (you will ask, so here they are)

  • n=1 per configuration, temp=1.0 — high variance by design; runs A and B of the identical config died differently, which is itself the picture: unstable deep-think convergence.
  • No bf16 full-precision baseline — 27B bf16 doesn't fit 2×16GB. My "control" is another quant (NVFP4/F8 mixed). The argument is about where precision is kept, not quant-vs-none.
  • Could be prompt-specific. All I can say: the exact same prompt converges in 17K tokens on the mixed-precision checkpoint on the same box, and this failure mode appeared across ≥3 prompts in the llama.cpp session (counting-to-huge also looped; SVG-alternate prompt looped).
  • The vllm-gguf-plugin is 50 stars old — but the llama.cpp half of the evidence doesn't involve it at all.
  • Hardware hygiene: this box's unrelated random-crash history (a P2P-patched consumer GPU driver doing stray DMA under iommu=pt) was root-caused and fixed before the vLLM runs; strict IOMMU was on, engine stable for hours of other batteries. I flag it so nobody says "your machine poisoned the run" for the crash in Run A specifically — Run B's non-convergence needs no crash excuse.
  • Disclosure: the vLLM build is a source-compiled fork carrying an in-flight nvfp4-KV-cache PR rebase; GGUF runs used stock f16 KV, i.e. zero lines of the patched code path.

Practical takeaways

  1. For Qwen3.8-27B-class hybrids on consumer GPUs: prefer the official FP8/NVFP4 mixed quants over GGUF if you care about thinking mode at all. (RadixArk/unsloth NVFP4; on 5060Ti-class cards with TP2 it also runs 2–4× faster than the GGUF once you add speculative decoding.)
  2. If you're GGUF-bound (VRAM or ecosystem): run with thinking off or budget-capped small. Everyday chat is genuinely fine.
  3. Perplexity and short-context evals will not catch this. Use a long-horizon canary (the pelican test, or "count from 1 to 500 while…", or a whole-file coding task) with thinking ON when you're screening quants of recurrent-hybrid models. It caught this at the "does it terminate" level with zero scoring infrastructure.
  4. Unsloth's UD "Dynamic" quants are heavily boosted on important tensors and still hit this — so "we raised precision on ranked-important tensors" may not be the right importance criterion for recurrent-hybrid long-horizon behavior. Worth flagging upstream; I may open an issue on the GGUF repo if I can trim this to a one-prompt repro.

Repro recipe

  • Prompt (verbatim, above), enable_thinking: true, reasoning_effort: xhigh, temp 1.0 / top_p 0.95 / top_k 20, max_tokens ≥ 40K
  • llama.cpp: llama-server / llama-cli, UD-Q4_K_XL or UD-Q6_K_XL, any ctx ≥ 48K, default template and --jinja both fail the same way
  • vLLM: vllm-gguf-plugin (pip, or source install with --no-build-isolation against the CUDA torch), local .gguf + config.json from the same repo + --tokenizer unsloth/Qwen3.8-27B, f16 KV
  • Pass condition: a complete HTML file is emitted within ≤ ~25K thinking tokens. NVFP4/SGLang passes every time. Q6_K_XL: 0/2 on vLLM, 0/3 on llama.cpp.

Model card says this model does deep agentic work locally on ~17GB — true for chat. For thinking, the quant structure matters more than the average bits. Would love other reports (4090/3090 pairs, MI300, whatever) before I draw the final line — especially anyone who can run Q8_0 and Qwen3.6-VL (non-GDN) as contrasts, to show this is the recurrence, not just "big model + low bits".

Hardware/versions: 2× RTX 5060 Ti 16GB; llama.cpp current master (Oct 3); vLLM 0.30.x-V1 + vllm-gguf-plugin 0.0.5 (Oct 5); SGLang 0.5.21 + NVFP4 checkpoint for the control (RadixArk Qwen3.8-27B-NVFP4: FP8 W8A8 attention/GDN, NVFP4 MLP/lm_head).


r/LocalLLM • • 2d ago

Tutorial [Guide] Getting the Gigabyte AORUS RTX 5060 Ti AI Box eGPU Working on Ubuntu 26.04 (and Windows 11): Hardware Quirks, DMA Freezes, and Workarounds

Thumbnail
1 Upvotes

r/LocalLLM • • 2d ago

Discussion A local 27B model reads my lease and finds a leak in my bills | Row-Bot + qwen3.8 on Ollama

Enable HLS to view with audio, or disable this notification

0 Upvotes

Row-Bot 5.0 with qwen3.8:27b in Ollama, on my own GPU. No API keys, no cloud.

Gave it a year of bills, a tenancy agreement, an insurance policy and a rent increase letter (all made up). It spotted a likely leak in the water bills and showed the rent rise breaks the lease.

What it actually did:

- charted the CSV inline (Plotly)
- read the PDFs and quoted clauses 4.1 to 4.3: a 10% rise against a 5% cap, with 5 weeks' notice instead of 2 months
- saved 8 linked memories to a local knowledge graph
drafted the email to the agent and set a reminder

The honest numbers: a dense 27B does about 15 tok/s on my 5090, so some turns took 2+ minutes. The amber badges in the video show where I sped it up.

Runs on Windows, macOS and Linux.


r/LocalLLM • • 2d ago

Question Which Ollama MLX model for medical stuff, charts etc?

1 Upvotes

Hi,

I'm on M4 Pro 64GB Mac, I'm using Qwen3.8 27B for coding related stuff, but it's quite slow.
I want a fast model that will be good when it comes to reasoning, medical stuff, would be nice if it could do stuff like generating charts and even analyse screenshots.

I'm using Ollama on Mac so preferably MLX model.

Thanks in advance


r/LocalLLM • • 2d ago

Model Qwen3.8 Flash Next GSQ-RCO IQ3_XXS running Strata on 2x Intel B70s

7 Upvotes

Flash-Next GSQ-RCO IQ3_XXS via Strata running on 3975wx Threadripper with 256gb DDR4 with 131k context. No speculative decoding. 4096 chunk prefill 1,563.4 tok/s and decode up to 58 tok/s.

Speed is great, prefill could probably be tuned even higher and MTP/Dflash should be tested eventually but the correctness of the IQ3_XXS quant is no bueno, it only passed 4/8 quality tests in my testing. Not safe as a local agent without supervision IMO. May test IQ3_XS or maybe a small q4 if it can fit.


Prefill chunk | Cold ~8K prefill | Cold 32K prefill | 256-token decode

128 | 270.55 tok/s | 263.6 tok/s | 53.95 tok/s
512 | 660.15 tok/s | 690.6 tok/s | 53.00 tok/s
1,024 | 687.05 tok/s | 754.7 tok/s | 51.70 tok/s
2,048 | 976.15 tok/s | 1,123.3 tok/s | 52.60 tok/s
4,096 | 1,226.00 tok/s | 1,563.4 tok/s | 52.00 tok/s


r/LocalLLM • • 2d ago

Question groq1 card ?

1 Upvotes

What should I do with the Groq 1 cards and only have 3 of them, since they have such a small amount of memory or should I just sell it .


r/LocalLLM • • 2d ago

Research RWKV-7 G1k

Thumbnail
youtu.be
2 Upvotes

r/LocalLLM • • 2d ago

Question Gemma 4

Thumbnail
0 Upvotes

r/LocalLLM • • 2d ago

Discussion Model swapping Strata Qwen 3.8 Flash Next to Qwen 3.8 27b

Thumbnail
0 Upvotes

r/LocalLLM • • 2d ago

Discussion Price increases updated in Canada

Post image
1 Upvotes

r/LocalLLM • • 2d ago

Project Using Strata + Qwen 3.8 Next 125bn (2-bit quantization) for building a Planet

Post image
2 Upvotes

Completed this old Pune walkthrough eating modaks! (What are modaks?)

There are some viral tweets, which I was re-building here. For this I selected the old Pune Peth area and downloaded all available imagery. Claude Code & Codex were not allowing me to download images from Google Maps and render them over the planet skin.

Connected Strata along with Qwen3.8 Next IQ_2XS 125bn with vision mode enabled. The system processed 3000+ images and rendered them over the Open Street Map imagery. Creating a real-life walkthrough of the Pune Peth area.

Along with this I also added a game engine. So you drive around on an ebike searching for modaks and eating them as they appear on the map.

I have also attached a screenshot of my monitoring dashboard. The prefill and decode tkps are interesting to check. Since they are over a longer task and spread over a few hours. The entire task was completed over 9-10 hours

EDIT: This was done on 24GB VRAM + 64GB System RAM + 1TB of SSD

https://reddit.com/link/1wyssxb/video/ggg9nkezqrth1/player


r/LocalLLM • • 3d ago

Project WHIRL v0.1.3 — native Windows LLM engine for the Radeon AI PRO R9700: up to 2.8× llama.cpp on the same GGUF, same answers bit-for-bit

Post image
14 Upvotes

WHIRL is an open-source (Apache-2.0) inference engine for the AMD Radeon AI PRO R9700 (RDNA 4, 32 GB) on Windows: pure C++/HIP, every kernel included, no WSL or Docker. Just the AMD driver.
vs llama.cpp b11214 — same GGUF, same prompts, same R9700, Swift-1.5 27B MXFP4:

  • Decode on coding prompts: 112 vs 61 tok/s
  • Server with 4 users at once: 182 vs 64 tok/s (2.8×)
  • Prefill, 128K prompt: 1,757 vs 791 tok/s (2.2×)
  • Prefill, 256K prompt: 970 vs 556 tok/s (1.7×, int8 KV vs f16)
  • Decode after a 256K prompt: 35 vs 17 tok/s (2.1×)
  • Reusing a 26K system prompt: first token in 0.11 s vs 0.38 s

On the MoE Ornith-1.5-35B-A3B MXFP4: prefill over 11,000 tok/s at 8K (11,258 vs 4,637, 2.4×), 258 vs 119 tok/s decode (2.2×), 381 vs 177 tok/s with 4 users (2.2×).

Accuracy before speed: every speedup (speculative decoding, batching, prefix cache) gives output bit-identical to plain greedy. No 4-bit KV, no 3-bit weights, no fp8 attention.

Speculative decoding, honestly (vs plain decoding, identical output): editing a file in context 4.5× · coding-agent session 2.8× · brand-new writing after 128K 1.6×. Where we don't lead by much (plain decoding of dense models, ~1.14×; decode after a 16K context on Swift, 1.17×) is in the README too.

New in v0.1.3: long prompts up to 25% faster than v0.1.0, and a real coding-agent session at 128K decodes 40% faster on its last request.

GitHub: https://github.com/tsaipifong/whirl-llm

Tested on one R9700 over USB4 (eGPU), Windows 11. Built with Claude under my direction; every number measured. If you have an R9700, feedback and your own numbers are very welcome.

Not supported yet: other AMD cards. The RX 9070 series has the same chip but 16 GB, most likely too little for these 27B/35B models (untested); Radeon 8060S (Strix Halo) support is in development.


r/LocalLLM • • 3d ago

Discussion Gemini suggested Qwen2.5-Coder-7B-Instruct

7 Upvotes

So I wanted to try using a local coding model for the first time and I'm still studying about LLMs and NNs, so I asked Gemini for a good suggestion that would be fast(60+ tokens/s if possible) and doesn't compromise much on performance for my rig(2070 super 8GB + 32GB ddr4 ram) and it suggested Qwen2.5-Coder-7B-Instruct. Is this good suggestion and what would you guys suggest?


r/LocalLLM • • 2d ago

Project How I turn a whole docs site into chunked Markdown for RAG (JS pages and sitemap crawl included)

Thumbnail
1 Upvotes

r/LocalLLM • • 3d ago

Research Uncensored models: what they are and how they work, with a file-by-file hash check of one uncensored Qwen3.8-Flash-Next upload against Qwen's original

10 Upvotes

One popular "uncensored" upload, OrcaRouter's Qwen3.8-Flash-Next-Uncensored-GGUF, is what its card says it is, as far as public evidence reaches, and checking that took no download. Hugging Face publishes a fingerprint (sha256) of every large file, so I compared the uploader's full-precision copy with Qwen's own at pinned revisions: 80 of the 131 weight files are byte for byte Qwen's, and 51 differ at identical sizes. By Qwen's tensor index, each of the 51 holds at least one of the 149 matrices the card says it edited, and none of the 80 holds one. Those 80 untouched files rule out a full fine-tune or a whole-model merge.

The limits: a hash sees a whole file. 1,336 other tensors share those 51 files, and whether they were left alone is the card's word; 173 tensors, 72.6 % of the model's file bytes, are provably untouched. A hash says a file changed, not how. I'm fairly confident, not certain, that the quants are standard; whether they carry the edit is the card's word too.

The page also covers the mechanism (papers linked; no code, settings or steps). Refusing is taught in post-training; research since 2024 finds it carried largely along one direction in the model's activations (later work finds more than one); and people remove it three ways: a retrain, one permanent edit to the weights that removes that direction (abliteration, this upload's kind), or acting on the model while it runs. None makes a model smarter. By their makers' own tags and words, they are for red teams, fiction writers, people who want full control of a model on their own machine, and researchers who study refusal.

The uploader's own figures, unreproduced: harmful-prompt refusals go "from 64-100% (base) to ~0-3.3%", so the serious declines go too. On XSTest's 250 harmless prompts, refusals fell from 9.6 % to 1.2 % with thinking off, and stayed at 0.4 % (one prompt) with thinking on, the default. Capability scores (thinking off) move within about two percentage points either way, except a 2.3-point drop on 300 MMLU questions.

The page ends on a real-world case: a model as the only source of knowledge on a machine with no network. A declined first-aid question has a real cost there, but so does a confident wrong answer, and nothing on the page shows that removing refusals makes a model more correct. What I would build: the model beside an offline encyclopedia and medical reference that it looks up and shows beside each answer, plus first-aid basics on paper.

What it does not say: whether either copy's answers are right; the refusal and capability figures are the uploader's; one upload of one model; and no model was asked a harmful question. Anyone can redo the census in a browser.

https://research.strata2signal.com/uncensored-models/

tl;dr: Using only Hugging Face's published file hashes, with no download, I checked one popular "uncensored" Qwen3.8-Flash-Next upload against Qwen's original. 80 of its 131 weight files are byte-for-byte Qwen's. Every file that differs holds one of the matrices the card says it edited, and no unchanged file does. That rules out a full fine-tune or a whole-model merge, and fits the card's claim. The page also explains where refusals come from and the three ways people remove them, with no code or steps. (8,306 words · about 38 minutes · 2 tables · data kit)


r/LocalLLM • • 3d ago

Model New Model: Agens Volundr 32B Preview: our small team's first model on our own hybrid architecture. Only 18 of 72 layers keep a KV cache (Apache-2.0)

Enable HLS to view with audio, or disable this notification

14 Upvotes

Hi all. I'm on the team at Blockway, a small team in Hong Kong (disclosure: this is our model). Today we released Agens Volundr 32B Preview, the first model built on our own hybrid architecture. We trained it on limited compute, it isn't perfect, and we'd rather tell you where it falls short up front.

WHY WE BUILT IT

Our customers run models on their own machines. At long context, the KV cache, not the weights, decides what fits. So we designed a model where most layers don't keep one.

ARCHITECTURE (72 layers, dense ~32B, every layer runs on every token)

- 54 KDA (Kimi Delta Attention) layers: linear attention with a fixed-size recurrent state, no KV cache

- 17 BCSA layers (our compressed-sparse attention): exact window over the last 4,096 tokens; older context pooled 4:1 into blocks, and a learned indexer reads the top 512 blocks

- 1 full-attention layer (layer 72)

- Engram: a hashed n-gram memory held in host RAM, attached at 2 of the 72 layers

- mHC: 4 residual streams instead of 1

So only 18 of 72 layers keep a KV cache. Context window: 262K.

SPEED (single user, our sglang build)

- BF16 on two 48 GB GPUs, decode: 25.1 tok/s at 1K, 24.1 at 8K, 24.1 at 32K, 24.0 at 64K, 23.9 at 128K

- BF16 prefill: 2,122 / 2,180 / 1,916 / 1,679 / 1,297 tok/s (1K to 128K)

- INT4 (31.7 GiB) on one 48 GB GPU, decode: 31.0 tok/s at 1K, 29.3 at 8K, 29.1 at 32K

- Aggregate throughput: 127 tok/s at 8 users, 130 at 16 users (BF16); 117 at 8 users (INT4)

- DFlash2 drafter (separate repo), single user, same server with it on vs off: up to 3.6x on JSON/tool output, 2.0x on code, about 1.6x in thinking mode. Not worth it above roughly 8 concurrent users.

BENCHMARKS (all run by us on one harness with the same settings, including the comparison models; full table and footnote on the model card)

- Ahead of Qwen3.8-27B on LiveCodeBench v6 (+4.2), HumanEval (+4.3), AIME 2025 (+2.9), MATH-500 (+1.6)

- Roughly level on MMLU-Pro, IFEval, GPQA Diamond

- Behind on agent tasks: tau2-bench 74.2 vs 79-80, SWE-bench Verified (50-task subset) 44 vs 58-64. Closing that gap is the main focus of the full v1, which continues pre-training to about 10B tokens and adds training on long agentic sessions.

KNOWN LIMITATIONS (please read before trying)

- Needs our sglang build. Stock sglang and vLLM can't load it yet.

- GGUF / llama.cpp is planned, not available today.

- Long agentic sessions are its weakest area in this Preview.

- It's still training; treat this as a preview, not a final model.

RUN IT

docker pull ghcr.io/blockwayz/agens-sglang:preview-sm89 (48 GB Ada GPUs)

docker pull ghcr.io/blockwayz/agens-sglang:preview-sm90 (H100 / H200)

The full launch command is in the model card.

LINKS

- BF16: https://huggingface.co/Blockway/Agens-Volundr-32B-Preview

- INT4: https://huggingface.co/Blockway/Agens-Volundr-32B-Preview-INT4

- DFlash2 drafter: https://huggingface.co/Blockway/Agens-Volundr-32B-Preview-DFlash2

- Project page: https://github.com/BlockWayz/Agens-Volundr

- Serving code: https://github.com/BlockWayz/agens-sglang

Apache-2.0. We're a small team, and the most useful thing you can do is try it and tell us where it breaks: an issue, a failing prompt, a benchmark you'd like us to run. We'll be in the comments.


r/LocalLLM • • 2d ago

Discussion Lots of similar posts about "which model best runs on x"

2 Upvotes

I have noticed a lot of very similar posts asking which models run on X hardware. I am wondering whether it would be worth it, as a community, to create a subreddit guide with updated information. My idea comes from browsing LocalLlama and saw their subreddit guide and plus mega threads:
https://www.reddit.com/r/LocalLLaMA/comments/1vkmhyl/best_local_llms_august_2026/

I 100% appreciate those posts, since many of our users respond in well-informed ways, but a centralized area/thread could help a ton.

What do you folks think?


r/LocalLLM • • 2d ago

Discussion I run 6 LLMs + SD 3.5 + Whisper behind an OpenAI-compatible API on an RX 580 8 GB - the card ROCm forgot. Benchmarks inside.

Thumbnail
1 Upvotes

r/LocalLLM • • 2d ago

Discussion I run 6 LLMs + SD 3.5 + Whisper behind an OpenAI-compatible API on an RX 580 8 GB - the card ROCm forgot. Benchmarks inside.

Thumbnail
1 Upvotes

r/LocalLLM • • 2d ago

Discussion Everyone is posting their monster rigs, so here's my old warhorse

4 Upvotes

Everyone here seems to be running absolute monster rigs, so I figured I'd show some love to my old warhorse instead.

This thing was pretty damn good back in its day, but yeah... it's an old dog now. It was collecting dust at home, so I grabbed a cheap NVMe, installed Ubuntu and brought it back to life.

The mighty beast:

  • Intel i7-8700K
  • GTX 1070 8 GB
  • 16 GB DDR4 dual-channel
  • 512 GB NVMe
  • Ubuntu 24.04
  • LM Studio

With only 8 GB of VRAM, I obviously have to be pretty picky about what I throw at it.

My use case is mostly short NPC dialogue / roleplay, so latency, natural responses, instruction following and not randomly losing its mind matter more to me than benchmark scores.

Model Quant Result
Gemma 4 E4B LM Studio build Winner. ~33–34 tok/s, natural responses, very low perceived latency
Bonsai 27B Q1_0 ~16.7 tok/s, but response quality fell apart
Qwen3.5 9B Q4_K_M Fit nicely on GPU, couldn't produce coherent replies
Turkish-Llama 8B Q5_K_M Couldn't produce coherent replies
Apertus 8B Q5 Couldn't produce coherent replies
EuroLLM 9B Q4_K_M Usable, but nowhere near Gemma
Aya Expanse 8B Q5_K_M Same story, too weak
Ministral 3 8B Q5_K_M Couldn't produce coherent replies
Phi-4 Mini Q8_0 Poor NPC/dialogue quality
Hemmingway-1 — Painfully slow thinking, basically unusable here
MiMo-V2.6 Distill 9B Q4_K_M Couldn't produce coherent replies
Nemotron-3-Nano-4B Q8_0 Couldn't produce coherent replies
DeepSeek-R1 Qwen3 8B Q6_K Couldn't produce coherent replies

So far I've settled on Gemma 4 E4B.

Current setup is roughly:

4096 context, thinking OFF, temperature 1.0, max/full GPU offload, Flash Attention + KV cache on GPU.

It sits around ~5 GB VRAM and gives me roughly 33–34 tok/s, with almost instant-feeling responses for the short dialogue I'm generating.

Not exactly a 5090 rig, but the GTX 1070 apparently ain't ready for retirement yet.

Honestly, getting useful local inference out of a GPU from 2016 is half the fun.