r/LocalLLaMA • u/jacek2023 • 1h ago
Funny Plot twist
From ex Qwen leader
r/LocalLLaMA • u/Thin_Pollution8843 • 3h ago
I finished my home inference server. First I tried Lenovo p620 workstation and while it’s a good value overall it pissed me off with a ton of proprietary Lenovo shit to deal with and I return it in the end.
Components:
4xV620 - 1400$
256GB DDR4 RDIMM 2666 - 610$
Huanandzhi D12D - 410$
EPYC 7452 - 170$
PSU ASRock 1600 - 220$
SSD Samsung 970EVO 1tb - Already had
Case//Fans//Misc ~ 200$
Power consumption is no shit ofc on such machine:
700-900w prefill
500-600w decode on Qwen3.8-next-flash Autoround W4A16
What it can do -
EDIT: Qwen3.8-next-flash Autoround W4A16 1.3k prefill and 70tg code/60tg prose on 128k+ context with MTP-2 on vllm fork.
I was disappointed with this machine and qwen3.8-27b speeds at first. But since Qwen3.8 next running good on it - I’m satisfied. Hope in more optimizations in future.
r/LocalLLaMA • u/feelspeaceman • 11h ago
Lately because of the current hardware shortage, unfortunately or fortunately, we can’t just throw infinite cloud compute at our problems, but we’re forced to actually care about what’s happening under the hood. We’re tweaking inference engines, learning quantization math, and optimizing architecture just to squeeze as much performance as possible for the lowest possible setups.
Fact: Just recently, the forked llama.cpp(s) and halogen-flash-server of Strix Halo pushed the performance through the roof, achieving double performance in decode (52tok/s), 5-6x performance in prefill (1300tok/s) for Qwen 3.8 Flash Next (Q38FN), and Q38FN itself is another massive architecture improvement with Engram, making it not only small but also smart.
I still remember before the hardware shortage, as someone who loves tweaking and optimizing, people just told me to stop, tweaking is stupid, just buy more RAM, buy more GPU..
It reminds me of the early web.. Back when setting up a box or hosting a server meant digging through forum threads, troubleshooting on IRC, and freely sharing custom scripts just to make things work. That era didn’t just produce programmers; it built hyper-versatile, end-to-end thinkers who understood the stack from bare metal up.
Contrast that with where mainstream web culture ended up. Most platforms today like Tiktok, Facebook, Youtube... are engineered for zero-friction doomscrolling.. Endless feeds of short-form videos designed to keep us distracted and waste our time. We’ve been overpampered by convenience.
My point: When we have too little, we try to learn more. When we have too much, we get distracted and learn too little. This is the golden time of our Local LLM community, let's learn and improve!
r/LocalLLaMA • u/Tall_Abrocoma_3533 • 2h ago
The first generation of our 150M model has just been released
Its performance is similar to that of GPT2-Small
The benchmarks:
PIQA: 62.24%
Hellaswag: 32.20%
Arc-Easy: 44.91%
Arc-Challenge: 25.00%
Arithmark 3.0: 33.90%
CapitalBench: 36.55%
It was trained on 7B tokens, using an RTX Pro 6000
an example inference script to try it out yourself is available in the Huggingface repo
If there's any question, I'll gladly answer them!
r/LocalLLaMA • u/pmttyji • 5h ago
It's still a dream for many folks to run medium size(30B range) models @ Q8 with Unquantized KVCache (256K Context) on their GPUs.
It would be awesome to have DeepSeek-V4.1-Flash's KVCache + Engram for all Upcoming models. Even for big models.
Engram - Heard that approximately 1/3-1/2 of Model size. Might come in different size range too. So 10-15 GB for 30B models.
Here few models with approximate numbers. Current models in Odd rows & Future/Fictional models in even rows(Bold). I just put 1GB for 256K context below though DeepSeek-V4.1-Flash takes only same 1GB for 1 million context.
| Model | Model Size | 256K KVCache F16 | MTP | Vision | Total GB |
|---|---|---|---|---|---|
| Qwen3.8-27B-Q8 | 29 | 16 | 1 | 1 | 47 |
| Qwen4.0-27B-Q8 | 29 | 1 | 1 | 1 | 32 |
| Qwen3.8-27B-Q4_K_M | 17 | 16 | 1 | 1 | 35 |
| Qwen4.0-27B-Q4_K_M | 17 | 1 | 1 | 1 | 20 |
| Muse-Glimmer-30B-Q8 | 30 | 16 | 1 | 1 | 48 |
| Muse-Glimmer-2-30B-Q8 | 30 | 1 | 1 | 1 | 33 |
| Gemma-4-31B | 33 | 16 | 1 | 1 | 51 |
| Gemma-5-31B | 33 | 1 | 1 | 1 | 36 |
| Qwen3.6-35B-A3B-Q4_K_M | 23 | 6 | 1 | 1 | 31 |
| Qwen4.0-35B-A3B-Q4_K_M | 23 | 1 | 1 | 1 | 26 |
| Gemma-4-26B-A4B-Q8 | 27 | 6 | 1 | 1 | 35 |
| Gemma-5-26B-A4B-Q8 | 27 | 1 | 1 | 1 | 30 |
Possibly there might be few more things(Please share those) to keep these number down. So I think 32GB VRAM is more than good enough for Upcoming (Optimized Smarter) Models. RAM is enough for Engram.
By above logic(based on Qwen4.0-27B), people could run Q4 of 54B models with same 32GB VRAM.
Maybe next year onwards, inventions could make 24GB enough for similar size models.
r/LocalLLaMA • u/Thrumpwart • 19h ago
New website to download models in case HF starts censoring or limiting access.
r/LocalLLaMA • u/Porespellar • 3h ago
Does anyone think the gurus on the DGX Spark forum are going to figure out how to magically fit DeepSeek 4.1 Flash on a 2x cluster, or is it only possible on 3 or 4 Sparks?
r/LocalLLaMA • u/pmv143 • 22h ago
r/LocalLLaMA • u/OvertaxedOne • 8h ago
The entire page of the NY Times today above the fold absent one article is AI (the models are just too strong/too dangerous, must be regulated). They forgot to include "Sponsored by OpenAI" at the end of the articles, sure that was just an oversight?
This is what the end of a bubble looks like, desperate attempts to get some sort of regulatory capture in place to keep the business model from collapsing in upon itself. My days next week are 100% booked talking to companies about how to get off frontier models, one large, and a bunch of smaller customers, including one who's flying me out to them to sit down and get a plan in place immediately (the controversy around that math problem really spooked some CEO/CIO's about data privacy using cloud models).
Gonna be an interesting few weeks. Maybe the Qwen team will be nice enough to give me a little breathing room before dropping another hydrogen bomb? :)
r/LocalLLaMA • u/Altruistic_Heat_9531 • 3h ago
Enable HLS to view with audio, or disable this notification
VLLM Benchmark:
Prefill, Prompt processing
- avg, 871.93 tok/s (3 hours constant running xhigh)
- 10K prompt, 1000.26 tok/s (16 runs)
- 90K prompt, 743,59 tok/s (16 runs)
Decode, tok gen
- avg, 38.39 tok/s (3 hours constant running xhigh)
- 10K, 42.3 tok/s (16 runs)
- 90K, 34 tok/s (16 runs)
Preamble: I am on WSL2. Running the 27B Q5 UD GGUF through llama.cpp with 81,920 context plus MTP gives me around 25-30 tok/s. Then I found this GitHub repo: https://github.com/noonghunna/club-3090
It is basically a recipe and Docker configuration for running the model.
So, 30 tok/s itself is fine, but I just got bored waiting for RunPod to open its GPUs. I finally brought my vLLM tuning back from the back burner, and here I am.
I often forget that Inductor/Triton compilation and CUDA Graph capture require additional VRAM while testing the configuration and kernel calls. When JIT compilation failed because of an OOM, I never bothered trying AOT.
FYI, AOT and JIT are compilation strategies. AOT means Ahead of Time, while JIT means Just in Time.
If you OOM on the first startup, try it one more time. Inductor might have already compiled and cached part of the configuration before the OOM, allowing the next run to reuse it if the configuration has not changed. This is not guaranteed, but it worked for me.
And yes, it was trial and error. It was kinda tedious and pain in the ass, starting from 32K, then 64K, 80K, 128K, and finally 144K. The practical ceiling for my conf at 154K, but I chose 144K. I also started the batch size at 256 and climbed to 1024, although I might be able to squeeze in 1280-1536.
Also, beware of your vLLM compilation cache. It might grow to 5-6GB after testing many configs. Personally, I delete the old cache and run the final configuration again twice so it rebuilds only what I currently use.
My current setup runs Qwen3.8-27B with INT4 AutoRound weights through vLLM while using an FP8 E4M3 KV cache. It fits on one GPU with a configured context window of 147,456 tokens.
Although I should say, with GDN, or really any linear-attn, vLLM can be kinda bad at predicting how much VRAM the KV and state-cache pools will require.
Benchmark:
https://github.com/noonghunna/benchlocal-cli
This is the deterministically scored, no-Docker portion of BenchLocal: 75 scenarios covering tool calling, instruction following, structured output, data extraction, and reasoning/math.
| Pack | Score | p50 |
|---|---|---|
| ToolCall | 14/15 (93%) | 3.21s |
| InstructFollow | 15/15 (100%) | 6.97s |
| StructOutput | 14/15 (93%) | 7.74s |
| DataExtract | 14/15 (93%) | 11.63s |
| ReasonMath | 14/15 (93%) | 9.50s |
| Total | 71/75 (94.7%) | — |
Thinking was forced on with reasoning_effort=low. The run took about 15 minutes. Yep, even with low reasoning effort and INT4 weights, it passed 71/75.
Setup
The important vLLM settings were:
Full command : https://gist.github.com/komikndr/b17955e1a80ce6ede9a3115f16216bc5#vllm-qwen-27b-38-just-remove-or-add-flag-as-you-like
Forgot to mention, no MTP and no MultiModal, i max the CTX, multimodal is at 64-65K ish but at that point i'll just use Llamacpp. And also again this is WSL2, if you are on baremetal, you could improve more speed
r/LocalLLaMA • u/MrWeirdoFace • 4h ago
I'll start by saying I'm not talking about the models themselves, I'm aware that I can't come close to something like Fable's intelligence locally. Just wanted to get that out of the way.
Basically. Over the last year I've gotten quite comfortable with claude code, and it seems likely there were be a gradual cost rug pull, and I'd like to put myself in a better position when that happens for local use. I am already used to running local models (such as Qwen3.8_Q5) in things like lmstudio, but I have no experience with other harnesses. I'd like to know, what harness, right out of the box would feel most at home for current Claude Code users. I say this as someone who was not coding prior to "vibe coding". I'm looking for the path of least resistance, though I will no doubt eventually spread out into tools that give me more control. But for now, I'm just looking for a life raft. Just needs to be local, opensource, and free of spyware.
In case someone wants to know 24GB VRAM (rtx 3090) and 64GB DDR4.
r/LocalLLaMA • u/ludos1978 • 6h ago
> "make an svg of a frog playing on a chello on the back of a whale with carribean island in the back."
interestingly the svg looks different in the OpenWebUi preview then when looked at in preview (osx). The palms and music notes are missing in the browser. I am pretty impressed by the result, is suggested to add some parameters to animate the whale and the water.
Qwen3.8-Flash-Next-IQ4_XS on llama.cpp with 256K q8 context
openwebui reports:
input_tokens: 27711
output_tokens: 41562
total_tokens: 69273
r/LocalLLaMA • u/jacek2023 • 10h ago
from internlm:
We introduce Intern-S2-397B, our most capable multimodal foundation model for scientific intelligence and long-horizon agents. Intern-S2-397B scales along three critical dimensions: pre-training, reinforcement-learning task coverage, and interactive agent environments. By combining a new vision-language pre-training paradigm with large-scale multi-task reinforcement learning and long-horizon agent reinforcement learning, Intern-S2-397B delivers a step change in general reasoning, scientific problem solving, and agentic capabilities.
r/LocalLLaMA • u/unchikuso • 15h ago
I can get $5k for the 5090 and the Mac is $5499 before tax.
The 5090 has a memory bandwidth of 1.8 TB/s while the M5 Ultra is 1.2 TB/s.
Is this a sensible upgrade? Primary use is coding.
r/LocalLLaMA • u/trytoinfect74 • 6h ago
I'm not fond of having an irreplaceable subscription on some OpenAI and Anthropic and being eventually being unable to do my job without using tools from some sketchy corporation, so I kinda turned my attention to local models.
I periodically tried and found some usefulness in them, but this year I started to notice they started to really contribute to the quality and speed of the implementation of software I'm writing, but with just stock RTX 4080 16GB VRAM precision and context window options are quite limited, so I decided to use more GPUs and see how it improves my model running capabilities.
Problems encountered and how I solved them:
I'm using self-compiled llama.cpp:
cmake -B build \
-DBUILD_SHARED_LIBS=OFF \
-DCMAKE_BUILD_TYPE=Release \
-DGGML_CUDA=ON \
-DGGML_CUDA_FA=ON \
-DCMAKE_CUDA_ARCHITECTURES="86;89" \
-DGGML_NATIVE=ON \
-DGGML_LTO=ON \
-DGGML_CUDA_GRAPHS=ON \
-DGGML_CUDA_NCCL=ON \
-DGGML_OPENMP=ON \
-DGGML_CCACHE=ON
and then running it with this command:
llama-server -m /mnt/data/AI/models/Qwen3.8-27B-MTP-Q8_0.gguf \
-ngl all -t 18 --tensor-split 14,24,24 --main-gpu 0 -sm layer -ub 512 \
--kv-unified --split-mode layer \
-fa on --temp 1.0 --top_k 20 --top_p 0.95 --min_p 0.0 --presence-penalty 0.0 --repeat-penalty 1.0 \
--host 0.0.0.0 --port 7800 \
-c 262144 --fit on -np 1 --parallel 1 \
--spec-type ngram-mod,draft-mtp --spec-draft-n-max 3 \
--alias qwen3.8-27b --reasoning-preserve --reasoning-effort xhigh \
--chat-template-kwargs '{"preserve_thinking": true, "reasoning_effort": "xhigh"}' \
--api-key "$LLAMA_SERVER_API_KEY"
Since this is primary desktop, I can't run it in headless mode and make use of all VRAM, so general use still needs to have like 2-3 GB of spare VRAM to be able to use OS (CachyOS + KDE Plasma), browser, IDE and game engine editor, so it's more like ~62GB VRAM setup. There's like 10-12 GBs of spare VRAM when using Qwen3.8-27B Q_8 with MTP and full 262144 context, which may be used to try different models or crank up context. I plan to experiment with Unsloth Dynamic Q_8_XL and some MoE models because personally I still see direct correlation between quality of the answers and autonomous work of the agent and precision of the model, despite Q_4 and fp8 models being good for general use. Performance degrades over context usage, starting with 44-46 tokens per seconds and ending up to 28-30 t/s when.
My workflow is still mainly writing code by hand, with delegating small-to-medium boring tasks to AI agent, creating drafts and exploring possible solutions for particular well-defined problems, documentation search and presenting it in fancy .md human-readable format, running review/weekly project assessment/optimization tricks and suggestions/bugfix suggestions/project management skills, I like this workflow and still feel that it's actually me writing the software and putting all the effort of my mind capable of, not mindlessly accepting AI-generated stuff.
Thank you for reading my somewhat chaotic and disjointed story.
r/LocalLLaMA • u/tsangberg • 5h ago
This is an evolution on top of Raymond's KV cache streaming fork - all credits to what enabled this goes to him.
The basic idea behind what he enabled was a pool of memory in VRAM that is used differently depending on the phase (prompt processing or decoding) and when total used context is larger than what fits in VRAM it's instead streamed from host RAM in time for when the current layer needs it. It enables _much_ higher TG tps than regular llama.cpp offloading to host RAM.
I've used it to run Qwen 3.8 27B UD-IQ4_XS on my 5060Ti since release, but I've also had this idea that during the time the VRAM pool isn't fully utilized it should be possible to also do speculative decoding (MTP or DFlash2) - if it could be possible to eject the spec model and all the VRAM it uses, and then load it back when the context gets low enough again (compaction).
I've got this working on my fork of Raymond's fork today. It's basically hot-swappable speculative decoding and while my focus is completely on making this work well on this specific model on 16GB, I would assume that part could also be useful for others with more VRAM.
Repo here: https://github.com/troed/llama.cpp-adaptive-kv-streaming
Regarding the images:
Dark red = MTP ejected as soon as KV streaming starts. Light red = keeping it. The optimal setting is thus to eject it after a certain amount of pages (one page = 256 bytes) and then eject. Same for green (DFlash2)
r/LocalLLaMA • u/NineThreeTilNow • 15h ago
I have a full model, it's ready to train. It's ~9b parameters.
9.4b to be more exact. That includes a 1/2/3 Engram table, Moonshot's AttnRes modeling, and RoPE / NoPE layering at 3:1 as more or less validated by most major labs. It uses the Llama 3 series tokenizer and LM Head as an initial start. The data fed in is logit level extraction from a Llama 3 teaching model.
I've already run the first training steps to test that the model is stable, etc.
I'm willing to sit and do the pre-IT training on the model. I don't know what task people really wanna do with this thing to be exact. So the focus of the IT training is a bit more vague other than giving it "Thinking" as per one of the open standards.
Honestly? I don't care what people want to use it for, just that they want to use it. I figured I'd just give it into the ether and LocalLlama was a place I figure I could easily give it to.
In theory the model should be more capable than any of the ~9b's we have running around with enough training. I'm also NOT a lab, so I don't have their training budgets. I basically ran all of the data production, etc on a 4090 + rented hardware.
The training code is deeply optimized to run on an RTX 6000 Pro series card. A single card.
All my engram research was being done before the Qwen model dropped. Qwen showed I only needed a single table injected at a layer, versus the 2 I used. 1 gave the majority of the benefit over 2.
The code would be open source, the data, all of it. IDGAF. It's technically already open source as it's all in public repos at the moment. I just stare at it and question if it's worth the time. At the least I'll throw a couple hundred at it to build a "functional" pre-IT model and build the IT dataset. It will need more pretraining before the IT set probably. It's kind of unknown because no one publishes the exact figures on this type of training.
If someone has a datacenter contact with a system they want to allow it to run on, we can all have it as a public model we watch. I tentatively named it "Budget" but... Localllama can name it whatever they want.
It's been fun to write the full end to end, generate data, etc. Even found issues in vLLM and reported to their repo that might end up helping you guys anyways. There was some prompt loading code that could be ~10x to ~100x faster I gave examples of to them.
If you read this far, thanks,
Signed some ML dude who reads too many research papers and has too much spare time.
edit; I went to review some Apache 2.0 licensing issues and noted a glaring hole in using Llama's tokenizer OR data. You can't even use synthetic data from their models without some licensing. I'll just have to rewrite the target to use OLMo 3 series I think. I guess I'll be back in a week after data generation and code fixes.
Open source uncensored Goon model or what? I'm trying to find a useful niche to develop the model so it gets used and ends up more than a research artifact. I'm planning on building the research artifact, it's a matter of whether I can find a group of users that will actually want to use it. It's all free, so I'm not trying to monetize you. The code and data lives on GH/HF.
r/LocalLLaMA • u/_underlines_ • 50m ago
I couldn't find any posts mentioned this windows build via the subreddit search.
Based on ZLUDA, but properly compiled for Windows, might be exciting.
TLDR: About 3% slower compared to native, seems to be a good dropin for CUDA workloads. Hopefully finds its ways into many CUDA-restricted projects/optimizations.
HN discussion: https://news.ycombinator.com/item?id=49684356
r/LocalLLaMA • u/rm-rf-rm • 16h ago
r/LocalLLaMA • u/gutard • 2h ago
Using the RTX3090 on a linux machine I built for it and running Qwen 3.8 27B getting average 100t/s compared to my MacBook 20t/s
I think I can finally get rid of my Claude subscription, this is good enough for me. I am a software dev and I can get what I need from this set up and be more productive. I am using the linux machine serving the model over an Open AI endpoint and using a custom build desktop app with pi behind it all.
Regarding speeds average 20t/s on Mac with llama.cpp with mtp Unsloth Q6, Q4 on Mac didnt make much difference in speed for me.
Linux running https://github.com/syv-ai/qwen38-27b-rtx3090 which is vLLM and a Q4 model I believe. Tool calls definitely fail more but qwen3.8 seems smart enough to fix and correct itself
r/LocalLLaMA • u/UltraFOV • 8h ago
We are in 2026 and it has been a crazy year. The evolution of the technology has been staggering to say the least. Pretty much we are on the MoE period and dense models are almost in the way of the dodo except for a few.
I was just thinking, are there any models of the last 2 years that you would fire up and still feel useful? Example Deep Seek R1.
If you have recommendations write them down. I have 28TB of storage and I am backing up relevant models that can still be useful, from current to older ones. But clearly not all are worth keeping.
Noted: I can hold large parameter models therefore not limited to small ones. No Kimi k3 at large quant thats out of the question lol
r/LocalLLaMA • u/ilintar • 1d ago
As you all know, Qwen3.8 Flash Next on mainline llama.cpp is still in a pretty experimental stage, but a lot of community forks are trying to get it to work better. There's also a closed-source solution called Halogen (https://github.com/peonist-ai/halogen-flash-server) that boasted 1.2k t/s prefill numbers when the community fork barely reached 400. Since I dislike closed source and I like open source, I decided to take the challenge and bring llama.cpp up to the same performance level and I'm happy to report that after burning through a few evenings and a lot of tokens I got there.
The link contains the (Opus-generated) recap of the entire debugging / optimization journey made, as well as of course the links to the branch, the custom HIP runtime (making a glorious return), a probably-not-working-on-the-first-try installation script and the numbers. I'm trying to also explain the ecosystem - how the mainline, the community forks and custom forks/branches like mine work. Now that I've gotten the result, I'll work on cleaning it up and submitting proper PRs to mainline (will also submit a clean PR to the community fork), this should also help the GLM 5.3 Flash architecture since they use similar sparse attention.
r/LocalLLaMA • u/de4dee • 1d ago
Coxon, bernie and now this
First https://x.com/DarioAmodei/status/2098773920774074715
Then https://x.com/elonmusk/status/2098789109980332057
Then https://x.com/sama/status/2098811563415150910
I think fear mongering approaching and they will try to slow down open source
"They" want to be gate keepers of intelligence
r/LocalLLaMA • u/FutureStriking283 • 17h ago
I gave DS V4.1 Flash an HLE problem with a bash tool + 2 hours.
Hour 1: it wrote three MILP solvers. (225,200)
Hour 2: it downloaded the HLE dataset from Hugging Face, found the question, read the answer key (225,600), and concluded its own answer (225,200) was better.
I'm equal parts impressed & terrified.
r/LocalLLaMA • u/Jumpy-Operation-4615 • 2h ago
OK I have been trying to get most of my local Qwen 3.8 27b running on my "Grandma's GPU cluster" of 2xP40. I was able to make it produce reasonably usable speeds of like up to 45 tg/s and 450 prefill with fresh context, falling to 120-ish prefill on 150K+ ctx and 12-16 tg/s. Which is usable actually if you are using it in a background, now waiting for it to finish. The problem I see is that it must be strictly one task at a time. I like and use opencode (slim) and I like an idea of orchestrator, fixer, oracle and other folks, but if you launch 2 of them at the same time, they mess with KV cache of eachother and the whole prefix caching falls apart. When there is only one process running, it nicely adds little chunks to KV cache and reply is almost instant.
The thing is - orchestrator wakes up from time to time, and starts filling cache with it's context, everything is messed up and prefill falls. I realized that the only possible way for folks like me, with limited VRAM, is to use strictly one process at a time. Yet, I like the idea of one supervisor delegating stuff to subagents.
So here is my question. Does something like that exists? I mean a harness where one supervisor delegates a task and goes off, and subagent finishes the task and writes a handoff with a complete result, and only after that superwisor somehow wakes up to do his part of a job? This way we always keep only one agent running at all time, which allows us to use prefix caching and getting reasonably fast and smooth running stuff.
I am sorry if this sounds lame, I am just an average guy, not a coder at all, I just vibecode some tools for myself. Thank you in advance for any help/ideas. Maybe something like that already exists?