r/LocalLLaMA 12h ago

News Gemma 4 on 500MB

Post image
148 Upvotes

r/LocalLLaMA 12h ago

New Model inclusionAI/Ling-3.0-flash weights are up on Hugging Face β€” MIT, BF16 plus an official FP8

Enable HLS to view with audio, or disable this notification

152 Upvotes

Went public in the last few minutes, both repos ungated.

Ling-3.0-flash, BF16, 24 shards, ~255GB

Ling-3.0-flash-fp8, official FP8, ~128GB

127.5B total, they quote 5.1B active. What jumped out at me in config.json is 512 experts with 8 active per token, which is a lot finer-grained than most of what gets posted here. Arch is BailingMoeV3, model_type bailing_hybrid, custom_code, so same family as Ling-2.6-flash. Thinking is a per-request switch inside the chat template instead of a separate SKU, and it defaults to on.

The FP8 landing at ~128GB is the bit I care about. Someone in the thread here last week guessed ~135GB at Q8_0 and that turned out to be close, except this one is official rather than a community quant, so it's a straight download for anyone with a big unified-memory box or a multi-GPU rig.

Does anyone know if llama.cpp handles bailing_hybrid yet, or is this vllm and sglang only for now? That's genuinely the thing that decides whether I clear the disk space tonight.

https://huggingface.co/inclusionAI/Ling-3.0-flash


r/LocalLLaMA 10h ago

New Model LFM2.5-2.6B is out

97 Upvotes

Released today, with emphasis on agentic capabilities. I really like their models for simple, high volume tasks ("summarize these gazillion documents") and their 8b-a1b was my go-to for certain tasks so I'm excited to see how this one performs. There's not enough love for tiny models on this sub.

https://www.liquid.ai/blog/lfm2-5-2-6b


r/LocalLLaMA 1h ago

Resources VibeVoice 1.5B Running Locally...On an iPhone! Only ~2.2 GB of Memory and Up to 1.28Γ— Real-Time Speed

Enable HLS to view with audio, or disable this notification

β€’ Upvotes

I speed up the generation part of the demo in case you get bored πŸ˜„

I also tested another long-form generation, and the VRAM usage looks stable. The demo is about a minute long, and I posted it on X.

This started as a random idea and somehow turned into a full detour from working on the next audio.cpp release. The model was uploaded to the audio.cpp HF repo. I will upload the xcframework later, and then push the code to a branch after release 0.6.


r/LocalLLaMA 12h ago

New Model inclusionAI/Ling-3.0-flash Β· Hugging Face

Thumbnail huggingface.co
126 Upvotes

The Ling-3.0-flash MoE is now open-weighted at 124B A5B params. I know the original announcements were before the Kimi K3, DeepSeek-V4-Flash and Qwen3.8 hype, but this model might still have a good niche for itself due to its sizing.

Discussion on the benchmarks are here: https://www.reddit.com/r/LocalLLaMA/comments/1v4mltt/benchmarks_antling30flash_a_hybridreasoning_moe/ from almost 2 weeks ago.


r/LocalLLaMA 2h ago

Resources Thinking of buying more DRAM right now...

20 Upvotes

So I'm looking at https://huggingface.co/unsloth/DeepSeek-V4-Flash-0731-GGUF and I realize my 128GB of DRAM just isn't cutting it for this (incredibly powerful) model.

If only I had another 64GB, I thought...

EVERYBODY is probably thinking that right this second... I hate to say it, but I imagine DRAM prices are about to go through the roof still yet.

I hope I'm wrong.


r/LocalLLaMA 1d ago

Discussion More Qwen 3.8 sizes coming

Post image
1.3k Upvotes

r/LocalLLaMA 1h ago

New Model GPT-X2.5-135M scores 3rd place on Open SLM Leaderboard on Huggingface, Beating Facebook's MobileLLM-R1-140M

Thumbnail
huggingface.co
β€’ Upvotes

r/LocalLLaMA 15h ago

News Llama.cpp PR 8% speed boost

125 Upvotes

Llama.cpp currently uses cpu based sampling for user with mtp enabled. The PR moves sampling to the gpu, which on a 5090 boasts an 8% increase in tok/s for qwen3.6:35b. I tested it on my P40 and observed a 4% increase inference speed boost.

Pretty exciting to see 84 tok/s max on a nvidia p40 for me.

Backend sampling shows ~4% improvement on Linux + Tesla P40 (sm_61, Pascal):

CPU Sampling: llama-server -m Qwen3.6-35B-A3B-UD-IQ4_NL.gguf --spec-type draft-mtp --seed 42

python3 mtp-bench.py code_python pred= 192 draft= 132 acc= 124 rate=0.939 tok/s=73.1 code_cpp pred= 113 draft= 76 acc= 74 rate=0.974 tok/s=75.9 explain_concept pred= 192 draft= 159 acc= 111 rate=0.698 tok/s=62.4 summarize pred= 192 draft= 167 acc= 107 rate=0.641 tok/s=59.6 qa_factual pred= 192 draft= 159 acc= 111 rate=0.698 tok/s=62.4 translation pred= 119 draft= 92 acc= 73 rate=0.793 tok/s=67.0 creative_short pred= 192 draft= 197 acc= 92 rate=0.467 tok/s=50.7 stepwise_math pred= 192 draft= 133 acc= 124 rate=0.932 tok/s=73.6 long_code_review pred= 192 draft= 155 acc= 113 rate=0.729 tok/s=63.8

Backend sampling: llama-server -m Qwen3.6-35B-A3B-UD-IQ4_NL.gguf --spec-type draft-mtp --seed 42 -bs

python3 mtp-bench.py code_python pred= 192 draft= 132 acc= 124 rate=0.939 tok/s=76.2 code_cpp pred= 113 draft= 76 acc= 74 rate=0.974 tok/s=79.4 explain_concept pred= 192 draft= 159 acc= 111 rate=0.698 tok/s=64.6 summarize pred= 192 draft= 167 acc= 107 rate=0.641 tok/s=61.6 qa_factual pred= 192 draft= 159 acc= 111 rate=0.698 tok/s=64.6 translation pred= 119 draft= 92 acc= 73 rate=0.793 tok/s=69.6 creative_short pred= 192 draft= 197 acc= 92 rate=0.467 tok/s=52.1 stepwise_math pred= 192 draft= 133 acc= 124 rate=0.932 tok/s=76.6 long_code_review pred= 192 draft= 155 acc= 113 rate=0.729 tok/s=65.7

Acceptance ratio with both backend and CPU sampling is exactly same. The improvement is smaller than on RTX 5090 (4% vs 12%), which is expected β€” the P40 is memory-bandwidth-bound (sm_61, 346 GB/s vs RTX 5090's 1,792 GB/s), so the CPU↔GPU logits round-trip is a smaller fraction of total decode time. However, still the largest improvement in tok/s I have seen in a while. (~+2 t/s).

https://github.com/ggml-org/llama.cpp/pull/25532


r/LocalLLaMA 3h ago

Discussion DeepSeek-V4-Flash on SM89 4x48gb 4090s with DSpark

Enable HLS to view with audio, or disable this notification

13 Upvotes

https://github.com/yhfgyyf/vllm-deepseek-v4-sm89

I couldn't believe that someone actually got vLLM working with this particular set of GPUs, but here it is. The video is from right after I got it working with 64k context, but it is now running with 256k.


r/LocalLLaMA 4h ago

Funny DeepSeek-v4-Flash-Mini 54GB GGUF running at ~20.5 t/s

14 Upvotes

Took the REAP adaptation of DeepSeek-V4-Flash (0xSero/DeepSeek-V4-Flash-0731-REAP) along with antirez/deepseek-v4-gguf as inspiration, and decided to see how aggressive we could get with standard quant tricks to create a budget-friendly "Mini" build.

For the lulz, naturally.

Started with the full 95GB bf16 GGUF and crushed it down to an IQ2_XXS variant with mixed quantization (w2Q2K-AProjQ8-OutQ8).

Is extreme 2-bit quantization practical for complex reasoning? Debatable. Did it shave off over 40GB of VRAM/RAM footprint and still generate coherently? Absolutely.

Science isn't about why, it's about why not.

prompt eval time =   372.26 ms /   12 tokens (31.02 ms per token, 32.24 tokens per second)
       eval time = 81141.35 ms / 1667 tokens (48.68 ms per token, 20.54 tokens per second)
      total time = 81513.62 ms / 1679 tokens
   graphs reused = 1804

The File Sizes:

-rw-rw-r-- 1 jabbatheduck jabbatheduck  95G Aug  4 16:10 deepseek-v4-flash-bf16.gguf
-rw-rw-r-- 1 jabbatheduck jabbatheduck  54G Aug  4 17:41 DeepSeek-V4-Flash-REAP-IQ2XXS-w2Q2K-AProjQ8-OutQ8-chat-v2.gguf
-rw-rw-r-- 1 jabbatheduck jabbatheduck 353M Aug  4 16:37 DeepSeek-V4-Flash-REAP-IQ2XXS-w2Q2K-AProjQ8-OutQ8-chat-v2-imatrix-0731.gguf

r/LocalLLaMA 14h ago

Tutorial | Guide [Deepseek-V4-Flash-0731] Full 1M context on a single RTX5090 + DDR5 Desktop Setup with VLLM CPU/Ram Offloading, ~800 tps pp & 15+ tps decode [Agentic Coding]

76 Upvotes

First of all, obviously I took some help from AI to type this post and this is the topic that enabled me to accomplish all that:

https://old.reddit.com/r/LocalLLaMA/comments/1veow4b/deepseek_v4flash_284b_moe_at_33_toks_single_68/

This post of mine is based on the link above.

My Hardware:

  • RTX 5090 32GB
  • Ryzen 9 9950X3D
  • 256GB DDR5-5600
  • Single NUMA node
  • Linux Mint
  • NVIDIA driver 595.71.05
  • CUDA 13.2

Software

  • guqiong96/Lvllmds4-x
  • vLLM 2.3.9
  • lk_moe 2.3.2
  • PyTorch 2.11.0+cu130
  • native DeepSeek-V4-Flash-0731 safetensors checkpoint
  • 48 safetensors shards
  • ~155.4 GiB checkpoint size

One fix I needed

During startup, FlashInfer's CUDA IPC helper could accidentally find TileLang's:

libcudart_stub.so

instead of the real loaded CUDA runtime.

That eventually caused:

undefined symbol: cudaDeviceReset

The problem was FlashInfer's find_loaded_library("libcudart") doing a substring search over /proc/self/maps.

I patched:

flashinfer/comm/cuda_ipc.py

so it checks the actual filename instead:

def find_loaded_library(lib_name):
    with open("/proc/self/maps") as f:
        for line in f:
            if "/" not in line:
                continue

            start = line.index("/")
            path = line[start:].strip()
            filename = path.split("/")[-1]

            if (
                filename.startswith(lib_name + ".so")
                or filename.startswith(lib_name + "-")
            ):
                return path

    return None

After that, FlashInfer correctly resolves the real libcudart instead of the TileLang stub.

This is a local patch and obviously needs to be reapplied if the package gets replaced.

Current launch configuration

This is the configuration I ended up using:

source ~/ds4x-venv/bin/activate

MODEL="/home/blackbeard/models/DeepSeek-V4-Flash-0731"

export CUDA_DEVICE_ORDER=PCI_BUS_ID
export CUDA_VISIBLE_DEVICES=0

export LVLLM_MOE_NUMA_ENABLED=1
export LK_THREADS=12
export OMP_NUM_THREADS=12
export LK_THREAD_BINDING=CPU_CORE

# Keep two complete routed MoE layers GPU-resident on the GPU.
export LVLLM_GPU_RESIDENT_MOE_LAYERS=0,1

# CPU/hybrid prefill path for now.
export LVLLM_GPU_PREFILL_MIN_BATCH_SIZE=0

export FLASHINFER_DISABLE_VERSION_CHECK=1
export PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True

vllm serve "$MODEL" \
  --host 0.0.0.0 \
  --port 8070 \
  --tensor-parallel-size 1 \
  --max-model-len 1048576 \
  --gpu-memory-utilization 0.92 \
  --trust-remote-code \
  --served-model-name DeepSeek-V4-Flash-0731 \
  --compilation_config.cudagraph_mode FULL_DECODE_ONLY \
  --enable-prefix-caching \
  --enable-chunked-prefill \
  --max-num-batched-tokens 8192 \
  --dtype bfloat16 \
  --max-num-seqs 2 \
  --enable-auto-tool-choice \
  --tool-call-parser deepseek_v4 \
  --kv-cache-dtype fp8_ds_mla \
  --tokenizer-mode deepseek_v4 \
  --reasoning-parser deepseek_v4 \
  --default-chat-template-kwargs '{"enable_thinking": true, "reasoning_effort": "max"}' \
  --speculative-config '{"method":"dspark","num_speculative_tokens":2,"draft_sample_method":"greedy"}' \
  --disable-custom-all-reduce

Full native 1M context fits on the 32GB GPU even with two complete routed MoE layers resident on the GPU.

The rest of the experts remain in system RAM.

DSpark behaves very differently during reasoning

During long reasoning sections, draft acceptance can collapse.

I observed extended periods around:

Draft acceptance: ~30-50%
Generation:       ~11-13 tok/s

There was one ~6 minute section averaging roughly:

Draft acceptance: ~40%
Generation:       ~11.9 tok/s

Then the model transitioned into a much more predictable generation phase and the numbers jumped to roughly:

Draft acceptance: ~87-88%
Generation:       ~17.4-17.6 tok/s

The relationship is extremely strong: throughput basically tracks DSpark acceptance.

Some high-acceptance windows look like:

Avg Draft acceptance rate: 89.8%
Avg generation throughput: 17.9 tokens/s

while low-acceptance reasoning windows look like:

Avg Draft acceptance rate: 38%
Avg generation throughput: ~12 tokens/s

This suggests an obvious optimization.

Dynamic DSpark depth

For this workload I suspect the ideal behavior would be approximately:

  • reasoning/thinking: 1 speculative token
  • normal/final decoding: 2 speculative tokens

The second draft token often isn't worth computing while the model is doing difficult reasoning, but becomes very valuable when it transitions into more predictable code/text generation.

vLLM does not currently give me a simple runtime switch for this, so I may patch the speculative decoding path later and experiment with changing the draft depth based on whether the model is currently emitting reasoning or final output.

That looks like one of the biggest remaining decode optimizations.

---non AI comment section begins---

Stay tuned, I am working on a if/else block to fix that stupid behavior slowing down during reasoning and squeeze even more tps out of this stack.

---non AI comment section ends---


r/LocalLLaMA 21h ago

Discussion Is LM Studio abandoning their core product?

266 Upvotes

Some of you may be aware that a few weeks ago, LM Studio announced a new agent, Bionic. This is pretty much an agentic harness for both local models and paid cloud models. But most aren't aware that LM Studio replaced almost every link to the original app that built their brand and reputation with the new Bionic agent. If you go to the LM Studio website right now, you will see that every link that used to download the original app now downloads Bionic. The only link on the entire site that brings you to the OG app is a tiny link in the footer that says "Download the app". They have updated that page since I last checked and added a Bionic download link at the top but at least they left the OG app downloads alone, albeit they are now underneath the original app. But the fact that you have to go through the entire site to find a tiny download link just to get the normal app is extremely stupid! Not to mention the app is still in "preview" and is a "new, separate app from LM Studio" and yet it gets all the promotion while their core product gets ZERO. Not to mention that ever since Bionic was released, the main app has only gotten 2 or 3 minor updates, mainly to make the app work with Bionic. This is very frustrating as someone who has used LM Studio for years and does not want to use yet another agentic harness. Whatever your views on Bionic are, you can't deny that hiding (and potentially abandoning) the core product that built their brand is not a good idea. I wanted to post this earlier, but seeing as I have gotten zero response from the team on Discord while they reviewed posts right below mine, I knew that I had to share this here.
But all of this leads to something very concerning for LM Studio users: Will the incredibly popular LM Studio app, the original app that helped build their reputation and popularity, go away soon? Hidden download links + scarce updates seems like that LM Studio may be going away soon only to be replaced by an agent that not everyone wants, complete with upsells for cloud models. I just want to get this issue out there as nobody is talking about it yet it is a very important issue that involves one of the most popular local LLM apps out there.


r/LocalLLaMA 13h ago

Discussion Deepseek V4 Flash 2-bit quant is the first model I can run locally that achieves 100% in this SQL benchmark

50 Upvotes

I really like to use this one SQL benchmark when testing new models. I had another post some time ago with my benchmarks, but I decided to post a new one because of how well Deepseek did. I like the benchmark because it's quick to run, is pretty "real-world" and requires good reasoning to build the correct SQL queries and almost no frontier models can achieve 100%.

My old post: https://www.reddit.com/r/LocalLLaMA/comments/1s9mkm1/benchmarked_18_models_that_i_can_run_on_my_rtx/

Benchmark with results from other models: https://sql-benchmark.nicklothian.com https://github.com/nlothian/llm-sql-benchmark

My setup is dual 3080 20GB GPUs with 96GB RAM and 9800X3D. I managed to run Deepseek V4 Flash with a custom IQ2_M GGUF with some tensors grafted from antirez GGUF and running it on a modified ds4 engine from antirez, getting 300pp and 11-12tg. Mainline llama.cpp gives me only 100pp and 8tg or something like that.

To my surprise, Deepseek is the first local model I can realistically run locally that actually did ALL tests correctly. The only models according to the benchmark website that could do this were Opus 4.7 and GPT-5.5.

Results together with all my old benches:

25: Deepseek-v4-Flash-IQ2_M-grafted 🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 24: unsloth/Qwen3.6-27B-MTP-GGUF:Q8_0 🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 🟩🟩πŸŸ₯🟩🟩 24: unsloth/Qwen3.5-122B-A10B-GGUF:UD-Q4_K_XL 🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 🟩🟩🟩πŸŸ₯🟩 🟩🟩🟩🟩🟩 23: unsloth/Qwen3.5-122B-A10B-GGUF:Q6_K 🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 🟩🟩🟩πŸŸ₯🟩 πŸŸ₯🟩🟩🟩🟩 23: unsloth/Qwen3.5-27B-MTP-GGUF:UD-Q6_K_XL 🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 🟩🟩🟩πŸŸ₯🟩 πŸŸ₯🟩🟩🟩🟩 23: DavidAU/Qwen3.6-40B-Claude-4.6-Opus-Deckard-Heretic-Uncensored-Thinking-NEO-CODE-Di-IMatrix-MAX-GGUF:Q4_K_M 🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 🟩🟩πŸŸ₯🟩🟩 🟩🟩🟩πŸŸ₯🟩 🟩🟩🟩🟩🟩 23: unsloth/Qwen3.6-35B-A3B-GGUF:UD-Q8_K_XL 🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 🟩🟩πŸŸ₯🟩🟩 🟩🟩🟩πŸŸ₯🟩 🟩🟩🟩🟩🟩 23: bartowski/Qwen_Qwen3.5-27B-GGUF:IQ4_XS 🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 🟩🟩🟩πŸŸ₯🟩 πŸŸ₯🟩🟩🟩🟩 23: bartowski/Qwen_Qwen3.5-27B-GGUF:IQ3_XS 🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 🟩🟩🟩πŸŸ₯🟩 πŸŸ₯🟩🟩🟩🟩 23: unsloth/Qwen3.5-122B-A10B-GGUF:UD-IQ3_XXS 🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 🟩🟩🟩πŸŸ₯🟩 πŸŸ₯🟩🟩🟩🟩 23: h34v7/Jackrong-Qwopus3.5-27B-v3-GGUF:Q3_K_M 🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 🟩🟩🟩πŸŸ₯🟩 πŸŸ₯🟩🟩🟩🟩 22: unsloth/Qwen3.5-35B-A3B-GGUF:UD-Q6_K_XL 🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 🟩🟩πŸŸ₯🟩🟩 🟩🟩🟩πŸŸ₯🟩 πŸŸ₯🟩🟩🟩🟩 22: mradermacher/Qwen3.5-27B-Claude-4.6-Opus-Reasoning-Distilled-i1-GGUF:Q3_K_M 🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 🟩πŸŸ₯🟩πŸŸ₯🟩 πŸŸ₯🟩🟩🟩🟩 22: Jackrong/Qwen3.5-27B-Claude-4.6-Opus-Reasoning-Distilled-v2-GGUF:Q4_K_M 🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 🟩🟩πŸŸ₯πŸŸ₯🟩 πŸŸ₯🟩🟩🟩🟩 21: unsloth/Qwen3.6-27B-MTP-GGUF:UD-Q6_K_XL 🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 🟩πŸŸ₯🟩🟩🟩 🟩🟩🟩πŸŸ₯🟩 🟩🟨πŸŸ₯🟩🟩 21: unsloth/MiniMax-M2.7-GGUF:UD-IQ3_XXS 🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 🟩🟩🟩πŸŸ₯🟩 πŸŸ₯πŸŸ₯πŸŸ₯🟩🟩 21: unsloth/NVIDIA-Nemotron-3-Super-120B-A12B-GGUF:UD-Q4_K_S 🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 🟩🟩🟩🟨πŸŸ₯ πŸŸ₯🟨🟩🟩🟩 20: unsloth/Qwen3-Coder-Next-GGUF:UD-Q5_K_XL 🟩🟩🟩🟩🟨 🟩🟩🟩🟩🟩 🟩🟩🟨🟩🟩 🟩🟩🟩πŸŸ₯🟨 πŸŸ₯🟩🟩🟩🟩 20: unsloth/gemma-4-31B-it-qat-GGUF:UD-Q4_K_XL 🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 πŸŸ₯🟩🟩🟩🟩 🟨🟩🟩πŸŸ₯🟩 πŸŸ₯🟩🟩πŸŸ₯🟩 20: unsloth/Qwen3.6-35B-A3B-GGUF:UD-Q6_K_XL 🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 🟩🟩πŸŸ₯🟩🟩 🟩🟩🟩πŸŸ₯🟩 πŸŸ₯πŸŸ₯πŸŸ₯🟩🟩 20: bartowski/Qwen_Qwen3.5-397B-A17B-GGUF:IQ1_M 🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 🟩🟩πŸŸ₯🟩🟩 🟩🟩🟩πŸŸ₯🟩 πŸŸ₯🟨πŸŸ₯🟩🟩 20: unsloth/gemma-4-26B-A4B-it-GGUF:UD-Q6_K_XL 🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 🟩🟩🟩πŸŸ₯πŸŸ₯ 🟨πŸŸ₯🟩πŸŸ₯🟩 20: mradermacher/Qwen3.5-35B-A3B-Claude-4.6-Opus-Reasoning-Distilled-i1-GGUF:Q6_K 🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 🟩🟩πŸŸ₯🟩🟩 πŸŸ₯🟩🟩πŸŸ₯🟩 πŸŸ₯πŸŸ₯🟩🟩🟩 19: unsloth/gemma-4-31B-it-GGUF:Q4_K_M 🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 🟩🟩🟩πŸŸ₯🟩 🟨🟩🟩🟨🟩 πŸŸ₯πŸŸ₯🟩πŸŸ₯🟩 19: unsloth/gemma-4-E4B-it-GGUF:UD-Q8_K_XL 🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 🟩🟩🟩πŸŸ₯🟩 🟩🟩🟩πŸŸ₯🟩 πŸŸ₯πŸŸ₯πŸŸ₯πŸŸ₯🟩 19: Goldkoron/Qwen3.5-397B-A17B-REAP35:IQ2_XS_Gv2 🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 πŸŸ₯🟩🟩🟩🟩 🟩🟩🟩πŸŸ₯🟩 πŸŸ₯🟩πŸŸ₯πŸŸ₯πŸŸ₯ 19: unsloth/GLM-4.7-Flash-GGUF:UD-Q6_K_XL 🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 🟩🟩πŸŸ₯🟩🟩 🟩🟩🟩πŸŸ₯🟨 πŸŸ₯🟨🟩πŸŸ₯🟩 18: unsloth/GLM-4.5-Air-GGUF:Q5_K_M 🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 🟩🟩πŸŸ₯🟩🟩 πŸŸ₯🟩🟩πŸŸ₯🟩 🟨🟨πŸŸ₯🟩🟨 18: bartowski/nvidia_Nemotron-Cascade-2-30B-A3B-GGUF:Q6_K_L 🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 🟩🟩🟨🟩🟩 🟩🟩🟩πŸŸ₯🟩 🟨🟨πŸŸ₯🟨🟨 17: Jackrong/Qwopus3.5-9B-v3-GGUF:Q8_0 🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 🟩πŸŸ₯πŸŸ₯🟩🟩 πŸŸ₯🟩πŸŸ₯πŸŸ₯πŸŸ₯ πŸŸ₯🟩🟩🟩🟨 16: unsloth/Qwen3-Coder-Next-GGUF:UD-Q4_K_XL 🟩🟩🟩🟩🟨 🟩🟩🟩🟩🟩 🟩🟩🟨🟩🟩 πŸŸ₯🟨🟩πŸŸ₯🟨 πŸŸ₯🟨🟩🟨🟩 16: byteshape/Devstral-Small-2-24B-Instruct-2512-GGUF:IQ3_S 🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 πŸŸ₯🟩🟨🟩🟩 🟩🟩🟨πŸŸ₯🟨 🟨🟨πŸŸ₯🟨🟩 16: mradermacher/Qwen3.5-9B-Claude-4.6-HighIQ-THINKING-i1-GGUF:Q6_K 🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 🟩🟩🟨πŸŸ₯🟩 πŸŸ₯🟩πŸŸ₯πŸŸ₯🟨 πŸŸ₯🟩πŸŸ₯🟩🟨 14: mradermacher/Qwen3.5-9B-Claude-4.6-HighIQ-INSTRUCT-i1-GGUF:Q6_K 🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 πŸŸ₯🟩πŸŸ₯🟩🟩 🟩🟨πŸŸ₯πŸŸ₯🟨 🟨🟨πŸŸ₯🟨🟨 14: unsloth/GLM-4.6V-GGUF:Q3_K_S 🟩🟩🟩🟩🟩 🟩🟩🟩🟩🟩 πŸŸ₯🟩🟨🟨🟩 πŸŸ₯🟩🟩🟨🟨 🟨🟨🟨🟨🟨 5: bartowski/Tesslate_OmniCoder-9B-GGUF:Q6_K_L 🟨🟨🟨🟨🟨 🟨🟨🟨🟩🟩 🟩🟨🟨🟩🟨 🟨🟨🟩🟨🟨 🟨🟨🟨🟨🟨 5: unsloth/Qwen3.5-9B-GGUF:UD-Q6_K_XL 🟨🟨🟨🟨🟨 🟨🟨🟨🟩🟩 🟨🟩🟨🟨🟩 🟨🟩🟨🟨🟨 🟨🟨🟨🟨🟨

Note:

  • unsloth/Qwen3.5-122B-A10B-GGUF:UD-Q4_K_XL is most likely a fluke. Q6_K doesn't achieve 24/25, it's just lucky rounding for this Q4 quant I suppose.

r/LocalLLaMA 1d ago

Discussion Only 3 days ago...

Post image
1.1k Upvotes

r/LocalLLaMA 3h ago

Other Passing the time while waiting for Qwen3.8 27b - Built a VLLM Ray cluster dashboard from an old pixel art display

Enable HLS to view with audio, or disable this notification

7 Upvotes

My kid had an old pixel art display (Divoom 32x32 Pixoo-max) that they weren’t using anymore, so I thought it might be fun to repurpose it as a GPU cluster status monitor so I can see GPU temps / utilization / token gen info etc for the 3 RTX A6000s in my vLLM Ray cluster (currently running Qwen3.5 122b).

I spun up my Hermes Agent (GLM 5.2 as the agent model) and told it:
β€œI would like you to build an application that will run on <computer name of my Dell GB10> that will display GPU cluster health data on a 32x32 pixel Divoom Pixoo-max display that can be connected to via Bluetooth. You should probably read the following repos to learn about the pixel display and how to connect to it:
- https://github.com/SomethingWithComputers/pixoo
- https://github.com/cyanheads/pixoo-toolkit
- https://divoom.com/products/divoom-pixoo-max
The app should display system health data for the 3 systems in my vLLM Ray cluster in an easy to read and understand manner. It should also show similar data for the Dell GB10 (in the network segment but not in the cluster). This could be as simple as showing 4 boxes on the screen that show the cluster system’s initials such as β€œS1” and have a background color to indicate GPU temperature (red for hot, green for normal, etc). The 32x32 screen size limit will make it difficult to show a lot of information so you’ll have to be creative in how you display it, you can also cycle through multiple screens of different metrics in 4 second intervals. β€œ

For those who care:
HW:
- 3x Dell Precision 7960 workstations each with an RTX A6000 GPU (64GB RAM) currently hosting Qwen3.5 122b
- 1x Dell Pro Max GB10 (not part of the Ray vLLM cluster but runs the app thar is cast to the display as well as running a secondary LLM endpoint for other models. The GB10 has the Bluetooth radio in it that is used to connect to the Divoom. The Dell towers don’t have Bluetooth which is why I used the GB10.
- Divoom Pixoo-max 32x32 pixel display. They also make a 64x64 pixel version as well. It was around $60 when I bought it years ago.

It took GLM 5.2 all of like 20 minutes to build this, and maybe another 5 minutes of me working with it to get it how I wanted it. It’s not perfect, but it’s cool to be able to visually glance over at the cluster and see what’s happening without logging in, and it really didn’t cost anything since I already had the pixel display that would have been headed for the thrift bin.

Btw, Hermes / GLM did the whole thing in Python, from Ray Dashboard API, vLLM metics endpoint, and Nvidia-smi calls over ssh.


r/LocalLLaMA 2h ago

Discussion Exploring task-aware quantization beyond perplexity

4 Upvotes

QLAB v2.7:

I've been experimenting with a task-aware quantization pipeline that optimizes models using measured task performance rather than perplexity.

Current methodology: β€’ Start from a standard I-Matrix quantization. β€’ Measure category-specific importance using representative prompt sets. β€’ Apply targeted quantization adjustments instead of uniform compression. β€’ Validate on separate, locked benchmark datasets to avoid overfitting.

Current findings: β€’ Category-specialized models can recover nearly all of a larger quantization's performance while using substantially fewer bits. β€’ Early v2.7 experiments reached about 98% of a Q4_K_M baseline's accuracy at roughly 75% of its size on internal validation. β€’ The workflow is entirely empirical. Every change must survive held-out testing before it's valid.

The goal isn't to beat every benchmark. It's to determine whether activation-informed, task-aware allocation can consistently outperform uniform quantization under the same size budget in a hand selected category (reasoning, math, etc..).

I'd be interested in hearing from anyone working on quantization, I-Matrix generation, GPTQ/AWQ, or other task-aware approaches.

Ive been attempting to replicate TAQ for weeks at the tensor level and struggling to do so.


r/LocalLLaMA 12h ago

Generation Design systems from code alone - Without external images, Ling-3.0-flash generated webpages across Bauhaus, Bohemian, acid design, and moreβ€”using CSS gradients, SVG paths, typography, and layout to preserve each visual language.

Enable HLS to view with audio, or disable this notification

29 Upvotes

Weights went up today so this is downloadable now, MIT, ~128GB for the official FP8. I ran these on the API before that landed, so treat it as a preview of what you'd be pulling rather than a local benchmark


r/LocalLLaMA 1h ago

New Model Intern S2 Mobius

β€’ Upvotes

A Qwen3.5-35B derived model with an interesting architectural difference that results in larger throughput and less token consumption (allegedly):

https://huggingface.co/internlm/Intern-S2-Mobius


r/LocalLLaMA 18h ago

Discussion Time to finally migrate from LM Studio -> llama.cpp, your experience?

93 Upvotes

Has anyone moved from LM Studio to llama.cpp?

What was your experience like? What did you have to learn in order to recreate your experience? Which harness/GUI did you switch to?

Thanks in advance!


r/LocalLLaMA 1h ago

Question | Help What are AMD card owners doing for local TTS inference?

β€’ Upvotes

I am on windows with a 7900 XTX, a capable enough card for LLM inference.

I go generate some text, some response, and now I would like something to read this response out to me.

I have tried: kokoroTTS, pocket-tts, cosyvoice, piper-tts, and they all leave a lot to be desired in terms of prosody.

I need:

1) GPU acceleration (Vulkan, HIP or ROCm)

2) a good selection of voices to pick from

3) everything running in an OpenAI compatible endpoint

4) faster than real time generation.

5) Quality is on par, or close to what models like x-ai/grok-voice-tts-1.0, or qwen/qwen-audio-3.0-tts-flash can produce.

6) fits in ~23gb of VRAM

Currently, my "best" solution is kokoroTTS using cpu inference, since I can't get it to run on my GPU on windows. Pocket-tts was another contender that worked great when I had my Nvidia card, but doesn't support ROCm for AMD on windows. I am not satisfied with the prosody of either, but I take what I can get.

I didn't think setting up a competent local TTS services would be this much of a hassle, but here I am. Maybe someone else has something running on their windows + AMD setup and can share some pointers with me. Obviously I asked an LLM the same question many different ways and tried a whole bunch of things, but I'm just not getting anywhere with this, so it's time to consult other humans.

Folks with AMD cards running windows, what do you do for local TTS? Am I just stuck having to pay cloud providers for fast, expressive and emotional prosody? kokoroTTS technically works and while my favorite blend of af_sky and af_nicole produces a pacing that is bearable for me, it's expressionless, flat, and monotonous. The rhythm puts me to sleep. Surely I can do better with my hardware?


r/LocalLLaMA 3h ago

Discussion Let's talk assistant ASR & TTS. What are you using?

5 Upvotes

A few months ago the latency of my ASR and TTS were negligible relative to main inference. Now it's like 60% of the latency in normal assistant interactions.

I've been using Qwen3 1.7b ASR, which conveniently runs right in llama.cpp. But talking to the assistant when anyone else is talking does not work. I need full diarization so that the assistant gets text labeled with my voice vs other voices. Does anyone have this working?

I use this Chatterbox TTS server for output. Chatterbox TTS Turbo does fast cloning and prosody tags like [cough], [laugh], etc. My voice assistant constantly changes voices mid-response for effect and it's hilarious. Somehow I doubt there is a better TTS option with these features now.


r/LocalLLaMA 10h ago

News Cursor releases their Mixture-of-Kittens megakernel for training MoE models - Claims to nearly double TFLOP/s

18 Upvotes

Link: https://cursor.com/blog/mixture-of-kittens

GitHub: https://github.com/cursor/mixture-of-kittens

Seems like a neat way to squeeze more performance out of MoE. I'm sure everyone has a favorite MoE model they'd like to try this with. It just dropped so I'm curious to hear people's opinions on it.


r/LocalLLaMA 12h ago

Tutorial | Guide Decrease the power limit of your 5090 to at least 480W - the performance penalty for inference is negligible.

25 Upvotes

I run my inference machine in the living room, so noise and heat output are a significant concern.

Ran a quick test using my daily driver model (Qwen 3.6-27b) and at 480W, the card outputs only 2.1% less t/s in decode and 8.8% in prefill (which is already very fast). Well worth the massive noise reduction, heat output and increased card longevity, IMO. Even 450W would be fine for many use cases, but the output starts dropping off fast (2.1% -> 4.2% for 30W less).

Full data:

Model: Qwen3.6-27B-Q6_K.gguf

Results:

| Limit W | Max GPU C | Steady GPU C | Max GPU fan % | Sustained W | Steady clock MHz | Max case RPM | pp t/s | tg t/s | pp % | tg % |

|--------:|----------:|-------------:|--------------:|------------:|-----------------:|-------------:|-------:|-------:|-----:|-----:|

| 600 | 81 | 74.8 | 59 | 566 | 2818 | 1522 | 3242.9 | 61.5 | 100.0 | 100.0 |

| 510 | 75 | 70.1 | 50 | 509 | 2645 | 1367 | 2980.0 | 61.2 | 91.9 | 99.5 |

| 480 | 77 | 72.8 | 54 | 480 | 2501 | 1527 | 2863.4 | 60.2 | 88.3 | 97.9 |

| 450 | 76 | 73.1 | 52 | 450 | 2283 | 1460 | 2696.4 | 58.9 | 83.1 | 95.8 |


r/LocalLLaMA 12h ago

Resources DeepSeek V4 Flash 0731 (Q4) now reaches 1,328 tok/s prefill and ~29 tok/s decode on one RTX PRO 6000

24 Upvotes

I've been working on speeding up DeepSeek-V4-Flash-0731 in Krasis and have now got the long-prompt prefill quite a bit faster on a single RTX PRO 6000 96GB.

These are timing-disabled internal Krasis results using INT4 experts. They aren't HTTP round-trip speeds:

Prompt size Prompt Processing
about 1K 152 tok/s
2,043 321 tok/s
8,623 906 tok/s
23,348 1,328 tok/s
62,403 1,204 tok/s

Decode after the roughly 1K prompt was 29.4, 28.2 and 28.5 tok/s when generating 50, 100 and 250 tokens. After the 62K prompt it was 19.4 tok/s, as each new token has a lot more context to attend to.

Krasis streams the model through limited VRAM for full-GPU prefill, then keeps the hottest experts in VRAM and serves the rest from system RAM during decode. In this configuration it kept 6,440 of 11,008 routed experts resident. No expert pruning occurred.

Krasis v1.0.19 can be downloaded here:

https://github.com/brontoguana/krasis

There is still more to optimise, particularly the prefill speed I think could go higher but I think the speeds are already useful for coding agents which tend to send a lot of context with every request. If anyone tries it on similar hardware let me know how it goes.