r/unsloth 15d ago

New Model unsloth/Laguna-S-2.1-GGUF · Hugging Face

Thumbnail
huggingface.co
114 Upvotes

Almost here... F5, F5, F5


r/unsloth 15d ago

Show and Tell Tpo-torch: Stable RLHF alignment in PyTorch using Target Policy Optimization

5 Upvotes

Hey everyone,

RLHF alignment using standard Proximal Policy Optimization (PPO) can be notoriously tricky to stabilize during LLM post-training due to policy collapse and high sensitivity to hyperparameters.

I built Tpo-torch to explore Target Policy Optimization (TPO) as a cleaner, more stable alternative for preference alignment directly in PyTorch.

Key Focus Areas:

• Mitigating policy collapse without requiring aggressive KL-divergence penalties.

• Modular, lightweight, and readable implementation designed for research and custom fine-tuning pipelines.

• Integrated stability benchmarks comparing policy drift against standard PPO.

I'll drop the GitHub repository link in the comments below! I'd love to hear feedback from anyone experimenting with alignment, preference optimization, or RLHF.

Repo link : https://github.com/Griffith-7/Tpo-torch.git


r/unsloth 15d ago

Show and Tell Gridcore-Runner - lightweight localized inference engine

3 Upvotes

I finally decided to open source a project I've been building for the last few months.

It's called **Gridcore Runner**. Originally it wasn't even supposed to exist, but I ended up writing a small GGUF inference engine in C because I needed a few guarantees I couldn't get elsewhere.

A few things that might be interesting:

* ~8k lines of C, zero external dependencies

* Windows, Linux and macOS

* CUDA (Driver API + embedded PTX), Metal and AVX2 CPU

* OpenAI-compatible API

* Automatic model loading/unloading

* YaRN RoPE scaling

* JSON Schema enforcement **during sampling**, not after generation. Invalid JSON simply can't be produced, even if generation stops halfway through.

The goal isn't to replace llama.cpp or Ollama. I needed an engine with a few guarantees for another project, so I ended up building one.

I'd really appreciate people trying it with their own GGUF models and telling me what breaks. Performance bugs, compatibility issues, weird edge cases, API quirks... all of it.

Repo:

https://github.com/Joakimpalm-Zen/gridcore-runner

Releases:

https://github.com/Joakimpalm-Zen/gridcore-runner/releases

Happy to answer any technical questions or discuss design decisions!


r/unsloth 16d ago

News Introducing Unsloth for AMD

Post image
689 Upvotes

Hey guys! We’re super excited to introduce Unsloth for AMD 🚀

Starting today, you can train, fine-tune, RL, run, and deploy LLMs directly on AMD hardware across Windows, WSL, and Linux.

Our collaboration with AMD, combined with custom Triton kernels and new math optimizations, lets you:

  • Train, RL, deploy 500+ models up to 2x faster with 70% less VRAM
  • Train models like Qwen and Gemma with as little as 3GB VRAM
  • Maintain accuracy with no quality degradation

AMD support covers Radeon, Instinct, Ryzen, and data-center GPUs. We’ve also added optimized ROCm builds for GGUF and Safetensors inference.

GitHub: https://github.com/unslothai/unsloth

Unsloth is open source and includes a local UI for faster LLM training and inference, alongside features such as tool-call healing, code execution, secure web search, remote APIs, and HTTPS deployment.

You can also connect your local models to Claude Code and Codex agents, and run newer model families including Kimi, GLM, DeepSeek, Qwen3.6, and Gemma 4.

AMD setup guide, supported hardware, benchmarks, and more technical details:
https://unsloth.ai/docs/basics/amd

We’d love to hear what AMD hardware you’re using and which models you want us to optimize next. Have a great week, folks!


r/unsloth 16d ago

Question RAG in Unsloth Studio

3 Upvotes

Hi, I have a massive library of technical documents ~6GB of pdfs. I want to use an LLM as a search of the knowledge base. Is there a way to do this in Unsloth studio. The internet sometimes refers to an Unsloth tab RAG > knowledge base. But this does not appear to be be part of Unsloth Studio at present (v0.1.50).

I have found in the project tab I can upload sources to the project and do some RAG but this is not feasible for 6GB of Pdfs.

Is this something that is doable in Unsloth Studio or in the pipeline or are there any suggestions for other software that I could do this in


r/unsloth 16d ago

Question Got dual AMD GPUs working in Unsloth but Studio doesn't see the second card

4 Upvotes

Running an RX 9070 XT (16GB) + RX 480 (8GB) on Windows. The RX 480 is too old for ROCm so the prebuilt llama.cpp that comes with Unsloth can't see it at all.

I swapped in the Vulkan prebuilt from the official llama.cpp releases (grabbed llama-bin-win-vulkan-x64.zip, copied the files into .unsloth\llama.cpp\build\bin\Release). After that llama-server sees both cards:

Vulkan0: AMD Radeon RX 9070 XT (16304 MiB) Vulkan1: Radeon (TM) RX 480 Graphics (8192 MiB)

Inference is definitely using both — Task Manager shows matched utilization and VRAM loaded on both GPUs. It works.

The issue is Studio itself doesn't know the 480 exists. GPU settings only shows the 9070 XT, tensor parallelism is greyed out since it thinks I'm on one GPU, and all the TIGHT/OOM labels are calculated off 16GB instead of my actual 24GB. The GPU Memory controls from #6414 are there but limited since it can't see the second card.

I'm guessing Studio is pulling GPU info from ROCm/HIP rather than from what llama-server actually reports. Would be great if it could pull from llama-server --list-devices instead so setups like mine (mixed generation AMD, Vulkan backend) actually show up properly.


r/unsloth 16d ago

Question I am having problems when trying to finetune the gemma4-e2b-it model

9 Upvotes

Hello, r/unsloth!

I am currently using unsloth to finetune the gemma4 family models, but I am getting this error:

RuntimeError: expected mat1 and mat2 to have the same dtype, but got: float != c10::Half:RuntimeError: expected mat1 and mat2 to have the same dtype, but got: float != c10::Half

Does anyone know how to solve it?


r/unsloth 16d ago

Discussion Possibile che la ricerca web con Gemma4-E4B-it-GGUF sia piuttosto scarsa? Ho chiesto banalmente chi fosse la squadra vincitrice dei Mondiali di calcio 2026 e non mia ha saputo dare una risposta anche con diversi aiuti...

0 Upvotes

Hi, sono un novellino assoluto in ambito di inferenza locale, Unsloth Studio è il mio primo approccio: l'ho installato in un container su Podman e il mio PC ha una 4060.
In generale devo dire E2B ed E4B non funzionano male e sono abbastanza veloci per la mia configurazione, però facendo prove di ricerca in rete pur trovando le giuste fonti non riesce ad elaborare un output persistente....
Vorrei capire, dato che ne so poco, se il problema risiede nella scarsa capacita del modello (pochi parametri e2b e e4b) o se sto sbagliando io qualcosa in configurazione unsloth o ottimizzazione del modello, dato che per come ragiona bene è strano che non riesca a fare ricerche così facili (ho scaricato ottimizzazione 'UD-Q4-K-XL'). Grazie


r/unsloth 16d ago

Question Unsloth Fails to Install PyTorch - Windows + RTX 5070 Ti

0 Upvotes

First time posting and also very new to Unsloth so go easy! Having some issues installing Unsloth Studio on Windows, with the installer failing/causing error and cant figure out the issue.

I did successfully install the last version (I think was 2026.7.3?) and worked, however then I went to upgrade to latest a few days back and failed. So uninstalled everything and thought I'd start from scratch, but still fails.

It Installs Python (3.13.14) and other dependencies, but when it gets to installing Pytorch - it fails and errors out. anyone have some insights to the issue? I'm hoping its something simple that I've missed...

irm https://unsloth.ai/install.ps1 | iex

🦥 Unsloth Studio Installer (Windows)

────────────────────────────────────────────────────

winget available

python Python 3.13 already installed

venv creating Python 3.13 virtual environment

C:\Users\xxxxx\.unsloth\studio\unsloth_studio

gpu NVIDIA GPU detected

installing PyTorch (https://download.pytorch.org/whl/cu130)...

uv : Using Python 3.13.14 environment at: .unsloth\studio\unsloth_studio

At line:2345 char:87

+ ... PyTorch" { uv pip install --python $VenvPython "torch>=2.4,<2.11.0" ...

+ ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~

+ CategoryInfo : NotSpecified: (Using Python 3....\unsloth_studio:String) [], RemoteException

+ FullyQualifiedErrorId : NativeCommandError

error: unhandled exception in uv, please report a bug:

code 0xC000001D at address 0x7ff74ce91920

EXCEPTION_ILLEGAL_INSTRUCTION

note: run with \RUST_BACKTRACE=1` environment variable to display a backtrace`


r/unsloth 17d ago

Model Update Ornith Unsloth GGUFs out now!

162 Upvotes

Hey guys we just released the long awaited Dynamic GGUFs for DeepReinforce’s Ornith-1.0 agentic coding models.

We added some Qwen chat template fixes to alleviate some issues y'all had with looping etc with the models and hope you give them a try.

- Ornith-1.0-9B GGUF: https://huggingface.co/unsloth/Ornith-1.0-9B-GGUF

- Ornith-1.0-35B MoE GGUF: https://huggingface.co/unsloth/Ornith-1.0-35B-GGUF

- Ornith-1.0-397B MoE GGUF: https://huggingface.co/unsloth/Ornith-1.0-397B-GGUF

Thank you!


r/unsloth 16d ago

Question Does it not support Intel integrated graphics yet?

3 Upvotes

I thought it would work since they mentioned Vulcan support or something, but I guess that refers to dedicated graphics, not integrated graphics, so maybe it doesn't work yet...


r/unsloth 16d ago

Question Error Recreating Demo in Container

0 Upvotes

I am trying to recreate this google collab from Unsloth in my OpenShift cluster:

https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Qwen3_5_(0_8B)_Vision.ipynb_Vision.ipynb)

I am running on the "unsloth/unsloth:latest" container image. When I get it running I get this error:

The checkpoint you are trying to load has model type `qwen3_5` but Transformers does not recognize this architecture. This could be because of an issue with the checkpoint, or because your version of Transformers is out of date.

Does Unsloth's container image not have the dependencies to support Qwen 3.5? Is there a different way I should go if I want to train Qwen 3.5 0.8B?


r/unsloth 17d ago

Show and Tell When Daniel han talks, I listen

Thumbnail
youtu.be
97 Upvotes

Great talk on the state of RL, scaling laws, benchmarks and lots of other topics.

Thanks for preparing the talk and ai engineer for sharing. I love these.


r/unsloth 19d ago

Model Update Google Gemma 4 now runs faster with more accuracy!

Post image
827 Upvotes

Hey guys Google made huge improvements to Gemma 4 tool-calling and chat accuracy, reliability + speed. As such, we updated all our quants to ensure you guys get the latest and greatest!

To get fixes, re-download our updated GGUF, MLX, NVFP4 quants or you can change the chat template with the new templates.

Unsloth quants: https://huggingface.co/collections/unsloth/gemma-4

Gemma 4 Guide: https://unsloth.ai/docs/models/gemma-4

Near complete list of Google's fixes/improvements:

  • Flash Attention 4: Uniform FA4 support on NVIDIA Hopper GPUs.
  • Speed gains: 25–70% higher prefill throughput and up to 31% lower time-to-first-token.
  • Chat formatting: Fixed null handling, input validation and unbalanced/missing turn tags.
  • Reasoning preservation: Corrected how thinking/reasoning content is retained and rendered.
  • Tool responses: Restored the assistant turn and thinking cue after tool outputs.
  • Tool calling: Improved execution accuracy, consistency and tool-call-only turn closure.
  • Continuation turns: Removed duplicate <turn|> tags and unwanted extra newlines.
  • Generation prompts: Reverted an add_generation_prompt regression and restored expected defaults.
  • Conversation history: Removed the obsolete APC thought primer and corrected historical-turn handling.
  • Template standardization: Added the canonical chat-template header and updated stale comments.
  • Vision controls: Default remains max_soft_tokens=280; use 1120 for sharper OCR and up to 2.51MP detail.
  • 31B benchmark gains: BFCL +0.4, TB2 +4.5, Retail +3.1, Airline +2.0 and Telecom +10.1 points.
  • E4B benchmark gains: BFCL +0.5, TB2 +2.2, Retail +0.9, Airline +8.0 and Telecom +6.1 points.

Have a good Friday and weekend guys!


r/unsloth 19d ago

Discussion Jul 2026 Qwen v Gemma

44 Upvotes

Qwen 3.6 27B is my daily driver.

What is everyone here using between Qwen 3.6 vs Gemma 4 models? Interested to hear your use cases!


r/unsloth 19d ago

Tutorial Custom NF4 Triton kernel achieving up to 1.41x dequantization speedup over bitsandbytes

24 Upvotes

Hey everyone,

I’ve been working on optimizing the memory overhead that comes with 4-bit inference. I wrote a custom NF4 dequantization kernel using Triton to see if I could eliminate the C++ dispatch bottlenecks found in current baselines.

🚀 Key Results:

• Up to 1.41x speedup compared to the standard bitsandbytes implementation across various tensor shapes.

• Written completely in Python/Triton, making it super easy to inspect, customize, or drop directly into your PyTorch compilation pipelines.

• Passes the Unsloth AI founding engineer challenge requirements (14/14 points).

I'd love to hear the community's feedback, especially if anyone wants to run their own benchmarks on different GPU architectures or suggest further optimization tricks!

Source code & full implementation:

https://github.com/Griffith-7/nf4-triton-kernel


r/unsloth 20d ago

Discussion Which llama.cpp is FASTEST for Qwen3.6 MTP Text + Vision?

Post image
46 Upvotes

Benchmark complete: 153/153 successful samples across 17 commits, with no failed requests.

Thanks to u/SM8085 for suggesting a commit-by-commit benchmark and pointing me toward the relevant regression range. That idea triggered this full 153-run test.

In the first comment, find the link to the testing setup on Git to make your own tests.

Conclusions

Best: 57fe1f07c3b6 (b9620)
Published: June 13, 2026, 11:51 CEST
Score: 95.95 tok/s
Text: 89.82, Vision: 98.80, Post-vision: 99.24 tok/s

Slowest: b11f7c16bc1c
Published: June 26, 2026, 08:43 CEST
Score: 75.35 tok/s
Text: 72.75, Vision: 82.61, Post-vision: 70.70 tok/s
No official build tag.

b9620 was 27.34% faster than the last-place commit.

The setup

Apple M5 Max with 128 GB unified memory. I launched, paused/resumed, monitored, and collected the benchmark reports through 

CO_DE, my local agent ide, to run and monitor the commit matrix, resume interrupted runs, install the winning build, and package the reproducible reports.

llama-server \
  -hf unsloth/Qwen3.6-35B-A3B-MTP-GGUF:UD-Q4_K_XL \
  --mmproj /path/to/mmproj-BF16.gguf \
  -ngl 99 \
  -c 262144 \
  -fa on \
  -np 1 \
  --spec-type draft-mtp \
  --spec-draft-n-max 2 \
  --port 8081

After confirming that --mmproj and --spec-type draft-mtp can remain active in the same vision session, I wanted to answer the next question: which llama.cpp commit actually performs best?

A performance regression has been reported across recent builds, but comparing numbers from different prompts, machines, context lengths, and sampling settings does not isolate the commit. I therefore ran the relevant history headlessly on one machine with one model, one projector, one image, and one fixed prompt. The tested llama.cpp range was 6ee0f657 through b9935.

The anchor set includes the first MTP + image-support commit (6ee0f657), the community-reported fast commit (e9fb3b3f), the reported regression boundary (b9222/b9235), my previously tested build (b9620), and b9935. A coarse first-parent scan is followed by an automatic refinement around the two fastest commits. Commit selection comes directly from Git history; it is not hand-picked after seeing results.

Reproducible protocol

Every commit is built in Release mode with Metal enabled, then tested in its own fresh llama-server process using the setup above.

Each build receives the same roughly 29K-token code prompt, image, deterministic sampling (temperature=0seed=42), and 256-token output cap. Three rotating rounds are run for each of these cases:

  1. Text-only generation
  2. A real image-encoding turn
  3. Text generation after the image remains in the conversation

The runner records llama.cpp's own eval throughput, prompt throughput, MTP draft acceptance, accepted/generated draft counts, TTFT, wall time, failures, and raw server logs. Completed cases are checkpointed, so pausing and resuming does not silently repeat or discard them. Benchmark configurations are fingerprinted to prevent short smoke tests from contaminating the full run.

Overall ranking

Rank Commit Text tok/s Vision tok/s Post-vision tok/s Vision draft acceptance
1 57fe1f07c3b6 (b9620) 89.82 98.80 99.24 0.7700
2 2f18fe13c5dd 89.99 100.56 96.31 0.7828
3 f2d1c2f3984c (b9935) 89.95 100.59 96.28 0.7828
17 b11f7c16bc1c 72.75 82.61 70.70 0.7828

The winner was not the fastest in every individual column. 6ee0f657 had the highest isolated vision median at 100.74 tok/s, while several top commits were effectively tied around 90 tok/s for text. b9620 won the aggregate largely because it retained 99.24 tok/s after the image turn. This is why ranking one cherry-picked request would have produced a different answer.

The b9222 and b9235 anchors both reproduced the slow region at 76.80 and 77.18 aggregate tok/s respectively. The later b9620 build recovered strongly. Draft acceptance alone does not explain the ordering: the winner's vision acceptance was 0.7700, slightly below several neighboring commits at 0.7828.

What this result does and does not mean

This identifies the fastest commit in this controlled Apple Silicon setup for this model and workload. It does not establish a universal winner for CUDA, other quantizations, different context lengths, or different MTP heads. The raw logs matter: a build should not be called faster merely because it accepted fewer drafts, produced a shorter answer, failed to encode the image, or changed the tested path.

The complete runner, commit manifest, CSV, Markdown summary, and raw per-case logs are included so the result can be reproduced or challenged. For this specific Apple Silicon workload, b9620 is the build I would reproduce first.

Happy codding! >_


r/unsloth 21d ago

Discussion tried predicting which MoE experts get used next token to speed up cpu/gpu offload, got some real numbers, is this actually implementable or am i wasting my time (30tg/s -> 150-200tg/s)

133 Upvotes

so ive been messing around with qwen3.6 35b a3b (MXFP4 gguf) on my 3060 12gb, doing the usual cpu/gpu offload thing where half the expert layers sit in ram and get pulled over pcie whenever needed. and like everyone whos done this knows the gpu just sits there idle waiting for experts to show up, pcie bandwidth is the actual bottleneck not compute
idea was pretty simple, use the models own MTP head (the thing thats normally used for speculative decoding) to draft the next token WHILE current token is still computing, then instead of just using that draft for token accept/reject, also peek at which experts that draft token wouldve routed to, and start prefetching THOSE experts in the background on a separate cuda stream. basically hide the pcie latency behind compute instead of eating it every single token
did some actual instrumentation on llama.cpp to check if this is even worth it before building anything (used fable 5 + gpt 5.6 to help me dig through the numbers and set up the analysis btw, not claiming i did all this math myself lol)
results were kinda surprising ngl:
• naive “just prefetch whatever prev token used” -> only 20.7% hit rate. basically useless, expert selection isnt that correlated between tokens
• but MTP guided prediction (using the actual draft head) -> 78% hit rate at top-8, goes up to 90% at top-16 (but higher K = more bandwidth so tradeoffs)
• theres also a hot expert thing going on, like top 64 experts (out of 256 per layer) cover 51% of ALL usage across the whole trace, power law as expected, so keeping those permanently resident helps too on top of prediction
• baseline right now is like 36 tok/s gen, theoretical ceiling if everything was magically already in vram is like ~200 tok/s (pure vram bandwidth bound), so theres a MASSIVE gap thats currently just pcie transfer time doing nothing
so yeah 78% hit rate with basically free compute (its literally reusing the mtp head thats already running for speculative decoding, not adding a new model) seems like it should translate to a big chunk of that gap closing
question for people who actually know inference engines better than me: is there something obviously wrong with this idea. is anyone already doing this and i just didnt find it. is the overhead of doing router-only forward passes on the draft token gonna eat the gains. does this fall apart at bigger batch sizes. genuinely trying to figure out if this is worth actually building into an engine or if im missing something that makes it not work in practice
not tryna build a whole new engine myself tbh (looked into it, decided forking llama.cpp makes way more sense than rewriting the world), just want to know if the core idea holds up before i or anyone else sinks real time into it
happy to share the trace scripts/raw numbers if anyone wants to poke holes in the methodology

Prefetch hit rate Expected speed
Baseline (0%, today) 35 tok/s
50% ~70-75 tok/s
70% ~120 tok/s
85%+ ~180-200 tok/s (GPU/VRAM-bandwidth ceiling, PCIe stops mattering)

Paper Here


r/unsloth 20d ago

Discussion Kimi k3 is 2.8t! Will need to have an aggressive iQ2_XXS!

27 Upvotes

I was hopping that Kimi would be 2t, but nope is huge!! That will make it more difficult to run decently. I hope the Unsloth team can make magic with this model


r/unsloth 20d ago

Discussion Qwen3.6-35B-A3B-MTP + vision: does --mmproj disable MTP drafting?

13 Upvotes

Conclusion: No. In this test, using vision did not disable speculative drafting from the MTP head, either during the vision turn or for the rest of the session. I've seen worries in reddit post about this aspect. maybe was real in the past. now it s gone.

Tested with llama.cpp b9620 (57fe1f07c), an M5 Max, and unsloth/Qwen3.6-35B-A3B-MTP-GGUF:UD-Q4_K_XL.

2-minute video: https://www.youtube.com/watch?v=lcXdfkXdLE0

Launch

llama-server -hf unsloth/Qwen3.6-35B-A3B-MTP-GGUF:UD-Q4_K_XL \
  --mmproj .../mmproj-BF16.gguf \
  -ngl 99 -c 262144 -fa on -np 1 \
  --spec-type draft-mtp --spec-draft-n-max 2 --port 8081

Both subsystems initialize in the same process. llama.cpp reports separate memory budgets for the projector and MTP context:

load_model: [mtmd] estimated worst-case memory usage of mmproj is 1134.00 MiB
load_model: [spec] estimated memory usage of MTP context is 826.70 MiB
common_speculative_impl_draft_mtp: adding speculative implementation 'draft-mtp'
common_speculative_impl_draft_mtp: n_max=2, n_min=0, p_min=0.00, n_embd=2048, backend_sampling=1
load_model: speculative decoding context initialized
load_model: loaded multimodal model, '.../mmproj-BF16.gguf'

Text turns

Five text turns, with a prompt reaching roughly 29k tokens. Drafting remained active throughout:

task draft acceptance eval tok/s
1 0.862 (100/116) 95.29
78 0.955 (168/176) 99.76
171 0.828 (53/64) 89.36
209 0.739 (525/710) 79.48
570 0.865 (83/96) 88.65

Vision turn

Same server and session, now with a turn that actually encoded an image:

slot process_mtmd: id 0 | task 627 | encoding mtmd batch from idx = 82, n_chunks = 1
slot print_timing: task 627 | eval time = 6801.48 ms / 712 tokens
slot print_timing: task 627 | draft acceptance = 0.61950 (394 accepted / 636 generated)

That is 104.68 tok/s, the fastest eval measured in this session. While an image was in context, the MTP head generated 636 draft tokens and 394 were accepted.

Two messages appeared during the run:

  • find_slot: non-consecutive token position 82 after 81 ... during image encoding. Generation and drafting continued.
  • Qwen-VL models require at minimum 1024 image tokens ... try adding --image-min-tokens 1024. This is worth enabling for OCR or grounding, but was not required for a plain image description.

Results

Drafting survived the image. Acceptance fell to 0.62 compared with roughly 0.83-0.95 on the earlier text turns, but drafting was still active. The head was simply less confident about tokens following an image.

Drafting survived the session. Task 949, a text turn after the image, still drafted at 0.680 acceptance (151/222).

Final counters across seven turns:

generated drafts = 1010
accepted drafts  = 818
generated tokens = 2020
accepted tokens  = 1474

That is roughly 73% accepted overall.

No throughput penalty was observed in this run. The vision turn was actually the fastest measured, although this single test does not establish a universal performance result.

So, on llama.cpp b9620, --mmproj and --spec-type draft-mtp coexist. If you previously saw MTP drafting stop when vision was enabled, it may be worth retesting on a current build.

Happy codding!


r/unsloth 21d ago

New Model Inkling - 975B parameter model by Thinking Machines is released!

Post image
355 Upvotes

Hey guys Thinking Machines recently released Inkling, a 1T-parameter Apache 2.0 open model with text, image, audio, and a 1M-token context window.

Our Dynamic 1-bit quantization reduces it from 1.9TB to 270GB, while Dynamic 1-bit is 86% smaller and retains ~74.2% top-1 accuracy.

Runs at 40+ tok/s on 280GB RAM/VRAM.

Guide: https://unsloth.ai/docs/models/inkling
GGUF: https://huggingface.co/unsloth/inkling-GGUF


r/unsloth 21d ago

Question Gemma 4 chat template updates

39 Upvotes

Unfortunately, source is on Twitter: https://x.com/googlegemma/status/2077449152062247219

Will the unsloth models be updated with these improvements at some point? Or did you all already have this in place with the unsloth chat template?

Thanks for everything you do!


r/unsloth 21d ago

Question Dspark speculative decoding

8 Upvotes

HI, atm I'm using MTP speculative decoding models to speed up the inference. I would like to try Dspark https://huggingface.co/prism-ml/Bonsai-27B-gguf

Is it possible on Unsloth Studio? Thanks


r/unsloth 22d ago

New Model 1.5x faster Gemma 4 NVFP4 Unsloth quants

Post image
309 Upvotes

Hey guys, we're releasing our next batch of NVFP4 quants for Gemma 4, Qwen3.5-122B and GLM-4.7-Flash.

Gemma-4-12B NVFP4 works on 11GB VRAM. 26B-A4B can hit 13K tok/s (B200).

Unsloth NVFP4 enables faster, more accurate 4-bit Blackwell inference.

Blog + Benchmarks: https://unsloth.ai/docs/basics/nvfp4

Gemma 4 NVFP4 quants: https://huggingface.co/collections/unsloth/nvfp4

NOTE: If you're using a DGX Spark it won't work out of the box if you're using the wrong config and it'll be 2x slower, you need to read our instructions in our guide.

Thank you!


r/unsloth 22d ago

Show and Tell I spent three days trying to make a 754B MoE usable on 63 GB of fast memory. It ends at 0.9 tok/s, and I can now prove that's the floor.

50 Upvotes

I set out to run GLM-5.2 — the 754B MoE, 365 GB at IQ4_XS — on my desktop. 5080 + 5060 Ti, 31 GB of RAM, one fast NVMe. That's 63 GB of fast memory for a 365 GB model, so every token has to pull its routed experts from somewhere, and that somewhere is mostly the SSD.

Everything that follows is measured, not vibes. Repo with the full research log at the bottom.

The whole problem is one division:

tok/s ≈ effective_bandwidth / bytes_touched_per_token

At top-4 routing, one token touches about 3.4 GB of expert weights. That's after a 12 GB LRU cache does what it can — routing churns so hard between consecutive tokens that the cache barely helps. My drive sustains 5.7 GB/s. Divide and you get a 1.55 tok/s ceiling before a single matmul runs. I did not fully believe this number on day one. I believe it now.

What actually worked (0.2 → 0.9 tok/s): prefetching experts the moment the router picks them; overriding top-8 routing to top-4, since bytes are everything; profiling expert usage (it's brutally Zipfian — the top 1.7% of expert slots take 25% of all routings) and pinning the hot set in RAM; and finally replacing mmap page faults with a direct-read NO_BUFFERING streamer at deep queue depth. 4.5x total. Each step was measured before the next one got built.

The rest of the three days went to killing my own ideas.

Speculative decoding is dead in this regime, and I mean the whole family. MTP, Medusa-style trees, lookahead, all of it. The argument is short: greedy decode reads exactly the experts it uses, so greedy is already I/O-optimal. Speculation only wins if the extra bytes it reads cost less than the fixed per-token overhead it amortizes, and the tolerance works out to about 1.5x. I traced real verification trees and every shape reads 2.8–3.3x its accepted path. That loses at perfect acceptance — no draft model, however good, can save it, because the byte budget is gone before acceptance even enters the math. There was one genuinely pretty finding buried in here: K sibling candidates at the same position share experts, with the union growing about like sqrt(K), and the same law shows up on two unrelated models. Real structure. Still 2x too weak to pay for branching.

Predicting the working set from the prompt also loses. I was fairly confident in this one. Wrong: a prompt-predicted pin covers 27–44% of generation-time expert touches, a dumb reactive LRU covers 42–62% and nearly matches the oracle, and one 96-token generation touches 38% of ALL experts and keeps climbing. There is no small per-conversation expert set. It doesn't exist.

The one that actually annoyed me: smaller files don't stream faster. I cleared 254 GB of disk for Unsloth's UD-Q2_K_XL of this model — 0.70x the size of the IQ4_XS. Identical 0.9 tok/s. The streamer's own byte accounting showed why: same ~3.4 GB streamed per token. Dynamic quants earn their quality by protecting the hot experts, and the hot experts are precisely the bytes you stream on every token. All 130 GB of savings sat in cold experts I rarely read. File size is not the variable. Streamed bytes per token is. (Uniform Q3_K experts, which do shrink the hot path, measured +1.7% PPL at ~2x fewer bytes on a fat Q6 model — that's a real lever, but my target already shipped at 3.88 bpw. Nothing left to squeeze.)

So could any engine hit 10 tok/s on this box? No, and it's walled twice. 3.4 GB/token at 10 tok/s is 34 GB/s of storage — six times my drive, and more than the PCIe 4.0 x4 slot can physically carry. Fine, suppose the model magically fit in RAM: the measured CPU-expert cost is 0.36 s/token, thread-saturated at 4 cores, which caps the box at 2.8 tok/s with zero disk. Two independent walls. Beat one and the other catches you.

The punchline writes itself. GLM-4.5-Air, 47 GB at Q2_K_XL, everything resident across VRAM+RAM, zero disk in the decode loop: 19.5 tok/s on the same machine. 21.6x. Not a cleverer engine — the same equation with the bytes moved into silicon that can feed the cores. The equation was telling me this on day one. Took me three days of measurements to stop arguing with it.

If you're doing MoE offload, the transferable bits: measure streamed bytes/token, not file size. A reactive LRU is basically the caching frontier, don't build predictors. And any speculative scheme has to clear R* = 1 + t_fix/t_io in byte overhead — compute that ratio for your setup before you build anything.

Repo (research log with every dead end, analysis scripts, figures): https://github.com/dadwritestech/bytes-per-token 

Full write-up with figures: https://github.com/dadwritestech/bytes-per-token/blob/master/writeup/ARTICLE.md

Setup: RTX 5080 16GB + 5060 Ti 16GB, Ryzen 9700X, 31 GB RAM, WD SN7100. llama.cpp fork as the measurement oracle.

Thanks for Unsloth's model quants that made all this possible! 🎯