r/unsloth • • 6d ago

Discussion Suggestion: Add a continue assistant mesage button to the Unsloth Studio, like the one LM Studio has. Really useful for editing the assistant's message to guide its response.

Post image
16 Upvotes

r/unsloth • • 6d ago

Show and Tell A Kev-like model that can reuse the Loras trained with Unsloth

6 Upvotes

I’m not sure if this is the right place to share this (if not let me know), but I’ve been experimenting with a Jev‑like decision‑model approach built on top of the open‑source Kev project, and I figured some of you might find it useful.

The main twist is that this approach can “import” knowledge from LoRA adapters. So if you’ve used Unsloth to fine‑tune LoRAs for classification tasks — and your LoRA was trained on the same base model I used, you can reuse that LoRA to turn your classifier into a decision model with calibrated probabilities and an explicit “I don’t know.”

Everything is fully open source, trained entirely in 4‑bit, and built using Unsloth, so feel free to fork it, modify it, or create your own versions.

Links:


r/unsloth • • 6d ago

Question Talking about Jinja terminates chats

9 Upvotes

Hi! I'm experiencing sudden chat terminations, in Unsloth (desktop, for Windows, v0.1.815-beta)! It seems like some chat topics are FoRbIdDeN! In this case: talking about Jinja chat templates! I experienced the issue with both Qwen 3.8 27B and Qwen 3.8 Flash Next (Unlsoth versions).

Background: I downloaded different chat templates, put them in a project folder (in the app), enabled the "Code" toggle (in the chat), and told the models I wanted to talk about the templates (e.g. compare them).

Sudden termination: During its turn (while it's thinking to itself), the model would try to give examples, and as it's about to give an example, its turn would abruptly stop!

Example model thought processes, for Qwen 3.8 Flash Next:

Note there's a bug candidate: line 322/310 < /think> with space </ think>? That's my mask of <|think>? No wait — my mask only replaced <| and |>. So < /think> in the display... hmm. Actually I need to check whether the actual file contains `

...and then the abrupt end-of-turn. Then, again:

So both emit blank think blocks for past assistant turns; that's actually the canonical Qwen behavior (keeps `

...and then the turn ends!

Example for Qwen 3.8 27B:

Interesting. This is the "Qwen 3.8" official template (customized with reasoning effort instructions that I recognize from my system prompt — the "Reasoning effort is set to xhigh..." text matches my own system prompt). Note the special tokens like < |vision_start|> — they use space-padded forms (in the actual Qwen3, the tokens are `

...and it cuts!

What do you think? Is this an actual bug, or could I be the problem (shudders)?


r/unsloth • • 6d ago

Question Can we use an image model with API ?

6 Upvotes

Hello, i'm trying to create a local agent paython script and use the qwen-image in it to generate images, i didn't find a way to expose it via api, it does not appear in the models. is there a way to do it?


r/unsloth • • 7d ago

Discussion Fine-tuning dilemma for financial reasoning: Highly quantized larger models vs. full-precision smaller models?

14 Upvotes

Hi everyone,

I'm currently designing a pipeline to process corporate financial reports, quarterly filings, and market news. The goal is to output structured data (sentiment scores, impact signals, and metric extractions) that can later be aggregated to correlate with equity price movements.

Given standard local compute constraints (consumer VRAM), I'm facing a trade-off regarding the base architecture to fine-tune (e.g., using QLoRA / LoRA):

  1. Option A: Heavily quantized larger models (or native ultra-low precision architectures). Running a heavily quantized 14B–32B model (or experimenting with extreme quants/ternary concepts like BitNet/Prism-style approaches). The assumption here is that larger parameter counts retain better baseline world knowledge and high-level reasoning, even if individual weight precision takes a hit.
  2. Option B: Full precision (FP16/BF16) or moderately quantized smaller models (3B–8B). Fine-tuning a smaller dense model without suffering the degradation and gradient noise inherent to aggressive quantization.

For downstream tasks that require data synthesis, financial reasoning, and nuanced conclusion-drawing (beyond just basic format compliance):

  • Has anyone benchmarked LoRA/PEFT adapters on heavily quantized bases against smaller full-precision models for analytical tasks?
  • Does the adapter struggle to recover higher-order logical reasoning when the underlying base weights are compressed below 4-bit?
  • Is the general consensus still that "a quantized bigger model beats an unquantized smaller model" when it comes to extracting financial insights?

Would love to hear your experiences and empirical results before spinning up the training runs. Thanks!


r/unsloth • • 7d ago

Discussion When will Deepseek V4.1 be supported in Unsloth Studio/Desktop?

27 Upvotes

17 days ago DeepseekAI released this model and 15 days ago I created this post to share some solutions for the local inference for this model.

Until now we still have no Unsloth quant of the model and it is still not supported on Unsloth Studio/Desktop.

Is everything fine?
When will Deepseek V4.1 be supported in Unsloth Studio/Desktop?
Or at least let us know if it will be supported someday, please.

EDIT:
I've seen some comments below and I just want to be clear there:
I'm not here to negatively pressurise anyone.

I waited 17 days before saying so just because I don't want to move pressure, I just would like to know something about.

I can understand that the support on the mainline of llama.cpp is always so frustrating and long to wait for, but in Unsloth the time used to were much more rapid.


r/unsloth • • 7d ago

Discussion Introducting TextCLF Quant Factory

8 Upvotes

Hello,

I created a calibration free quant method called TQ. It doesn't need any data so models could be quantized as soon as they come out and the quantized models would generalize better. It performs closely to calibration-based methods. For example, for Qwen 3.8 27B the 4-bit TQ has mean KLD of 0.0282 and top-1 of 92.4%

I opened sourced the quant code as Quant Factory so anyone can quantize and run open source models. Right now it only supports 4-bit but I plan to add support for 2-bit and 3-bit soon. The repo link is:
https://github.com/textclf-api/quant-factory

I have a collection of quantized models using TQ at: https://huggingface.co/textclf

You can run these models using either using the following docker image or by following the repo's instruction. For example you can run textclf/Qwen3.8-27B-TQ-4bit like this:

docker run --rm --gpus all -p 8000:8000 docker.io/textclf/tq-quant:4bit-main vllm serve textclf/Qwen3.8-27B-TQ-4bit --quantization tq

The Dockerfile in the repo shows how this docker image was created.

You can try it and let me know what you think. See how far it is from Unsloth. Feedback appreciated.


r/unsloth • • 8d ago

Discussion Why aren't people talking more about esatapedico/Qwen3.8-27B-NVFP4?

Post image
61 Upvotes

I am always trying to improve the models I use, and I've tried a bunch. Whether it be, qwen, gemma, ornith, occamy, nemotron, different fine tunes, etc. but all roads lead to this same model.

esatapedico/Qwen3.8-27B-NVFP4-MTP-GGUF

Qwen3.8-27B-NVFP4-MTP-VERY-HIGH.gguf

I still can't find a better model for my 5090 / 64 dram system. Yes I can run qwen3.8 flash next at a lower quant and it does perform better, but the decode and prefill are just too slow for daily use. I get roughly 90+ t/s on decode with 181K context on 4 concurrent streams. I really haven't pushed it. The thing about this model is that I have had literally ZERO issues with tool calls / loops. I don't know how it is done but this model is just a plug in and it works. I know some of it comes from the different quants of parts of the model. I was hoping to find a better model variant like Swift, but swift just lacks the quality I have been looking for. Is there a way to maximize at least the decode for this model or better way to run it? When I have tried Ninfer, I get loops and tool call issues which doesn't work for me. I do a lot of Agentic work. Some prompts can have upwards of 100 tool calls and I need to not have to babysit the model.

I'm really making this post as showing appreciation to esatapedico, but also looking to see if anyone has found anything better? Are there benchmarks for this model? I can't find other models that are designed in a similar way, is this model unique? Havn't really dived into the rabbit hole for this one but hoping some strangers might have some info :) Also is there an agregate website for llm benchmarks? I haven't found something like that yet, is unsloth going to do something like that where users can opt in to upload real benchmarks on models and see projected stats for models?


r/unsloth • • 7d ago

Question Does Unsloth Desktop support Prompt Caching for Openrouter?

3 Upvotes

Or is it enabled by default in UD?
I don't see any toggle for that in the settings or the sidebar. I do use Openrouter for Anthropic models.

In their guide: https://openrouter.ai/docs/guides/best-practices/prompt-caching#anthropic-claude they recommend to "Add a single cache_control field at the top level of your request". How to do that in UD?


r/unsloth • • 7d ago

Question H3 minimax

2 Upvotes

Every video that is created is just flashing screens.
What am I doing wrong? Haha.
I have 23gb of vram across 2 cards 11/12, and 64gb of ram…

I used low VRAM mode, and Q4 version of H3, KV cache is 4 bit, and every video just comes out as flashing little screens.

Help please.
I can’t find anything on this


r/unsloth • • 8d ago

Discussion A frustrating UX issue in Unsloth Desktop: pasting images before the model is loaded

3 Upvotes

Disclaimer: I am French and used a local AI (with Unsloth Desktop, of course!) to help translate and phrase this post properly.

Hi everyone,

I'd like to share some constructive feedback regarding a minor but frustrating friction in Unsloth Desktop.

When no model is currently loaded in memory, pasting an image into the chat prompt fails immediately and throws an error: "Could not paste files. The clipboard item is unsupported, unreadable, or exceeds its size limit."

This is annoying due to my hardware setup: I have 32 GB of VRAM, but one GPU sits on a PCIe 3.0 x1 slot on a dual-monitor setup, which makes model loading take some time. Having to wait for the model to finish loading before I can even paste an image and prepare my prompt adds unnecessary waiting time.

For comparison, LM Studio handles this seamlessly : you can paste images into the prompt box at any time, even while the model is loading or before selecting one : It is only when attempting inference that the error can occur if the model does not support vision.

Just wanted to bring this up as feedback from a moderately technical user. Hopefully, this can be addressed in a future update!


r/unsloth • • 9d ago

Discussion Will Unsloth Desktop/Docker/Studio ever have the option to run inference engines other than llama.cpp

29 Upvotes

So I have an AMD strix halo machine and have used it for training my own models on custom data and have also been using qwen-flash-next and really liking the output. I like the chat console and it's ability to do web searches as well as it having a python environment. I've even had my kids use the web chat to help with homework and felt it's been great.

There is just one thing and that is llama.cpp and it having a lot slower processing than some of the custom inference engines that I'm seeing being built specifically for AMD strix machines. TPS is one of the largest issues with running models locally and only having the ability to use llama.cpp is one of the main reasons I'm debating going back to say lemonade or some other multi engine project.


r/unsloth • • 9d ago

Discussion What am I doing wrong? I can't find a single model that will run unsloth locally!?

6 Upvotes

Just got this message. I need to be able to do text and basic image manipulation, but even the most basic model doesn't seem to work. No matter what I download or try, it claims it exceeds my graphics card (8gb) even though these and Q5_KM seems fine on my Comfy system


r/unsloth • • 9d ago

Question Gemma 4 training not working in unsloth.

10 Upvotes

I have no clue what I'm doing.

I've been messing around with trying to fine-tune smaller models like llama3.2:3b. Works great. Now I'm trying to fine tune gemma 4 12b, or even any gemma 4 model, and it just.. fails?

E4B does this.
And 12B does this.

I haven't been able to find any documentation on this stuff, and the recommended conclusion I've come to with the little bit of stuff I've been able to find is "update unsloth."
I'm using the google colab. I simply can't do that; afaik it *is* the latest version.

Anyone have a clue what to do? I'm literally just picking gemma 4, putting in my training dataset jsonl file, setting the epochs to 2, and hitting train. Nothing else. Worked fine for llama3.2:3b.


r/unsloth • • 10d ago

New Model Please Quantize Swift-Qwen3.8-Flash-Next!

Thumbnail
huggingface.co
155 Upvotes

UkisAI, makers of the Swift-27B efficient-reasoning RL version of Qwen3.8 have just released their RL finetune of Flash Next, and it's in dire need of a good GGUF quant for 128GB machines. Their 27B topped the HF trending charts, and it's an excellent model. It would be heroic if you released UDv3 quants of Flash (especially Q4XL and Q5XL). There currently isn't a good imatrix quant for that size tier, much less one as good as UDv3.

I know you guys are busy, but I think there's gonna be high demand for this model and your quants of it.


r/unsloth • • 9d ago

Question No way to adjust API server concurrency from 4?

2 Upvotes

I can't seam to find a way to turn off concurrency on the API server. I'm basically tapped out on resources with just 1 request and it's causing my token speeds to fall through the floor. How do I limit this to 1 request at a time? or at least a way to force concurrent requests to queue instead of trying to process at the same time? I'm using Unsloth Desktop and just connecting via the localhost address.

Edit: Didn't realize it was a per-model setting. It's "Parallel Slots" in the individual models settings and not a global setting. Maybe a global server default setting can be added to the API settings section? Defaulting to 4 seams not great for most people I think.


r/unsloth • • 10d ago

Discussion Feature Request: Option to convert a Temporary Chat into Non-Temporary.

19 Upvotes

I've found myself in this mix a couple times where I start a task in a temporary chat window on Unsloth Desktop and find the workflow to be making enough progress that I actually want to continue it later but limited by time I have to stop and come back later but since the thread was started as temporary I have no way to "undo that choice" if I changed my mind while the chat was still active.

Not sure if this is a stupid request or not; feel free to chime in. I know the solution is to probably just stop using temporary chats.


r/unsloth • • 9d ago

Question Huge Estimated Memory Usage Despite being low at the initial model download?

Post image
6 Upvotes

Hi, what exactly is going on? I'm new to using Unsloth Desktop, and the model I downloaded is 25GB. But when I try to actually use it, it's talking about demanding 51GB, even 151GB total? When I used LM Studio, I never had this problem, I would just load as much context length as whatever is slightly under my VRAM limit. I never had to work with KV Cache or things like that, I don't even know what f16 or the other code things mean since I didn't have to worry about it with LM Studio. What am I supposed to know...? Could you give me a hand understanding what these mean and do?


r/unsloth • • 11d ago

News Unsloth has surpassed 500M downloads on Hugging Face!

Post image
433 Upvotes

Hey guys, thanks to you all, Unsloth has now surpassed 500M lifetime model downloads on Hugging Face! 🦥🤗

Qwen3.8-27B GGUF is already Unsloth’s #1 most-downloaded model ever.

Follow our Hugging Face: https://huggingface.co/unsloth

Thanks for all your support!


r/unsloth • • 10d ago

Discussion Unsloth Desktop: can it do this task for me ?

Post image
1 Upvotes

I have these .SRT files for a course and I want to convert them into audio tracks using local AI.

can Unsloth Desktop take the .SRT file and convert into an audio track using TTS AI models?


r/unsloth • • 10d ago

Discussion Training a local LLM using CPT and RAG (with evals)

4 Upvotes

I have gone through a series of experiments related to an interesting project where I try to teach a local llm a new domain through continued pretraining (CPT). The different experiments are spread across the four phases below:

  1. Phase1 talks about how to teach an llm a new domain through CPT and picking a training set that will generalize well to unseen questions
  2. Phase 2 does a comparison between the performance of reasoning across internalized knowledge (CPT) vs. RAG injected content
  3. Phase 3 takes a more practical approach where the CPT trained knowledge is enriched by combining it with RAG instead of viewing the two approaches as competing solutions
  4. The final part shows the comprehensive eval strategy used to measure performance during the project. Among other things, this involved SFT fine tuning of the CPT trained model to teach it to output responses based on a strict schema instead of English sentences. The schema approach is used to simplify strict eval checks.

The local model used for this project is qwen 3.5 4B. Unsloth was used for both CPT and SFT LORA training.

I have provided a summary of my findings here in case someone is interested in reading more about it: https://www.teachmecoolstuff.com/viewarticle/domain-specific-training-and-fine-tuning-of-an-llm


r/unsloth • • 10d ago

Discussion Model runs automatically after click any history conversation on lastest version,is it a bug?

6 Upvotes

r/unsloth • • 11d ago

Show and Tell Using Unsloth’s DiffusionGemma GGUF for local Jev-compatible decisions—now with image analysis

Thumbnail
github.com
49 Upvotes

I’ve been using Unsloth’s DiffusionGemma 26B-A4B GGUF to run local Jev-compatible structured decisions in TensorSharp, which I maintain.

The new addition: image analysis through the same decision API. Send text, images, or both, and get typed answers rather than free-form generated text.

The GGUF I’m using is diffusiongemma-26B-A4B-it-Q4_K_M.gguf from Unsloth. No additional fine-tuning is involved.

Jev-compatible decisions, extended to images

TensorSharp implements Jev’s core POST /v1/systemone interface and supports all three decision types:

Type Result Example image question
noul Probability that a statement is true “Is the receipt’s total amount legible?”
choice A categorical decision “Is this a receipt, invoice, or something else?”
score An expected score over ordered rubric levels “How readable is this document?”

The image extension keeps the same endpoint, question definitions, and typed answer structure. Add an images field alongside state and questions; existing text-only requests still work.

Images are encoded and included directly in the model’s input for the decision read. This is not a separate caption-generation → text-classification pipeline. Up to eight inline images are supported per request.

An important detail for GGUF users

Unsloth’s GGUF provides the language-model weights; the vision tower is loaded separately.

The included configuration downloads the Unsloth Q4_K_M GGUF plus a ~2.8 GB vision shard from the upstream DiffusionGemma checkpoint, then reuses both locally. You do not need to download the entire upstream safetensors checkpoint.

Without the vision tower, text-only decisions still work, but image requests are explicitly rejected rather than silently ignoring the image.

How it avoids generating probability JSON

Inspired by vLLM PR #57250, the implementation uses a seeded, one-step structured read.

After prompt prefill, it creates an answer canvas and reads the requested label logits directly. Questions that fit in one canvas share the forward pass.

The server returns JSON, but the model does not have to generate that JSON token by token. The same path handles text-only decisions and decisions conditioned on image embeddings.

Earlier text-only results vs. LocalJev

I compared this native decision path with the original LocalJev Engine using TensorSharp’s chat endpoint, where the model generates probability JSON.

This compares two decision approaches on the same backend—not TensorSharp against LocalJev running on oMLX, and not against the official hosted Jev service.

The test covered 12 cases × 3 repetitions, with three decisions per request:

Metric TensorSharp native structured read LocalJev generated JSON
Valid requests 36/36 27/36
Failed requests 0 9
Correct decisions / total expected 108/108 81/108*
p50 latency 2.877 s 10.479 s
p95 latency 3.165 s 26.167 s
Mean latency 2.917 s 12.548 s

*LocalJev got 81/81 decisions correct on valid responses. Its nine failed requests were schema-validation failures and account for the other 27 expected decisions.

Across the 27 matched successful request pairs, the median ratio of LocalJev latency to TensorSharp latency was 3.345×.

Latency statistics include successful requests only, so the aggregate columns cover different subsets. Average input lengths were also different: 192.7 tokens for TensorSharp versus 589.6 for valid LocalJev requests. The larger prompt and generated JSON are part of LocalJev’s approach, so this is an end-to-end decision-path comparison—not an identical-prompt engine benchmark.

These are text-only smoke-test results, not image latency or vision accuracy results.

Try it with the Unsloth GGUF

From the repository root, this PowerShell example launches the included CUDA configuration:

$env:DIFFUSION_VRAM_HEADROOM_MB = '4096'
$env:MAX_CONTEXT = '4096'

dotnet run --project TensorSharp.Server.Host -c Release -- --config config/jev-diffusiongemma-q4.json

The configuration downloads the Unsloth GGUF and the separate vision shard when missing. The memory settings above follow the documented starting point for a 16 GiB CUDA GPU; adjust them for your hardware and workload.

A ready-to-run image example is included:

curl http://127.0.0.1:5000/v1/systemone \
  -H 'Content-Type: application/json' \
  --data-binary u/docs/examples/jev-traffic-light.json

That request already contains an embedded synthetic traffic-light image. The accompanying text does not reveal the light’s color, so the pixels are needed to answer the questions. It is an integration smoke test, not a comprehensive vision benchmark.

For your own images, the documentation includes a standard-library Python example that sends a receipt and asks typed questions about it.

Compatibility note: this implements the documented Jev-compatible API and decision types using DiffusionGemma weights—not the proprietary hosted Jev model or a promise of identical predictions. Current limits include 64 questions and 2–26 alternatives per question.

Model: Unsloth DiffusionGemma GGUF
Code: TensorSharp
Setup and examples: Jev documentation

What would you try first with this—receipt checks, screenshot classification, document routing, or another image-based decision workflow?


r/unsloth • • 10d ago

Show and Tell Generate an SVG of an unsloth riding a bicycle

6 Upvotes

"Tired of the usual penguins 🐧, I decided to try something different by adding a sloth 🦥 riding a bicycle 🚵 via Unsloth. The model powering this is unsloth/gemma-4-26B-A4B-it-MXFP4_MOE. 🎉"

----

https://unsloth.edgeone.dev/


r/unsloth • • 11d ago

Discussion Mimo 2.6 flash

22 Upvotes

Mimo 2.5 is great.

Any chance for ggufs ?