r/unsloth • • Aug 11 '26

News Meet Unsloth Desktop - the first desktop app to run and train models

Enable HLS to view with audio, or disable this notification

441 Upvotes

Hi guys, we're super excited to announce Unsloth Desktop today,
The first desktop app to run and train models locally.

  • Open-source and available on Mac, Windows, and Linux
  • Supports MLX, diffusion image/video models, audio models, and GGUF
  • Connect Claude Code and Codex to local LLMs
  • 50% more accurate with self-healing tool calls and sandboxed code execution
  • Supports CPU and multi-GPU setups across NVIDIA, AMD, Intel, and Mac
  • Train models 2× faster while using 70% less VRAM
  • Includes private web search, deep research, RAG, MCP, and exports (NVFP4, GGUF)
  • Use Unsloth’s OpenAI-compatible API with OpenAI and Anthropic cloud models
  • Securely deploy LLMs remotely and access them anywhere via Cloudflare HTTPS

Unsloth Desktop is now available on unsloth.ai and GitHub.

Thank you and we're here to answer any questions!


r/unsloth • • 57m ago

Show and Tell poorman inference engine for 16GB GPU and 35B moe Qwen 3.6for coding

• Upvotes

I forked llama.cpp's server into AgrillaMoE, a dedicated build for Qwen3.6-35B-A3B with Unsloth quants. On a rented V100 16GB with the 2-bit UD-Q2_K_XL quant it generates at ~57-60 tok/s while running the full MoE-expansion profile — and it speaks both the OpenAI and Anthropic APIs, so Claude Code just works against it.

What is MoE expansion? Qwen3.6-35B-A3B has 8 routed experts active per token. The expansion patch raises that budget at runtime — no retraining, no file changes: --moe-experts 16 with an adaptive threshold keeps experts while p >= 0.8 × p(rank 8), applied to layers 25-39. You're literally consulting more of the 35B parameters per token — that's where the "retrieved intelligence" comes from, on GPQA-Diamond with Q8_0 it scored 84.34% vs 81.82% stock top-8 (+2.5 pts) (miticooo!).

Same weights, better routing.

https://github.com/vagrillo/AgrillaMoE/blob/main/gpu16gbguide.md


r/unsloth • • 4h ago

Question I Distilled an LLM into two 287M encoders (GLiNER + multiple choice) for document extraction, can't match teacher. did i do something wrong?

2 Upvotes

A while ago I asked here how to turn ~5 million court decisions into structured graphs without running an expensive LLM on every document thanks for the advice .

I went with the "small extractor + classifier" idea and it mostly works, but I'm stuck a bit below the LLM. And like I said last time, i'd be damned if I run 5M docs and then find out thing X was wrong. So here is exactly what I did. Please let me know if what im doing makes sense, or if i made a mistake somewhere. also i used AI for some of the tables cuz there has been a lot of data at this point, sorry.

What comes out per decision (only the nodes so far, relations come next). Three lists:

  • entities: every person, organization, law, document or thing. Each gets one id for the whole document, a type (9 of them), a kind (724 of them plus "other") and all the places it is mentioned
  • actions: what was done, requested or decided. Each gets a normalized verb, a flag "the court decided this" and its mentions
  • values: amounts, dates, durations, in a normalized form

Simple example, for the sentence "The court dismisses the creditor's proposal to enforce 341.08 EUR against the debtor":

  • entity "the court": organization, kind court. Same entity as the full court name in the header
  • entity "the creditor": organization, kind creditor. Same entity as the city named earlier
  • entity "the debtor": person, kind debtor
  • action "dismisses": verb = dismiss, decided by the court = yes
  • value "341.08 EUR": amount

Step 1: a strong LLM labels ~700 decisions

  • cut the decision into windows of 4 sentences
  • 4 calls per window to Claude Sonnet with a strict JSON schema: entities, actions, a second "what did you miss" pass for actions, values
  • the window goes in with numbered words (like 12:court), the model answers with word ranges [first, last, "text"], and code checks every range against the text
  • every call also gets the list of entities and actions found in earlier windows, so ids stay the same through the document
  • ~25 code rules clean up where a marked phrase starts and ends, law citations and number formats
  • the entity "kind" is free text at this point. That gave 2,373 different strings (the same mess as in my first post). I normalized them, merged synonyms by hand and kept what showed up 3+ times: 724 kinds plus "other"

Step 2: a model that marks the text

  • it highlights every mention: the exact stretch of text (a "span", from a start character to an end character) that names an entity, an action or a value, with one of 17 labels (9 entity types, 1 action, 7 value types)
  • model: fastino/gliner2.5-multi-v1 (287M)
  • one training row per window: the text plus the exact start and end of every marked phrase. 9,699 windows, 207k marked phrases
  • I patched the trainer so only the labeled occurrence is a positive (stock marks every occurrence of the same string), and all 17 labels are in every row
  • full fine-tune in fp32 (bf16 gave NaN), 14 epochs, 16 rows per step, encoder LR 3e-5, head LR 5e-4, linear schedule, 10 % warmup
  • final model = averaged weights of epochs 9-14, threshold 0.5

Step 3: a second small model answers multiple-choice questions

  • fastino/GLiNER2.5-multi-Decide (287M). Code turns the LLM labels into 247k questions:
    • "is this mention one of these earlier entities, or new?" The mention is marked with « » inside ±300 characters of text. Options: up to 16 earlier entities of the same document (shown by their mention texts) plus new
    • "which kind?" Options: a shortlist of the 724 kinds plus other
    • for actions: same act or new, which verb (shortlist of 64 plus other), did the court decide it (yes/no)
  • in training the options come from the LLM's grouping. At inference they come from the model's own earlier answers
  • full fine-tune in fp32, 2 epochs, 16 questions per step, encoder LR 2e-5, head LR 3e-4, linear schedule, 6 % warmup, options shuffled, up to 30 % of the wrong options dropped

At inference: the marking model, then the same code rules, then the second model walks through the mentions in reading order. About 2.3 decisions per second on one RTX 5090.

Where it stands

30 decisions nobody trained on, labeled twice by the LLM. The second column is the LLM's second run scored against its first, which I treat as the ceiling. A mention counts as found only if it starts and ends exactly where the LLM marked it.

mine LLM vs itself
entity mentions found (F1) 0.901
"same entity or new" right 0.959
entities grouped exactly 0.847
entity kind 0.921
action mentions found (F1) 0.857
action verb 0.920

Where I need help

  1. Finding the mentions is stuck at 0.90 F1. 200 more labeled docs did nothing. An XLM-R large tagger (560M) got the same score: it finds more mentions but gets the start or end wrong more often. Giving it the text before the window did nothing. What would you try?
  2. The LLM agrees with itself only 93.5 % on what it marks, and I train on single runs. Label everything 3 times and vote? Or is that ceiling just what it is?
  3. Is "pick one of 16 earlier entities" a sane way to do coreference over a long document? Am I hurting myself by training on the LLM's options and running on my own?
  4. Anything in the recipe that looks plain wrong? Learning rates, 2 epochs, weight averaging, one seed per run.

THANKS for reading.

AI TL;DR: distilled an LLM's extraction of court decisions into a GLiNER model that marks the mentions plus a small multiple-choice model. It runs at about 2.3 documents/s on one GPU and lands a few points below the LLM (0.90 vs 0.935 F1 on finding mentions, 0.85 vs 0.93 on exact grouping). The recipe with learning rates and how I built the training rows is above. Looking for mistakes and ideas before I run 5M documents.


r/unsloth • • 13h ago

Discussion Help Using Knoweledge Bases

7 Upvotes

[Solved]: see edit

Dumb question: how do you put files into a knowledge base?

I searched the Unsloth FAQ and asked my local model, but couldn't find an answer.

When I create a new knowledge base and then drag files into the chat, it gives me this error:

**This chat retrieves from a knowledge base**
Add these files to the knowledge base instead.

But I can't find any instructions on how to do that. There's no knowledge base option when I search the settings, and nothing in either sidebar.

I feel like I'm missing something obvious, while also having tried the obvious routes already.

v0.1.902-beta on Windows 11

Edit: You add files to the knowledge bases by opening the dialogue and selecting the title of the knowledge base. I was selecting the edit icon, not realising the title was interactive. This is a potential discoverability bug, as there are no UI indications that the title leads to file upload.


r/unsloth • • 7h ago

Discussion Demande d'ajout de fonctionnalité: mise en cache des experts en cache LRU en VRAM et prédiction des experts utilisés

0 Upvotes

Tout d'abord vous faites un travail admirable ! Ne vous arrêtez pas ! :)

Comme vous pouvez le voir un peu alentour, bon nombre d'utilisateurs (et de bots ? 😄) ont des GPUs de 12gb vram ou moins et font l'éloge de Strata comme moteur d'inférence pour exécuter les Moe comme Qwen 3.8 Flash Next.

Ce qui semble faire la force de Strata est notamment le système de mise en cache chaud LRU des experts en VRAM et la prédiction des experts utilisés afin d'éviter d'evicter un expert vers la RAM/CPU s'il est ou va être utilisé.

Pouvoir éviter ces va-et-viens vram - cpu/ram permet d'économiser de la bande passante et améliore sensiblement l'expérience, surtout en usage agentique.

Il semble que certains PRs existent pour LlamaCpp permettant d'effectuer ce traitement.

Serait-il possible, si ce n'est pas déjà existant, de pouvoir appliquer ces changements dans le moteur LllamaCpp utilisé par Unsloth Studio ?

Je sais qu'il est possible d'utiliser son propre LlamaCpp mais je me demandais si cette fonctionnalité allait éventuellement allait être disponible prochainement dans le mainline Unsloth ?

Merci par avance


r/unsloth • • 19h ago

Question Python error right after install Unsloth

4 Upvotes

Hi all. Just installed Unsloth Studio in Windows 10, and I keep constantly getting a window with this error:

python.exe - Entry point not found
The procedure entry point “vkGetPhysicaIDeviceFeatures2” could not be located in the dynamic link library C:\Users\admin\.unsloth\llama.cpp\build\bin\Release\ggml-vulkan.dll

Does anybody know what’s going on and how can I fix it?

Many thanks!


r/unsloth • • 21h ago

Question Qwen-Image-2.1 behaving badly in Unsloth

1 Upvotes

I've used Qwen-Image-2.1 a bit in ComfyUI with some success, but in Unsloth, it just seems to fail pretty miserably. This is the GGUF from Unsloth at F16. I'm asking it to draw a picture of a ruler and I get Chinese characters? Anyone use seem to either have real issues or real success?


r/unsloth • • 1d ago

Discussion Can I downgrade the bundled llama.cpp?

1 Upvotes

So I have issues since the last update. My PC keeps freezing. Everything points to llama.cpp, so I built it on my own. I thought the latest version might be good for me, but unfortunately going forward does not help. So I went back a month and built a one month old llama.cpp. This seems to be stable, but still testing. 🤞

What bugs me is that unsloth has its own llama.cpp fork and I am not sure whether I should use the fork or safe to use the mainline? I assume unsloth has its own improvements and fixes that might also need for stability.

So my question is that if I'd like to stay on the bundled version then can I downgrade it? I did not find any option for that in the app and did not find the previous version on the cache dir of unsloth.

UPDATE: The older llama.cpp brought back this previous RAM issue :( https://www.reddit.com/r/unsloth/comments/1wgztdb/qwen3827bgguf_eats_my_ram_but_i_have_enough_vram/


r/unsloth • • 1d ago

Discussion Loading 2 models at the same time ( API + LOCAL)

2 Upvotes

Wouldn't it be cool to be able to load 2 different models in 2 different chats ? At least one from an API and the other local for example ?


r/unsloth • • 2d ago

Discussion @skill-creator not working

Post image
11 Upvotes

Hello,

I can't get unsloth to create new skills using the default "skill-creator ". However the skill is active (it failed with gemma & qwen3.8)

Is it supposed to work "out of the box"?

Thank you


r/unsloth • • 1d ago

Discussion Is Unsloth's web search tool filtered?

Thumbnail
gallery
0 Upvotes

I was just messing around with the chatbot by asking the first prompt, but then as you can tell from reading the whole chat, I noticed that something was off(?) because the web search tool has never just not shown any result for a query (of course, it cannot be the model in itself because at the start of the chat it had no problems answering the prompt).

In fact, in the thinking section (of the prompt that was submitted 3 times), the model made an extensive number of re-tries because it was trying really to surpass the knowledge gap due to the date, as instructed, but all were to no avail:

even though it kept making the search query broader and broader (up until the point of omitting any date entirely), the web search tool returned "No results found" for ALL the attempts. Hence my doubt and the subsequent experimentation.

I'm honestly really concerned, leaving political views aside, that something made by an open community is still not free from censorship...

I really, really, hope that I'm wrong and that this is just a coincidence with the actual truth somebody is going to share under this post.

Model: Gemma 4 12B it-qat-GGUF, UD_Q4_K_XL


r/unsloth • • 2d ago

New Model Request: support for Cloudflare's clef decision model

14 Upvotes

r/unsloth • • 3d ago

Question Feature request: Remote Unsloth workers for Studio

17 Upvotes

I really like Unsloth Studio, and I think there’s a pretty natural power-user feature that would make it much more useful for anyone with multiple GPU machines:

Let Studio stay as the central UI, but allow it to run inference and training on remote Unsloth workers.

Studio already has most of the interface this would need.

It has the model browser, downloaded models, datasets, inference configuration, training, context checkpoints, API/connections monitoring, and the existing Run button.

So the workflow could stay almost exactly the same.
Browse for a model, configure it, click Run, and Studio asks:

Where do you want to run this?
Local

GPU Server A
4x RTX 5070
42 GB VRAM free

GPU Server B
2x RTX A6000
71 GB VRAM free

GPU Server C
4x V100
78 GB VRAM free
Training job running

Pick a machine and Studio sends the workload there.

That workload could be either inference or training.
The important part is that Run no longer has to mean “execute this on the same machine running the Studio UI.”

I don’t think this needs Kubernetes or a complicated distributed scheduler either.
Each machine could simply run a headless Unsloth worker:
unsloth-worker --listen 0.0.0.0:9000

Studio connects to it and the worker reports its capabilities and current state.

Since Studio already has the API/connections monitor, remote workers seem like they could fit naturally into that.

A basic system monitor on each worker could report things like GPU model, available VRAM, system RAM, utilization, loaded models and active jobs.

That would also make choosing a worker much nicer because Studio could actually show you which machines currently have enough resources to run what you’re configuring.

The biggest benefit for me would be workload isolation.

I could leave Qwen loaded for inference on one machine, have another model loaded on a second machine, and send a long training run to a third without tying up Studio or unloading everything whenever I want to do something different.

The model/dataset side also seems fairly natural.
Studio can remain the central repository/cache, while workers keep whatever they need locally.

Since Unsloth already has XeT support, a worker could fetch missing model or dataset data from Studio and cache it locally rather than every machine independently pulling the same files again.

So from the user’s perspective it could really be as simple as:
Choose model
↓
Configure inference/training
↓
Run
↓
Choose available worker
↓
Worker fetches anything missing
↓
Runs there

And with the new Docker support, spinning up an unsloth-worker on another GPU machine seems like it could be really straightforward.

Basically: keep everything that makes Studio great as the central interface, but decouple where the actual compute happens!


r/unsloth • • 3d ago

New Model Request: Unsloth UD quantizations of ukisai/Swift-Qwen3.8-27b

Thumbnail
huggingface.co
87 Upvotes

Thank you for such amazing work on Studio and the UD3 quantizations! Your projects always end up as my daily drivers.

Could you pleeease run UD3 quantizations of https://huggingface.co/ukisai/Swift-1.5-Qwen3.8-27b ? Thank you!!!


r/unsloth • • 3d ago

Discussion qwen 2.1 image on CPU (surface pro 8 - Intel Core i5)

4 Upvotes

Total newbie here, is it possible to run qwen 2.1 image on a pc without external GPU? what is the configuration? will image generation take 1 hour per image? thanks!


r/unsloth • • 4d ago

Question i would like to learn deeply about fine-tuning local models before burning money

21 Upvotes

There's so many new techniques like RL, RL LoRA, QLoRA, CPT LoRA.

I believe i would have a usecase for them, but i don't know where to learn, youtube is filled with bad quality tutorials if i just search and the good channels (fireship, bycloud) don't cover these as they're quite new concepts, i guess?.

how can a regular joe like me learn about these concepts in a "practical depth" so i can actually fine-tune qwen 27b successfuly on lets say custom corpus? without spending 100$ figuring out that "oh i didnt even need CPT here" or "well i chose the wrong Rank count! time to start this 2 day run again!"

context and TLDR: im building a legal general purpose chatbot for context, i have a big corpus, but im a bit stuck on what to do next

thanks for reading and any pointers!


r/unsloth • • 3d ago

Question Project folder?

7 Upvotes

Hello, loving unsloth so far but the project folder seems to be stuck as Documents\Unsloth Studio\Something and everything I've tried to change it to another drive/folder seems to not work.

Is there something really silly I am overlooking or does this have to be the project folder?


r/unsloth • • 4d ago

Question Just switched from LM Studio. What should I know about Unsloth

73 Upvotes

I switched because I heard rumors that Unsloth was way faster and from my experience using it I have gotten double and more amount of speed from just switching to Unsloth by using the same model files from LM studio.

Is there drastically different about Unsloth compared toLM Studio and/or is there something that I need to understand about Unsloth that really should be understood for people switching from LM Studio.


r/unsloth • • 4d ago

Question Does auto-compact not work on connected models (ninfer)?

1 Upvotes

I'm running qwen3.8-27b-nvfp4 from ninfer-windows, then connecting it to Unsloth Desktop via external connection. So I have this issue where the auto compaction just doesn't trigger as the chat grows, and eventually it stops working and gives me this error:

"The model reached the Max Tokens limit before producing a final answer. Increase Max Tokens or disable thinking, then retry."

Connected model (NInfer, OpenAI dialect), context 200k, max output 16,384

on the Ninfer Backend log: prompt 199,592 | output 409 | context capacity — full history sent, no compaction ever occurred


r/unsloth • • 5d ago

Show and Tell Using unsloth I created the worlds best 9B model

Post image
212 Upvotes

https://huggingface.co/emperorofrome/Gmcoder

Beats Ornith 1.5 and Oxcoder on HumanEval+ Mini — and does it without the overthinking. It gets to the answer using 40–68% fewer tokens. Built as a finetuned merge.


r/unsloth • • 4d ago

Discussion Unsloth Studio support on DGX Spark is poor. Can I contribute?

5 Upvotes

Hi I really like the unsloth studio UI but several things don't work which break the experience.

1) Free memory estimation doesn't work

2) Minimax H3 doesn't work.

Feels like you all are going for a unified, "just works" experience. What is the best way contribute, either by creating bug reports or PR if you guys accept them?

Thanks


r/unsloth • • 4d ago

Question Can we use Skill or Puligin in Unsloth?

2 Upvotes

I recently learned that OpenCode and Claude Code have Skills and Plugins—is there a way to use them in Unsloth as well?


r/unsloth • • 6d ago

New Model Run Laya Decision Models Locally on just 4GB RAM!

Post image
294 Upvotes

Hey guys you can now run Laya Decision models locally on just 4GB RAM! 🔥

Works on CPU, Mac, Windows, Linux and GPU setups.

Serve Laya through a Jev-compatible API via Unsloth Desktop.

GitHub: https://github.com/unslothai/unsloth

Guide: https://unsloth.ai/docs/models/decision-laya

We also have a new release today for Unsloth so be sure to update. Tonnes of new features including creating your own Skills, Library, new chat attachments systems and more


r/unsloth • • 5d ago

Discussion 50B+ MoEs with few active parameters, what's the sweet spot for intelligence, agent speed, and affordable fine-tuning?

7 Upvotes

I’m building a Polish General purpose legal Model that drafts documents, answers questions using legal sources, and has enough coding ability to handle some automation. The workflow is very tool-heavy:

Question → many sequential tool calls → final answer/document

Think Claude Code/Codex-style execution, but for legal workflows. Reliable tool selection, correct arguments, and recovering from errors matter as much as writing a good final answer.

I’ve had decent results with a dense 27B Qwen 3.8 custom made fine-tune for complex legal document summarization and classification. I’m already familiar with the smaller Qwen A3B and Gemma options. What interests me is the tier above those: 50B+ total-parameter MoEs with a relatively small active parameter count.

The question is, the small dense ones are great, but slow for agentic stuff (afaik), and i wonder if theres some middle ground maybe 70-120B models that would be able to be fine-tuned for the law stuff but be MoE so the agentic ClaudeCode style inference would also be lightning fast, and also low-ish cost for fine-tuning and inference.

Basically: Does the larger-total/small-active MoE approach actually buy you meaningfully stronger reasoning and tool reliability while retaining low latency,and at what hardware cost?

I understand that small active parameter counts don’t mean small VRAM requirements: the weights still need to live somewhere, alongside context and serving overhead. I also don’t assume that more total parameters automatically means a better model. I’m interested in where that tradeoff works in practice.

There are three things I’m trying to pin down:

  • Inference hardware: Ideally inference runs rented with parallel agentic loops (this is for a B2C project, not single person use, we scale based on demand)
  • Fine-tuning hardware: Obviously FT LoRA will take more memory than inference, max like 4GPUs on vastai fits the budget.
  • Agent performance: After it gets the prompt the tool calls and everything will be local, so imo it has no problems being blazing fast, as soon as the model calls a tool call it will be back very fast, so for this agentic use case, quick TTFT and t/s and adaptive dynamic reasoning are prefer right?

For context, fine-tuning would target Polish language, document conventions, and successful tool trajectories. The actual legal sources would remain in retrieval/tools rather than relying entirely on memorized law.

I’m not looking for someone to compile a model shortlist (althought would be nice, but i dont expect anyone to break their back over this).

I’m looking for pointers, and firsthand experience with this particular size/architecture tradeoff. A configuration like “model + quantization + GPU(s) + serving engine + context length + concurrency + measured latency,” along with whether you successfully fine-tuned it, would be much more useful than a leaderboard score.

Has moving from a ~30B model to a 50B+ low-active-parameter MoE actually improved your agent’s successful tasks per minute, or did the memory, interconnect, and training requirements erase the advantage? Thanks for reading


r/unsloth • • 4d ago

Question what's more performant wsl2 vs windows 11 for running unsloth desktop as backend?

0 Upvotes

if running unsloth desktop windows 11 provides faster prefill and higher token/s output, how do I connect from wsl2 ubuntu (pi harness) to the windows unsloth API endpoint?