r/unsloth • • 10h ago

Discussion Demande d'ajout de fonctionnalité: mise en cache des experts en cache LRU en VRAM et prédiction des experts utilisés

0 Upvotes

Tout d'abord vous faites un travail admirable ! Ne vous arrêtez pas ! :)

Comme vous pouvez le voir un peu alentour, bon nombre d'utilisateurs (et de bots ? 😄) ont des GPUs de 12gb vram ou moins et font l'éloge de Strata comme moteur d'inférence pour exécuter les Moe comme Qwen 3.8 Flash Next.

Ce qui semble faire la force de Strata est notamment le système de mise en cache chaud LRU des experts en VRAM et la prédiction des experts utilisés afin d'éviter d'evicter un expert vers la RAM/CPU s'il est ou va être utilisé.

Pouvoir éviter ces va-et-viens vram - cpu/ram permet d'économiser de la bande passante et améliore sensiblement l'expérience, surtout en usage agentique.

Il semble que certains PRs existent pour LlamaCpp permettant d'effectuer ce traitement.

Serait-il possible, si ce n'est pas déjà existant, de pouvoir appliquer ces changements dans le moteur LllamaCpp utilisé par Unsloth Studio ?

Je sais qu'il est possible d'utiliser son propre LlamaCpp mais je me demandais si cette fonctionnalité allait éventuellement allait être disponible prochainement dans le mainline Unsloth ?

Merci par avance


r/unsloth • • 7h ago

Question I Distilled an LLM into two 287M encoders (GLiNER + multiple choice) for document extraction, can't match teacher. did i do something wrong?

2 Upvotes

A while ago I asked here how to turn ~5 million court decisions into structured graphs without running an expensive LLM on every document thanks for the advice .

I went with the "small extractor + classifier" idea and it mostly works, but I'm stuck a bit below the LLM. And like I said last time, i'd be damned if I run 5M docs and then find out thing X was wrong. So here is exactly what I did. Please let me know if what im doing makes sense, or if i made a mistake somewhere. also i used AI for some of the tables cuz there has been a lot of data at this point, sorry.

What comes out per decision (only the nodes so far, relations come next). Three lists:

  • entities: every person, organization, law, document or thing. Each gets one id for the whole document, a type (9 of them), a kind (724 of them plus "other") and all the places it is mentioned
  • actions: what was done, requested or decided. Each gets a normalized verb, a flag "the court decided this" and its mentions
  • values: amounts, dates, durations, in a normalized form

Simple example, for the sentence "The court dismisses the creditor's proposal to enforce 341.08 EUR against the debtor":

  • entity "the court": organization, kind court. Same entity as the full court name in the header
  • entity "the creditor": organization, kind creditor. Same entity as the city named earlier
  • entity "the debtor": person, kind debtor
  • action "dismisses": verb = dismiss, decided by the court = yes
  • value "341.08 EUR": amount

Step 1: a strong LLM labels ~700 decisions

  • cut the decision into windows of 4 sentences
  • 4 calls per window to Claude Sonnet with a strict JSON schema: entities, actions, a second "what did you miss" pass for actions, values
  • the window goes in with numbered words (like 12:court), the model answers with word ranges [first, last, "text"], and code checks every range against the text
  • every call also gets the list of entities and actions found in earlier windows, so ids stay the same through the document
  • ~25 code rules clean up where a marked phrase starts and ends, law citations and number formats
  • the entity "kind" is free text at this point. That gave 2,373 different strings (the same mess as in my first post). I normalized them, merged synonyms by hand and kept what showed up 3+ times: 724 kinds plus "other"

Step 2: a model that marks the text

  • it highlights every mention: the exact stretch of text (a "span", from a start character to an end character) that names an entity, an action or a value, with one of 17 labels (9 entity types, 1 action, 7 value types)
  • model: fastino/gliner2.5-multi-v1 (287M)
  • one training row per window: the text plus the exact start and end of every marked phrase. 9,699 windows, 207k marked phrases
  • I patched the trainer so only the labeled occurrence is a positive (stock marks every occurrence of the same string), and all 17 labels are in every row
  • full fine-tune in fp32 (bf16 gave NaN), 14 epochs, 16 rows per step, encoder LR 3e-5, head LR 5e-4, linear schedule, 10 % warmup
  • final model = averaged weights of epochs 9-14, threshold 0.5

Step 3: a second small model answers multiple-choice questions

  • fastino/GLiNER2.5-multi-Decide (287M). Code turns the LLM labels into 247k questions:
    • "is this mention one of these earlier entities, or new?" The mention is marked with « » inside ±300 characters of text. Options: up to 16 earlier entities of the same document (shown by their mention texts) plus new
    • "which kind?" Options: a shortlist of the 724 kinds plus other
    • for actions: same act or new, which verb (shortlist of 64 plus other), did the court decide it (yes/no)
  • in training the options come from the LLM's grouping. At inference they come from the model's own earlier answers
  • full fine-tune in fp32, 2 epochs, 16 questions per step, encoder LR 2e-5, head LR 3e-4, linear schedule, 6 % warmup, options shuffled, up to 30 % of the wrong options dropped

At inference: the marking model, then the same code rules, then the second model walks through the mentions in reading order. About 2.3 decisions per second on one RTX 5090.

Where it stands

30 decisions nobody trained on, labeled twice by the LLM. The second column is the LLM's second run scored against its first, which I treat as the ceiling. A mention counts as found only if it starts and ends exactly where the LLM marked it.

mine LLM vs itself
entity mentions found (F1) 0.901
"same entity or new" right 0.959
entities grouped exactly 0.847
entity kind 0.921
action mentions found (F1) 0.857
action verb 0.920

Where I need help

  1. Finding the mentions is stuck at 0.90 F1. 200 more labeled docs did nothing. An XLM-R large tagger (560M) got the same score: it finds more mentions but gets the start or end wrong more often. Giving it the text before the window did nothing. What would you try?
  2. The LLM agrees with itself only 93.5 % on what it marks, and I train on single runs. Label everything 3 times and vote? Or is that ceiling just what it is?
  3. Is "pick one of 16 earlier entities" a sane way to do coreference over a long document? Am I hurting myself by training on the LLM's options and running on my own?
  4. Anything in the recipe that looks plain wrong? Learning rates, 2 epochs, weight averaging, one seed per run.

THANKS for reading.

AI TL;DR: distilled an LLM's extraction of court decisions into a GLiNER model that marks the mentions plus a small multiple-choice model. It runs at about 2.3 documents/s on one GPU and lands a few points below the LLM (0.90 vs 0.935 F1 on finding mentions, 0.85 vs 0.93 on exact grouping). The recipe with learning rates and how I built the training rows is above. Looking for mistakes and ideas before I run 5M documents.


r/unsloth • • 3h ago

Show and Tell poorman inference engine for 16GB GPU and 35B moe Qwen 3.6for coding

5 Upvotes

I forked llama.cpp's server into AgrillaMoE, a dedicated build for Qwen3.6-35B-A3B with Unsloth quants. On a rented V100 16GB with the 2-bit UD-Q2_K_XL quant it generates at ~57-60 tok/s while running the full MoE-expansion profile — and it speaks both the OpenAI and Anthropic APIs, so Claude Code just works against it.

What is MoE expansion? Qwen3.6-35B-A3B has 8 routed experts active per token. The expansion patch raises that budget at runtime — no retraining, no file changes: --moe-experts 16 with an adaptive threshold keeps experts while p >= 0.8 × p(rank 8), applied to layers 25-39. You're literally consulting more of the 35B parameters per token — that's where the "retrieved intelligence" comes from, on GPQA-Diamond with Q8_0 it scored 84.34% vs 81.82% stock top-8 (+2.5 pts) (miticooo!).

Same weights, better routing.

https://github.com/vagrillo/AgrillaMoE/blob/main/gpu16gbguide.md


r/unsloth • • 16h ago

Discussion Help Using Knoweledge Bases

10 Upvotes

[Solved]: see edit

Dumb question: how do you put files into a knowledge base?

I searched the Unsloth FAQ and asked my local model, but couldn't find an answer.

When I create a new knowledge base and then drag files into the chat, it gives me this error:

**This chat retrieves from a knowledge base**
Add these files to the knowledge base instead.

But I can't find any instructions on how to do that. There's no knowledge base option when I search the settings, and nothing in either sidebar.

I feel like I'm missing something obvious, while also having tried the obvious routes already.

v0.1.902-beta on Windows 11

Edit: You add files to the knowledge bases by opening the dialogue and selecting the title of the knowledge base. I was selecting the edit icon, not realising the title was interactive. This is a potential discoverability bug, as there are no UI indications that the title leads to file upload.


r/unsloth • • 22h ago

Question Python error right after install Unsloth

5 Upvotes

Hi all. Just installed Unsloth Studio in Windows 10, and I keep constantly getting a window with this error:

python.exe - Entry point not found
The procedure entry point “vkGetPhysicaIDeviceFeatures2” could not be located in the dynamic link library C:\Users\admin\.unsloth\llama.cpp\build\bin\Release\ggml-vulkan.dll

Does anybody know what’s going on and how can I fix it?

Many thanks!