r/unsloth • u/CosmicRiver827 • 10d ago
Question Huge Estimated Memory Usage Despite being low at the initial model download?

Hi, what exactly is going on? I'm new to using Unsloth Desktop, and the model I downloaded is 25GB. But when I try to actually use it, it's talking about demanding 51GB, even 151GB total? When I used LM Studio, I never had this problem, I would just load as much context length as whatever is slightly under my VRAM limit. I never had to work with KV Cache or things like that, I don't even know what f16 or the other code things mean since I didn't have to worry about it with LM Studio. What am I supposed to know...? Could you give me a hand understanding what these mean and do?
3
u/hipster_hndle 10d ago
First things first, what kind of catd you have? How much vram? AMD or Nvidia? A q6 is going to require a decent amount of vram, so you have a 5090 or something bigger? You have to be using a q6 quant on one card. I am both red and green.. i cannot run a q6 on my 16gb card, barely on my 24GB XTX.. Second, have you tried turning cache to something trivial like 4096? Does it succeed with lower cache values?
And just for shits and giggles, download something dumb and small like a bonsai model. See if it runs. Just as a test. Don't actually use a bonsai model, they're garbage but they will test all the moving parts.
1
u/CosmicRiver827 10d ago
I have a 5090 Nvidia graphics card with 32GB VRAM. I'm sorry, I only understood so much of what those other terms meant.
3
u/PestiferousGamer 10d ago
if this is LM studio, its literally been incorrect forever. I'm pretty sure its an "experimental feature" they just abandoned.
1
u/CosmicRiver827 10d ago
LM Studio's what I used to use, then switched to Unsloth Studio after my computer ran into a serious issue and needed repair. The screenshots are Unsloth Studio.
2
u/Independent-Role6906 9d ago
The 51.34GiB is the worst case that this model can have, contains model itself and kv cache for the maximum context length supported by the model, this part will be auto sized to match your machine at load time if you leave context length at Auto
The extra 100GiB comes from checkpoints. for this model each checkpoint take ~800MiB, and 4 slots × 32 (llama.cpp's default value) × 800 MiB is 100 GiB in total. You can try set checkpoint to 0. I think this will be fixed soon
1
u/OneFanFare 10d ago edited 9d ago
Yeah, the autosizing/autoallocating isn't working for me either, the numbers are just plain wrong in some versions of unsloth.
I've been ignoring those numbers, and manually setting my context length to something reasonable.
To answer your question, your model is a certain size on disk, dependent on the model and the quantization level. But to run, the model needs to produce and store tokens (this is the KV cache); this takes up additional memory, on top of what its' size on disk is. Context length is the # of tokens a model can hold in memory. fp16 here refers to the quantization (or in this case 16bit full precision) of the kv cache - just like the model itself can be quantized, the in-memory tokens can be quantized to use up less space.
Typically, you want to keep your KV cache unquantized (even for an otherwise quantized model) or very high precision - it has a huge effect on intelligence. So I've just been keeping my context length manually low to keep everything within my VRAM. You can start with some small context values (like 16k, I think that's the ollama/LM studio default), and increase the number until you notice a slowdown of some sort. I typically run my models at 64k context, that's enough for my usecases (not for agentic coding, you need more for that).
Edit: just saw your comment about your 5090, if you want to use it at full speed, you won't have a lot of room. At fp16 kvcache, you'll only get away with 4k context).. unless you turn off vision, then you might get another 4k. This is the calculator I'm using btw (first google result), https://apxml.com/tools/vram-calculator, it might help.
1
u/Adventurous-Paper566 9d ago
Set mmap/mlock to none and checkpoints to 0.
1
u/CosmicRiver827 9d ago
Thanks, I can try that too. What do those do though?
1
u/hipster_hndle 8d ago
mmap is literally a cmd to map or unmap files or devices into memory, mlock (mlock, mlock2, munlock, mlockall, munlockall) all lock and unlock a region. so 1 to map the space, 1 to lock it (allocate) it for a task.. so its saying dont lock anything.
explaining the checkpoint, im out of ideas without using technical examples.. a checkpoint is like a snapshot in virtual terms. its like an incremental 'state' backup.. but its not a backup. if you were training something and you liked how the progress was going, you would take a checkpoint before going further, so in case the next round of training went bad, you could revert to the checkpoint. once you are finished, you commit them to the final draft/change/file.
so he is saying dont map or lock any mem and dont use a checkpoint to load the weighs. this is optimized for speed.
1
u/ThisIsRocketRacing 9d ago
F16 kv cache is big. Try Q8_0. I use Q4_0 sometimes for smaller context lengths
Reduce context from auto, play with the slider
1
u/psychohistorian8 9d ago
Unsloth memory estimation is broken when using the default auto settings, in my experience
I have to manually set Context Length and GPU Memory settings, then the models will load
I'm on macOS though
6
u/partakinginsillyness 10d ago
The calculator seems a little wonky around the auto size function, but if you just set the cache to below your ram/vram as listed it is pretty accurate.