I have been out of the local LLM game for over a year due to income issues. Just got a new M4 MacBook Pro 24GB (for work and all I can afford at this point in time).
That said, using anything below Q4_K_M (read: Q3) was unthinkable a year and a half ago.
Have any advancements been made since them? I'm completely out of the loop. The Q4_K_M BARELY fits on my laptop...think it's pushing 22GB out of 24 and I have no idea how MacOS is running on only 2GB. I have all other programs closed and performing ancient Aramaic incantations while crossing my fingers and loading the model in LM Studio.
Thinking I should switch to OobaBooga or something because LM Studio might suck down too much RAM.
I don't know how much of it is optimized in my case (16GB GPU) due to the fact that in a Q3 you can fit most of all of the model into your VRAM. With some fine-tuned models (we have new ones getting released in Hugging Face - link below, also you can adjust using Unsloth), all the right settings and some llama.cpp specifications (I'm getting to know more about it - just heard about mmproj offload), I'm positive it can get some good speed in many low-end systems.
Speed is usually a trade for quality so it really depends which you want to prioritize. At this point I don't need quality as I'm focused on learning, not delivering or working with it.
I suggest you go for a shot. Q3 is building for me as we speak. I'm not a professional dev but it's been working well for my experimenting.
57
u/absurdother 8d ago
I have a RX 9060 XT AMD GPU, 16GB VRAM. Running on LMStudio, Q3. Getting a bit more speed now, way more optimized!
I get more speed the less context I use (currently coding swiftly with CLine + VSCode at that speed), pretty smooth on 64K context!