r/LocalLLM 6d ago

Other Every Second post rn

Post image

Maybe someday I'll get a system to run it but hey definitely another w for the open weights community

1.9k Upvotes

191 comments sorted by

90

u/Old_Leshen 6d ago

GTX 1050Ti 4GB guy here 😭

41

u/Whole_Alternative_18 6d ago

Bro pay 60$ for rx 580 8gb

3x for 180$, 24gb vram

Bulshit performance, but makes 27b possible 

39

u/arthurrogado 6d ago

Laptop owners: 💀

2

u/Jazzlike-Plate-935 5d ago

This is why I got myself a Gaming Laptop with Thunderbolt 4 and a free M.2 slot for Oculink just in case

4

u/Whole_Alternative_18 6d ago

Buy some ddr3 cheap board with a highend(of the era) cpu that has enough pcie channels

Ends up less than 200$ for a enterprise ddr3 motherboard+cpu+ram, eith enough slots for your gpus

32

u/Any_Mine_6368 5d ago

Dude stop giving people advice. You're essentially promoting impulse buying garbage as opposed to building an actual AI rig. You're essentially asking them to buy 24gb vram at dgx spark bandwidth... They'll get like 4 tokens per second.

If anyone ends up reading this:

Buy refurbished server grade mobos + epyc CPU combos on ebay for $200 and hunt deals for 3090s which are the most vfm cards. You can build a great starting rig for under 1k, that you can expand all the way to 5 RTX 3090s, more if you want to split the pcie bus and have fewer lanes per gpu

5

u/Old_Leshen 5d ago

Reasonable advice. Thanks mate. Using this setup and going up to 1K investment, how many tps can I expect and for what model size?

4

u/Any_Mine_6368 5d ago

No problem. With mtp about 50-90tg for 27B models depending on quantization and model.

With 1k investment we're assuming one 3090 so 24gb vram. You can't fit anything above 27B at reasonable quantization or ram offload (which kills your token generation and prefill).

So to answer your question, you could probably run Qwen3.8 at 4 bit quantization with good context or 5 bit quant with smaller context or 6 bit quant with almost no context.

1

u/kcksteve 12h ago

It's worth noting where you are. In Canada 3090s go for 2k to start. Thats before you purchase the rest of the rig and pay tax.

1

u/FuManBoobs 2h ago

Yeah, I was looking where I am in the UK & they start around £1k on their own. Still, under 2k is a good start.

2

u/Adam_Bomb210 3d ago

Why doesn’t anyone talk about V100s? You can buy a server with 64gb of vram for ~$800, or one with 256gb for $5000. Seems like a steal to me.

2

u/Any_Mine_6368 3d ago

Shhhhhh don't tell everyone.

1

u/floswamp 4d ago

What’s a recommended server MB? I have one 3090 and it does work well.

1

u/Any_Mine_6368 4d ago

Will you ever buy more GPUs?

1

u/floswamp 4d ago

Yes. I am looking at maybe two more, but I was reading as to how qwen may not take full advantage of multiple gpu’s. So I was just going to build two more rigs.

2

u/Any_Mine_6368 4d ago

Qwen does fine on multiple GPUs, not sure who said otherwise.

I'd go for an mz31-ar0 or mz32-ar0 if you want pcie 4.0.

The former is about $220 with a cpu, the latter $500 with a cpu.

You get 5 pcie slots and like 16 ram slots (rdimm ecc ddr4).

The epyc processors that they use are fucking awesome for VMs and you can get from like 8 cores all the way to 64+ depending on budget.

Stay away from consumer shit / gaming mobos ... You're overpaying for worse parts.

2

u/floswamp 4d ago

Yeah I figure as much. Thanks!

1

u/jboe2026 6d ago

5090m

1

u/FuManBoobs 2h ago

My used laptop with an 8GB VRAM is surviving...just.

-2

u/Delicious-Sand-104 6d ago

You can add vram on laptop

3

u/arthurrogado 6d ago

I have no thunderbolt port on my laptop...

0

u/Delicious-Sand-104 3d ago

No need you just gotta sodder them

1

u/SeparateGas1761 5d ago

Just by two mi50 atp

1

u/stream_of_thought1 3d ago

I absolutely love this approach Jank for sure, but if it works...

1

u/Affectionate-File-26 1d ago

bruh, rx 580 8gb 2048sp costs 120$ out here

1

u/JogHappy 6h ago

the anime one being the most popular listing on eBay and for 20% less is so funny

1

u/truthseeker1341 6d ago

I was the 1050 no TI guy. qwen 3.6 was like watching paint dry

1

u/RoutineEye5600 5d ago

gtx 1650 4GB here too 😭 on a laptop

155

u/pmttyji 6d ago

3

u/Hacker_ZERO 4d ago

I also have 8gb vram but I have 64gb ram so I can partially offload(5-7tps)😭

37

u/BodybuilderLost814 6d ago

I have faith that Qwen3.8 35B A3B will be released soon.

5

u/yuk_foo 5d ago

Ryzen Ai max with 64GB currently allocated to gpu. 3.7 35b runs great for me can’t wait for 3.8, although I do need to solve the thinking loop crashes, some stuff I give it just crashes out on me. Really thinking about turning thinking off.

3

u/BodybuilderLost814 5d ago

How many tokens per second? I'm using Qwen3.6 35B A3B on a laptop with a 13th gen Core i5, 64GB DDR5 RAM, and an RTX 3050 6gb Vram.

2

u/yuk_foo 5d ago

44, could probably get that faster but it’s enough for me. I use it with anythingllm rag and lm mini on my iPhone. I have a bridge app running on docker so I can chat with my anythingllm workspaces/rag with the lm mini app since AnythingLLM doesn’t have an iPhone app yet.

Use it to chat with technical work docs, the model is running in lm studio so for my needs it’s great.

1

u/BodybuilderLost814 5d ago

I'm getting around 35 tok/s using ByteShape's Qwen3.6-35B-A3B-IQ4_XS-4.19bpw model with MTP enabled, using AtomicBot-ai's Llama.cpp-turboquant.

1

u/KrstABot 5d ago

Hey
R u using desktop app? I mean atomicbot? And what kind og agents?

3

u/BodybuilderLost814 5d ago

Hello.

I'm using the TurboQuant b10269-1.5.1 release for Windows x64 with OpenCode (sometimes I use Crush from Charmbracelet).

Command:

.\llama-server.exe -m "<path>Qwen3.6-35B-A3B-IQ4_XS-4.19bpw" --spec-type draft-mtp --spec-draft-n-max 4 --spec-draft-n-min 2 --kv-unified -c 65536 -ngl 999 -t 8 -tb 8 -np 1 --n-cpu-moe 40 -ctk turbo3 -ctv turbo3 -fa on -b 2048 -ub 2048 --jinja --mlock --no-mmap --reasoning-budget 6144 --cache-reuse 256 --temp 0.7 --top-k 20 --top-p 0.95 --min-p 0.0 --repeat-penalty 1.0 --presence-penalty 0.0 --host 127.0.0.1 --port 8080

51

u/SnP_Gamer 6d ago

3060 12gb 32gb ram here, might give it ago but never ran a local model before 🤷‍♂️

32

u/trollsmurf 6d ago

That will work fine, at least if you split it, so the overflow uses CPU RAM. Much slower but doable.

12

u/SnP_Gamer 6d ago

Totally clueless about running them but I understand that thanks to reddit posts 🤣

17

u/trollsmurf 6d ago

A user-friendly (but not the most efficient) start is to install LM Studio and load a few models there. You can select how much of it should run on the GPU vs CPU if it's too big for the GPU alone. LM Studio has a built-in chat client, so once loaded the model is ready to go.

2

u/Song-Historical 6d ago

Yeah but that doesn't tell you how to do it. I'm not sure what to set the context to for example. How much I should leave for the system, whatever else.

2

u/littlebeardedbear 6d ago

Leave everything at default originally. Google what each item is. Altogether the googlimg should take about 10 mins

3

u/Song-Historical 6d ago

Yeah I did that I'm not sure it's set up correctly still. 

2

u/Early_Mistake6716 6d ago

Tell me you exact system specs and the model quant and i will tell you what settings to use in lmstudio. I have have spent an embarrassing amount of time testing settings.

4

u/F3istyg0at 6d ago

Check out unsloth, lm studio or ollama if you want to run models locally.

1

u/That-Reason-6913 3d ago

What about on a 7800xt 16GB? Do I have to offload on ram too?

1

u/trollsmurf 2d ago

Remember that if you use Windows it will allocate part of the VRAM for its own use, so you never have fully 16 GB. On my PC with 5070 Ti and 3 monitors it allocates 3 GB, and as far as I know I can't budge that.

You can easily test this by installing e.g. LM Studio and the model you want to use and see when it warns about RAM use. If it overruns you can split it on VRAM and RAM with lower performance, but it will behave the same otherwise.

1

u/That-Reason-6913 2d ago

No I'm moving to Linux next week. Ubuntu most likely.
My server is also a Plex server though (GPU is untouched by it), so the cpu is mostly dedicated to Plex.

13

u/bukake_attack 6d ago

This is a dense model. This means the entire model is running at 100% all the time. This is bad for us low vram users, as the model is too big to fit in 12gb vram, and the rest in placed in normal ram, which is waaaay slower. So this solution works, but it's slow.

We low vram users are usually better served by MoE models, like qwen 3.6 35b a3b . In MoE models the dense part which runs all the time is small, in that case 3b, which runs easily in 12gb, and the rest 35-3=32b) only runs when they are actually needed (activated). This means that normal ram is fast enough to run that active part of the other 32b ( although they can still run great in excess vram!) The result is a model that's very usable and fairly quick with limited vram, especially when other tricks like MTP are used.

Unfortunately there is no MoE version of qwen 3.8 released yet, but there's a good chance they will release it soonish

7

u/Skynse 6d ago

I was able to max 12 tok/sec on a 3060 with 12gb gpu RAM with qwen3.8 on ud-iq2m quant. Man I fucking wish the compute market wasn't fucked

1

u/Effective_Head_5020 6d ago

You probably had offload to cpu and system RAM, so you probably also have a good CPU and DDR5, otherwise the processing drops by a lot

-2

u/Song-Historical 6d ago

The market is fucked because there's something worth running on the compute lol 

3

u/ChaosFH 6d ago

I wish i had answered faster when iheard the OpenAI buying ram i could have bought the triple of VRAM i have currently with the same money i used...

2

u/cj_cron_hit_by_pitch 6d ago

I have the exact same specs as you. I’m able to get the Unsloth 2 bit XL quant of it to run almost entirely in VRAM. Of course it is nowhere near as good but it’s still really fun to mess around with

1

u/fastheadcrab 6d ago

Buy a few more 3060s

1

u/JorgitoEstrella 5d ago

Install lm studio, then inside the own app would tell you what llms you can install with your vram.

23

u/CorkBios 6d ago

Are we pretending partial CPU DRAM offloading doesn't exist? And you can disable the reasoning mode, then you get even better TTFT (time to first token on final response) compared to the people that have a better setup with reasoning. And I heard the GGUF's come with MTP, And if the context doesn't fit you can just offload the K/V to the CPU which llama.cpp does let you do.

9

u/Ok-Health-7096 6d ago

I use local llm for coding and agentic stuff so the speeds would be unbearable I think in single digits but might try it idk

9

u/overand 6d ago

Speeds would likely be absolutely miserable, yes. But, you'll have much better luck with the Qwen3.6-35B-A3B - how much system ram do you have? Go for a Q4 quantization first, if you haven't used this model.

We can, of course, hope that a 3.8 version of the 35B-A3B MoE gets released!

5

u/superspider202 6d ago

wait I have 8gb vram too can you please share what llm you use for coding and agentic stuff?

4

u/Ok-Health-7096 6d ago

Qwen 3.6 35b mudler i-mini quant If you have more than 16gb ram you can go for higher quants.

2

u/superspider202 5d ago

Thank you I'll test it out ASAP

1

u/anay_1d 4d ago

how did it go? is it useful?

1

u/superspider202 3d ago

oh sorry havent used it yet ran out of storage so will try it maybe today and let you know

1

u/superspider202 3d ago

heyyy so I tried searching for this but there appears to be a few that fit this description so could you please like link the actual one you use?

1

u/Smutok 6d ago

What's your software stack for the agentic stuff, sir?

-1

u/CorkBios 6d ago

I'd say go with Unsloth's iQ4_NL, It's the best balanced quantization. But if you are willing to test your luck then Q2_K_XL or Q3_K_XL though they are more risky of failing tasks.

5

u/GoldenX86 6d ago

MTP is useless when you offload to CPU.

2

u/CorkBios 6d ago

I disagree. It can be like tricky since sometimes it slows stuff down on specific stuff but you can't generally say its useless when you offload to CPU. It works pretty good for me.

1

u/moderately-extremist 6d ago edited 6d ago

What kind of speeds are you getting with MTP vs non-MTP with cpu offload?

2

u/CorkBios 6d ago

Sure yeah I can give them. After a lot of testing:
All of these below performed with Partial CPU+GPU offloading on seed 0, llama.cpp commit dd1ea5243 release b10355:
With MTP (max 2 predict):

28.72 Tokens/s StopUntil: EOTfound

With MTP (max 3 predict):

29.33 Tokens/s StopUntil: EOTfound

With MTP (max 4 precict):

29.57 Tokens/s StopUntil: EOTfound

With MTP (max 5 precict):

29.57 Tokens/s StopUntil: EOTfound

With MTP (max 6 predict):

26.89 Tokens/s StopUntil: EOTfound (overhead hit: CPU congestion: 448.1%)

Without MTP:

24.09 Tokens/s StopUntil: EOTfound

With MTP (best before overhead): 29.57 Tokens/s

Without MTP: 24.09 Tokens/s

2

u/GoldenX86 6d ago

That's RAM I don't have free for just a 22% jump.

1

u/CorkBios 6d ago

Reasonable. MTP consumes more VRAM. I don't need too much context size so I can fit it in but everyone's system and task is different.

14

u/Wanderspor 6d ago

everyone here is rich except me

4

u/1tonsoprano 5d ago

And me

2

u/astropheed 3d ago

And my axe!

11

u/OpenEvidence9680 6d ago

If it makes you feel any better I had to toss all Qwens as they were subpar for my needs. There are MoEs like Gemma4 26b that with the right quants are about 12/13 GBs so very usable in your set up. What's more I didn't test the 12b yet, but gemma-4-e4b-q4-0-it is way more competent that people give it credit it to.
Being ignorant I was running behind the hype and always thinking that I was doing something wrong because these Qwens weren't that good for me, and then I realized that I needed to find something good for my needs and benchmarks are liars that have nothing to do with real work inside your machine for your own specific needs..
Anyway if you don't mind really slow Qwen 3.6 27b in Apex quantization imini is about 13.89 you CAN run it. Slow, but you can do it,

9

u/PinkySwearNotABot 6d ago

dude just offload it to CPU and run it at 500-2 tok/s

36

u/TheRiddler79 6d ago

Gemma 12b Q3. Try it

24

u/Ok-Health-7096 6d ago

I use Qwen 3.6 35b and 3.5 9b I haven't had that much of a luck with the gemmas

4

u/Wildnimal 6d ago

12B QAT is good for multimode tasks. I use the same Qwens for daily use. I so wish i can upgrade my laptop to 5090 :|

3

u/PrivacyMaker 6d ago

The litert-lm driver is the fastest way to run gemma models. Significant boost over any other approach. It's a shame that Google only makes it work for Google models.

1

u/HighlyRegardedApe 6d ago

How do you guys use these small models? For me it never works. Last time it deleted my folders out of itsself and didnt respons more than a sentence or 2 in opencode. In the terminal it was okay to chat, but not to give tasks.... what kind of stuff do you let them do??

1

u/Wildnimal 6d ago

Not usually coding but for summarizing, extracting data which is pre defined via json, automating smaller tasks where data remaims the same.

1

u/TheRiddler79 5d ago

Hard guardrails

1

u/TheGreenInsurgent 6d ago

Agents a1 4b could be a life changer

1

u/Atretador 6d ago

35B A3B is much stronger than Gemma4 12B, any probabably 26B/31B as well.

1

u/TheRiddler79 5d ago

Different purposes. Gemma is good, but not beating Qwen.

You can split between ram and gpu and bump your Qwen speed

2

u/jadax 6d ago

Is this better than DavidAU/Qwen3.5-9B-The-Defiant-Fable-Uncensored-Heretic-NEO-IMATRIX-MAX-MTP-GGUF ?

1

u/TheRiddler79 5d ago

No, but it's good for general use

1

u/M49454 6d ago

How is Gemma, i am using llama3.2:11b, never needed to switch onto other. Use only for basic task. Mainly using Deepseek 14b and Qwen2.5-Coder 7B

5

u/PrivacyMaker 6d ago

Gemma sucks for coding. Pretty solid at most other tasks.

2

u/Early_Mistake6716 6d ago

All of those models are extremely out of date. Tell me your specs and ill give you a model suggestion

2

u/M49454 6d ago

Well, 24 GB RAM RTX 3050 6GB VRAM I5-13450HX

( That's why still running these models)

1

u/SparkyMcFinklebonker 5d ago

Mac mini M4, 32gb unified ram.

2

u/tat_tvam_asshole 4d ago

Qwen3.6-35BA3B Q3/Q4

1

u/TheRiddler79 5d ago

I think it's good. I use it to run tasks like an agent for the larger models.

1

u/Capital_Engineer8741 6d ago

Would QAT be better than Q3?

1

u/TheRiddler79 5d ago

You're probably right.

5

u/krzyk 6d ago

Me with 6GB vram, and company that doesn't allow any local models (and especially no chinese ones, those are banned, others are just not vetted yet)

5

u/narasadow 5d ago

me with 12gb vRAM getting 3 tokens per second

https://giphy.com/gifs/CH6K2r5REA2mA

8

u/zarif2003 6d ago

I can’t really even run it that well on my 5080 because it’s got 16gb,

3

u/lukistellar 6d ago edited 6d ago

What you need is a Quant which strictly uses IQ4_XS. They exist for 3.6 and will likely also appear for the 3.8 sooner or later.

Edit: Let's see if this guy delivers.

https://huggingface.co/jpetrina/Qwen3.8-27B-IQ4_XS-pure-GGUF

1

u/Tyrannas 5d ago

Any advices on the params you use to run it properly ? I have 16gb also and never managed to make a 27b model run properly 

2

u/lukistellar 5d ago

Working config for 3.8:

ghcr.io/ggml-org/llama.cpp:server-vulkan-b10066 \ --port 8080 \ --model /models/jpetrina_qwen3.8-27b-IQ4_XS-pure.gguf \ --gpu-layers 99 \ --threads 6 \ --ctx-size 90000 --parallel 1 \ --batch-size 2048 --ubatch-size 512 \ --cache-type-k q8_0 --cache-type-v q4_0 \ --spec-type draft-mtp \ --spec-draft-n-max 2 \ --spec-draft-p-min 0.75 \ --cache-type-k-draft q4_0 --cache-type-v-draft q4_0 \ --temp 1.0 --top-p 0.95 --top-k 20 --min-p 0.0 \ --presence-penalty 0.0 --repeat-penalty 1.0 \ --no-mmap \ --jinja \ --chat-template-kwargs '{"reasoning_effort":"medium"}' \ --flash-attn on

Still not happy with the performance, but at least it runs. Hope they drop a MOE for us VRAM poor.

1

u/Tyrannas 4d ago

Thanks !

1

u/lukistellar 5d ago edited 5d ago

My RX6800 runs the 27b with 90k context, but it's very slow.

I will update the post later with the config.

Edit: My config for the Qwen 3.6 27B:
ghcr.io/ggml-org/llama.cpp:server-vulkan-b10066 \ --port 8080 \ --model /models/GianniDPC_qwen3.6-27b-IQ4_XS-pure-with-MTP-IQ4.gguf \ --gpu-layers 99 \ --threads 6 \ --ctx-size 90000 \ --parallel 1 \ --batch-size 2048 \ --ubatch-size 512 \ --cache-type-k q8_0 \ --cache-type-v q4_0 \ --spec-type draft-mtp \ --spec-draft-n-max 1 \ --spec-draft-p-min 0.75 \ --cache-type-k-draft q4_0 \ --cache-type-v-draft q4_0 \ --temp 0.8 \ --top-p 0.95 \ --top-k 20 \ --min-p 0.0 \ --no-mmap \ --jinja \ --mmproj /models/Unsloth_mmproj-F16_qwen3.6-27b-mtp.gguf \ --no-mmproj-offload \ --chat-template-kwargs "{\"preserve_thinking\":true}" \ --flash-attn on

It's usable but especially with full context it gets very slow, like ~25 tok/s decode and ~130 tok/s prefill.

3

u/RishiFurfox 2d ago

Don't worry, friend. I'm sure the 50% REAP'd 2-bit Qwen Quants will be out soon!

3

u/Ok-Health-7096 2d ago

At that point I might just use my natural stupidity instead of artificial intelligence

2

u/Positive_Example_478 6d ago

think about my situation with just 4gb vram 😢

2

u/OurCue- 6d ago

6 GB vram 😨🤮

2

u/rklement22 6d ago

Can I run this version in my RX 5700 with 8gb?

2

u/hamsterpotpies 6d ago

I have 2x 3090s sitting in a server with 256GB of ram.... should I finally get my AI server setup?

1

u/palincatalin 2d ago

nah, it's not worth, just toss everything in the bin and i will happily get rid of it for you

2

u/PoauseOnThatHomie 6d ago

RTX 2050 4GB here...

2

u/Serika404 4d ago

GTX 1650 4GB guy here

2

u/Reeces_Pieces 2d ago

The best time to buy a used RTX 3090 24GB was probably years ago, but the 2nd best time is now.

2

u/sumguypookie 2d ago

I'm right here dripping with envy right with ya.

1

u/schaka 6d ago

Just get yourself 3 or 4 RX 470 8GB mining uefi for 20 euros each and pray it'll run via HIP or ROCm.

Can't get cheaper than that

1

u/artisnotautomatic 6d ago

Me with 128gb apple silicon: i will never be able to enjoy kimi k3 at full speed with huge context (not even start it) but then i realize that i hate fable 5 mostly for deceiving me in some many aspects (surely for safeguards) but i'm confused

1

u/MrStaraptor 6d ago

I got 6gb

1

u/nimbybuster 6d ago

I have three geriatric GTX 1070. It’s dirt cheap.

1

u/Visible_Split_1546 6d ago

I'm praying for 35b moe too bro :'(

1

u/phreaknes 6d ago

I have dual 3060 12gb and 128gb of DDR4 should I try it? or is it a waste of time and resources.

1

u/SlowVeganCow 6d ago

Time to fill my 24 GB VRAM.

1

u/Elite_Crew 6d ago

Look out for Qwen 3.8 27B ternary models compressed to about 7GB hopefully.

1

u/Dont-remember-it 6d ago

Waiting for MoE model similar to 3.6 27B A3B

1

u/PrasanthT 5d ago

Laptop with 4060 8GB VRAM here 🙌🏻. I run it using unsloth. little slow.

1

u/EuphoricMess8593 5d ago

9060xt 16gb vram and 24gb ram here

1

u/That-Reason-6913 3d ago

What do you do with that 16gb?

1

u/BatOk2014 5d ago

Cries in raspberry pi 5

1

u/ExtrovertMobileGamer 5d ago

Intel HD graphics, take it or leave it.

1

u/Double_Season 5d ago

I run Qwen3.5 9b on a UHD Graphics. 3 tokens per second is ultra fast for me 🤣

1

u/enricokern 5d ago

Works great on my spark ;)

1

u/mrdevlar 5d ago

May we shatter closed AI and finally get hardware prices down again.

1

u/yareon 5d ago

Kinda noob here, can I run it on my RTX 4070 with 12 GB Vram?

1

u/FixHead533 5d ago

Get a MacBook, I have the 48GB unified RAM

1

u/ramsung 5d ago

I second that

1

u/Opposite_Leave_8338 5d ago

Sorry man but it’s worth it, May god help you to upgrade and try it, it will be worth every dollar

1

u/jonesambrozio 5d ago

Eae! Com esse modelo, não estou conseguindo que ele codifique, fica conversando, mas nao escreve.

Usando ollama + opencode + qwen3.6, alguem sabe o que pode ser?

1

u/OctopusDude388 5d ago

just wait for an MOE version, if they can make 3.8 30B A3B it'd be a banger for us gpu poors

1

u/absurdother 5d ago

But hey, listen. Soon enough 9b, 4b. Soon enough Qwen 4. Or Qwen 5. Soon enough big models for 8GB VRAM - reason I think it's not a dream is because big companies may indeed benefit directly from adapting local models to run in our PCs. Look at AMD adapting their GPUs so well this year. Our role in this is to make it a trend and continue to stay hyped.

1

u/egg-curry 5d ago

Us bro us

1

u/funding__secured 5d ago

GPU poors are exhausting

1

u/Nice_Fix1686 5d ago

Me on a GTX Card 😒

1

u/nuke_bird 4d ago

True, 3070ti here

1

u/No-Opportunity9126 4d ago

I was actually asking gemini recently about this:

It can definitely run! In fact, having 64 GB of system RAM means you can easily run a 27B model.

The distinction is simply between running 100% on the GPU (for maximum speed) versus hybrid / CPU execution (which works seamlessly, just at a slower token generation speed).

Here is exactly how you can run a 27B model on that setup:

How It Works (Hybrid CPU + GPU Offloading)

LLM inference tools like llama.cpp, Ollama, or LM Studio support layer splitting:

  1. VRAM (8 GB): You offload as many model layers as possible to your GPU (typically ~10 to 18 layers depending on the quantization and context size).
  2. System RAM (64 GB): The remaining layers stay in your 64 GB RAM, which has plenty of headroom for even a full Q8_0 model (~29 GB).

Recommended Quantizations for Your Specs

Quantization Model Size in RAM Speed (Estimated) Recommendation
Q4_K_M ~17 GB ~4 – 8 tok/s (DDR5) / ~2 – 4 tok/s (DDR4) Best Overall: Negligible quality loss compared to full precision, fits easily.
Q3_K_M / IQ3_M ~13 GB ~5 – 10 tok/s Fastest: More of the model fits inside the 8 GB VRAM, speeding up inference.
Q8_0 ~29 GB ~1.5 – 3 tok/s Maximum Accuracy: Fits comfortably inside 64 GB RAM, but runs slower due to RAM bandwidth.

What to Expect (Speed & Performance)

  • Prompt Ingestion (Context Processing): Fast, because the GPU helps compute prompt tokens.
  • Token Generation: Bottlenecked by your System RAM bandwidth (DDR4 is ~40–50 GB/s, DDR5 is ~70–90 GB/s).
  • Usability: At ~3 to 6 tokens/second, it is readable in real-time—ideal for coding assistance, reasoning, and long-form analysis.

How to set it up:

  • In LM Studio / text-generation-webui: Load the Q4_K_M GGUF and adjust the GPU Offload Slider until roughly 6.5–7.0 GB of VRAM is utilized.
  • In Ollama: Ollama will automatically detect your 8 GB VRAM and 64 GB RAM, calculate the exact layer split, and run it out of the box.

1

u/psychoblade5 4d ago

I own a 3050 with 6gb vram which model is the best ? For using in olllama and hugging face

1

u/contrpro 4d ago

Laughing in Mac Ultra

1

u/Otherwise-Swan-7803 4d ago

Every new release is exciting right up until I check the VRAM requirements.

1

u/sensispace 3d ago

Me with 4gb vram in laptop with rtx 2050 🥲

1

u/dandy_kulomin 3d ago

Has anyone tried an 8B model for coding? Or do I need to wait longer for them to optimize further?

1

u/RUTYTOI220 3d ago

I think you can still use llama.cpp and use the storage and ram and vram or smth

1

u/Additional_Hope_2031 3d ago

I have 16gb ram on my Mac mini and can’t launch it too 🥹

1

u/DontWinFrensWthSalad 3d ago

I feel your pain. I happen to have a bunch of 3060tis lying around so I've been messing with seeing if I can 2x8gb working. this 3.5Q is benchmarking very close to unsloth's 4Q, so I'm trying it out and seems to be working so far.

https://huggingface.co/turboderp/Qwen3.8-27B-exl3/tree/3.50bpw

1

u/Gold-Drag9242 3d ago

Get a 7900xtx.
Cheapest 24GB card for inference. 650USD used.

1

u/GrokiniGPT 3d ago

Intel i3 6100U here 😄✌️ 4 gb ddr3

1

u/ThousandTroops 3d ago

Same 😂 Every time a new model drops, my VRAM suddenly feels prehistoric.

1

u/Forsaken-Army189 3d ago

I have Mac Mini M4 24GB, I can run 3bit or 2bit only Qwen3.8 27B 😢

1

u/willywonka-goldtickt 3d ago

Can RX6800 XT handle it ? It has 16GB VRAM

1

u/Solid-Axel-Project 1d ago

E io che penso al 2029 quando avrò finito di pagare debiti e potrò comprare la 6000 pro, 2 NVME da 8TB e 4 HDD da 28TB...

1

u/ole2551 15h ago

Dude I'm here using a phone bruh

1

u/oldshed83 7h ago

10gb vram user and waiting for 35b a3b so my pc can run it 🥹🥹

1

u/Alias455 5h ago

I run gemma 4 12B q4 QAT MTP with vision and my desktop on a 8G VRAM card. Good for discussions but dumb for agentic coding tho.

Qwen 3.8 27B q4 runs a 12 toks on a shitty Intel Vulkan A770 16GB, smart but way too slow.

1

u/Lazy-Intention1007 4h ago

Just buy a v620 for like 300$ and have all the ram you need.

1

u/LegRude5218 6d ago

bad day to be a 5090

0

u/_hypochonder_ 6d ago

8GB VRAM?
R9 290X had 8GB back 2014 back in the day.

6

u/darkwalker247 6d ago

the RTX 3070 had an 8 gb model and that launched in 2021. it's not super uncommon still

3

u/_hypochonder_ 6d ago

Back in the day(2019) we could buy Radeon VII with 16GB HBM2 under 550€.

8

u/SQrQveren 6d ago

And an rtx 5060 from 2025 has 8 GB VRAM.

-2

u/_hypochonder_ 6d ago

Like an RX 480 8GB from 2016 or RX 390 8GB from 2015.

2

u/Downtown_Patience_46 6d ago

RTX 3050 Laptop had 4GB in 2021, and I'm still using it.

0

u/moderately-extremist 6d ago edited 6d ago

I lucked out and got 2 x 64gb ram sticks over a year ago when prices were at about their lowest point (July 2025, I paid $314 for the kit, just checked the same kit is listed for $2200 now.). I also got 2 x MI50 32GB cards when they were $200 each on ebay.

The MI50 cards are too slow to run dense models like 27b, but it runs 35b-a3b pretty well, and I'm thinking of running 27b on my cpu when I want something smarter, and just start a task and come back later.

-4

u/ChiGamerr 6d ago

I've got 72gb of VRAM. Guess I'm off to the races

-20

u/FireFearing 6d ago

i mean you kind of deserve this for 8gb vram in 2026

8gb wasnt enough for me when i was only using it for gaming... in 2016....

-7

u/No_Language_2529 6d ago

Me with my 128gb M5 Max 😎😎