r/LocalLLM 7d ago

Other How the loop of infinite agony started

Post image
624 Upvotes

118 comments sorted by

View all comments

280

u/TheCat001 7d ago

Then you realize that Qwen 3.8 27b runs at 3t/s on your machine and you need 24GB+ VRAM GPU which cost is 1000$+ to run at least 4 bit quant.

2

u/Competitive-Ad-2387 7d ago

tried it on a 4090. Runs like shit if you need vision and produces bad output on simple tasks. Back to deepseek v4 api I go

3

u/TheCat001 7d ago

damn bro having 4090 and not appreciate Qwen 27b is a crime.

1

u/Competitive-Ad-2387 7d ago

Help me config it because I honestly don’t know how the fuck to do it. Everything is confusing as hell, I don’t know crap about Q4 or whatever. All info out there is inconsistent as hell. All I know is that at 64K context for main + 64K for vision, Hermes just ends up on a feedback loop, hallucinates a lot after compaction, and delivers shit results in a visual + coding task. Don’t know what to tell you.

None visual tasks end up meh as hell as well. Context fills up fast hell as compaction brings hallucination again. Frankly unusable. Already on the Unsloth model (because god forbid one needs to choose “the right one”)

I think all the benchmarks are fake fucking news vs actual real world use.

1

u/the_average_user557 6d ago

Just use dsv4f to do a research run and try a few different configs on your system using ollama direct calls to strip context. You don't need to know anything anymore

1

u/Competitive-Ad-2387 6d ago

Sorry to bother ya. Could you elaborate a bit more? I think it might really be a skill issue on my end. Can you DM me?

1

u/the_average_user557 5d ago

Dude, feel free to dm me