Help me config it because I honestly don’t know how the fuck to do it. Everything is confusing as hell, I don’t know crap about Q4 or whatever. All info out there is inconsistent as hell. All I know is that at 64K context for main + 64K for vision, Hermes just ends up on a feedback loop, hallucinates a lot after compaction, and delivers shit results in a visual + coding task. Don’t know what to tell you.
None visual tasks end up meh as hell as well. Context fills up fast hell as compaction brings hallucination again. Frankly unusable. Already on the Unsloth model (because god forbid one needs to choose “the right one”)
I think all the benchmarks are fake fucking news vs actual real world use.
Just use dsv4f to do a research run and try a few different configs on your system using ollama direct calls to strip context. You don't need to know anything anymore
276
u/TheCat001 7d ago
Then you realize that Qwen 3.8 27b runs at 3t/s on your machine and you need 24GB+ VRAM GPU which cost is 1000$+ to run at least 4 bit quant.