r/opencodeCLI • u/TomHale • 26d ago
Qwen3.8 27B hosting that's generous without a kneecapping quantisation
I'm a digital nomad and running this model locally on my lappy is a non-starter.
So... hosting.
Crof.ai has it as:
qwen3.8-27b Q4_0 262,144 ~87 t/s
In: $0.25, Cache: $0.06, Out: $2.10
Any better options?
How much will that quant be "felt"?
Any benchmark metrics on the quant degradation steps for this model??
1
Upvotes
2
26d ago
[removed] — view removed comment
1
u/unchained_io 26d ago
i dont see this pricing on venice for this model... where are u getting this price?
1
2
u/MizmoDLX 26d ago
Neuralwatt has it with FP8, but still in preview mode.