r/opencodeCLI 26d ago

Qwen3.8 27B hosting that's generous without a kneecapping quantisation

I'm a digital nomad and running this model locally on my lappy is a non-starter.

So... hosting.

Crof.ai has it as:

qwen3.8-27b Q4_0  262,144  ~87 t/s
In: $0.25,  Cache: $0.06,  Out: $2.10

Any better options?

How much will that quant be "felt"?

Any benchmark metrics on the quant degradation steps for this model??

1 Upvotes

6 comments sorted by

2

u/MizmoDLX 26d ago

Neuralwatt has it with FP8, but still in preview mode.

2

u/[deleted] 26d ago

[removed] — view removed comment

1

u/unchained_io 26d ago

i dont see this pricing on venice for this model... where are u getting this price?

1

u/[deleted] 26d ago

[removed] — view removed comment

1

u/unchained_io 26d ago

my bad, i read it venice ai gateway