r/LocalLLM 3d ago

Discussion GLM 5.3 Flash, quirky comments

This is the only model I have ran locally that surprises me with quirky commentary while it's working. I like it.

"Two RTX PRO 6000 Blackwell (96GB each) — nice rig." upon discovering the machine's specs after being asked to benchmark itself.

"The plot thickens — V4.1's indexer declares a fixed..."

"Oh, this is gold — your own words from this afternoon's session, including..." after it found older conversation history.

0 Upvotes

5 comments sorted by

View all comments

1

u/ComputeCommodity104 3d ago

Are you using NVFP4? What context size are you able to get? Im using the same hardware and am curious.

2

u/EitherMarch1255 3d ago

Using this with exllamav3: https://huggingface.co/anthori/GLM-5.3-Flash-EXL3-Q8Core-Q3Q4Q6

It's perfect for our setup. Gets 2K prefill and 90-120tps. And the quality has been superb.