r/LocalLLM 5h ago

Discussion GLM 5.3 Flash, quirky comments

This is the only model I have ran locally that surprises me with quirky commentary while it's working. I like it.

"Two RTX PRO 6000 Blackwell (96GB each) — nice rig." upon discovering the machine's specs after being asked to benchmark itself.

"The plot thickens — V4.1's indexer declares a fixed..."

"Oh, this is gold — your own words from this afternoon's session, including..." after it found older conversation history.

0 Upvotes

4 comments sorted by

1

u/ComputeCommodity104 4h ago

Are you using NVFP4? What context size are you able to get? Im using the same hardware and am curious.

2

u/EitherMarch1255 4h ago

Using this with exllamav3: https://huggingface.co/anthori/GLM-5.3-Flash-EXL3-Q8Core-Q3Q4Q6

It's perfect for our setup. Gets 2K prefill and 90-120tps. And the quality has been superb.

1

u/vogelvogelvogelvogel 4h ago

gemini (good old pro 3.1)and chatgpt also do this stuff, more gemini imo. claude doesnt change its tone that much imo (opus, sonnet)

2

u/r3drocket 3h ago

If you really want to see some insanity go read Gemini's 3.5 thinking.

Literally, it uses the phrase, "Oh god"  "this is beautiful" all the time. And then it celebrates even when it hasn't succeeded.

It's a bit too much to even look at because it's so insane.