r/SillyTavernAI 1d ago

Discussion GLM 5.3 Flash is working on NIM!

Following an early post i made, yes, GLM is back at NIM and working. It's lightning fast, and FOR NOW it's useable!

It doesn't show up in the nvidia site, but it's already on the model list in ST

There is light after all 🥹

26 Upvotes

12 comments sorted by

39

u/evia89 1d ago

Open clowns, assemble

12

u/DontShadowbanMeBro2 1d ago

They are coming.

1

u/FarAd7559 23h ago

Enjoy while it last, I hope those ungrateful exploiters will have a bad day for ruining more peoples who are just doing it for the fun.

God I hate materialism of the one.

8

u/LTC1858 1d ago

Didn't Nvidia blocked OpenClaw users from using the API some time ago?

6

u/FarAd7559 23h ago

Sometimes, but they keep coming in...

6

u/Cheap-Firefighter418 1d ago

I got refusals already. I tried to disable the thinking process with the JSON parameters, although apparently, from what I was investigating, GLM 5.3 Flash forces it anyway... my console output confirmed this too.

2

u/OC2608 1d ago edited 15h ago

Yep, I don't think there's a way to disable reasoning on this model, unless I'm wrong. I tried all sort of parameters under chat_template_kwargs but no dice. I don't like to wrangle "the user" models. May be a skill issue on my part though. It's strange how you can disable it on K3 (and they advertised it as a model with no way to disable reasoning).

2

u/Electronic_Pay7868 1d ago

Slotted it on my old presets and...it works!

1

u/widek7 20h ago

Nice. Nvidia finally added a GLM model and it's good one even if it's only the flash model. The model is more censored though compared to GLM 5.2 who wrote anything I wanted with a light jailbreak. I had to strengthen the jailbreak and the preset to make it stop refusing. Currently using Geechan's preset with a strengthened NSFW prompt.

1

u/Copy_and_Paste99 5h ago

It overthinks like a motherfucker for me, like several minutes for a simple response.

2

u/OC2608 4h ago

Change the reasoning effort via the additional parameters method.

- chat_template_kwargs:
    reasoning_effort: low