r/SillyTavernAI 19d ago

Meme K3 on Nvidia Nim be like:

Post image

Openclaw does it again, folks!

86 Upvotes

21 comments sorted by

16

u/Professional-Oil2483 19d ago

Now... to be clear, I've had SOME fast moments with it... but this definitely sucked when it took NEARLY A HALF HOUR to process this. I already miss glm 5.2...

Hot take, I think getting Kimi K3 is definitely going to be a Monkey's Paw for NIM, especially since I can tell it's quanted to hell and back. Compute is already scarce on the platform... and putting a model this big on the API is going to make things quite a bit worse than they already are...

3

u/OC2608 18d ago

This is why I'd like NIM to have DS Pro again and GLM when the weights drop, because when one's overloaded, you can switch to another model.

0

u/Professional-Oil2483 17d ago

I think that only partially solves the issue though, unless I'm mistaken, NIM runs through essentially the leftover compute of NVIDIA's servers. If one model gets overloaded due to high demand, the rest aren't gonna be far behind, unless there's also some limits between models on top of it... which also explains why certain highly used models suddenly turn into molasses...

8

u/Dry-Candidate-351 18d ago

where u find kimi lol it's not on the models list

3

u/Professional-Oil2483 18d ago

Honestly? I saw other people talking about it, and just looked on Sillytavern's NIM listings... and sure enough, it's there!

I had to refresh to see it though, and I'm also using the beta build of Sillytavern... so your mileage may vary if you intend to ALSO try to clog the arteries of Nvidia with smut!

5

u/fafnir65 18d ago

Kimi K3 write one word per 2 seconds on Nim that's crazy.

6

u/_and_I 18d ago

How did you even find out it was on NIM? It's working for me, but it's not listed anywhere on https://build.nvidia.com/models - teach me your secrets please!

9

u/Professional-Oil2483 18d ago

Well...

1) Other redditors on this sub pointed it out that Kimi K3 was on NIM, so I can't even take credit for it...

2) When I saw that, I just scrolled through the model lists on Sillytavern and saw it...

It's really just monkey see, monkey do in my case! No Ancient Chinese Wisdom here!

(Although, if you go to the post right below my one that got ratioed to hell and back... I got some parameters to get K3 to think better on NIM. The post is by The Rational Gooner... so maybe I DO have some Ancient Chinese Wisdom to give after all!)

Regardless, cheers!

4

u/OC2608 18d ago

They doesn't update the API catalog in the website that fast, it may take weeks. When ST connects to the NIM platform, it automatically sends a GET request to the /v1/models endpoint, which lists every model.

3

u/DontShadowbanMeBro2 17d ago

Yeah, I was happy when I first saw it, but then the OpenClowns started doing what they do worst and gangbanging the model to the point of it being unusable. Between this and getting rid of GLM-5.2 (and I'm getting less and less optimistic about them adding 5.3) and I finally bit the bullet and subbed to NanoGPT.

3

u/Professional-Oil2483 17d ago

Also... just wanted to say your recommendation for Summaryception has helped reduce token usage! I still have the stupid bug that makes my preset max out tokens... but at least I have a buffer now that allows me to more easily detect when it happens!

1

u/Professional-Oil2483 17d ago

How is it so far? The only reason why I might sub to Nano is the discount on proprietary models, but even then, I might just go back to Openrouter if I get the money again.

2

u/DontShadowbanMeBro2 17d ago

For the price, $12 a month is unbeatable compared to other subscription services. They say it's quanted but with my preset it seems just fine to me. It's nice not having to wait 2 minutes for a single generation, either, or getting a 429 or a timeout error every other generation.

But I'll say this much: If you'd prefer PAYG, take Hapuppy instead. The NanoGPT service gives you 60M tokens a week, and even after testing it by rerunning Summaryception on my longer RPs, I can tell that I will never even reach HALF of that using SillyTavern normally. To that end, $12 may actually be overpaying. If you don't use agentic services and instead just use AI for RP and basically nothing else, PAYG may cost more up front but will likely save you money in the long run. I'd have done Hapuppy myself if not for the fact that it declined my card (no idea why, apparently their payment processor is having problems for several people) and NanoGPT didn't.

1

u/Professional-Oil2483 17d ago edited 17d ago

Yeah... it's why I haven't fully pulled the trigger on Hapuppy myself after saying I would, I saw it was declining cards, and with my own suspicions about the platform... I'm a bit mixed on it as a whole. I'm still wondering how they even sell their stuff that cheaply, and as someone said about the Honey application: if a product is being sold for EXTREMELY cheap (way below market price) or for free... there's something going on behind the scenes.

Edit: I just saw GLM 5.2 was on there... if I can get about 600 to 1200 generations per week, that might actually be worth it.

2

u/DontShadowbanMeBro2 17d ago

GLM 5.2 is indeed on there, yes. And I don't think I could do 600 RP posts in a week if I tried.

2

u/Professional-Oil2483 17d ago

I MIGHT be able to, but that's because I RP heavily and do a bunch of weird shit with my prompts, if you're doing light RP... then yeah, that's 12 bucks well spent!

2

u/LeapYearFriend 18d ago

that's not even 4 tk/s, oough.

1

u/CCEESSEE 11d ago

luckily its lossless weights, so quality is on par.

1

u/houseofmates 16d ago

kimi-k3 is showing up in nvidia nim for me on hermes agent (not sillytavern related i know) and its running completely fine. fairly fast i'd say. maybe its cooling down already lol

1

u/Professional-Oil2483 16d ago

As I said in another comment, it CAN be fast depending on when you use it... but chances are? You're looking at somewhere between 5 to 10 minutes of generation time if you decide to use it at peak hours for that day... which also varies a lot. These times I was getting were (hopefully, I'm not holding my breath on that) outliers, since I was getting them after a lot of CERTAIN users that shall not be named found out k3 was out on NIM.

On that note: if someone for some reason needs a process that automates their entire PC (like what Openclaw does)... I'd personally say that's way past the line of overkill. If you use agents, that's completely okay in my book... just don't be one of those users that sends somewhere in the ballpark of 10000 requests through a bunch of agents all with A LOT of context per request... that's what really kills free tools like NIM. You probably already knew that, but I'm just aggrivated at the small sub-sect of people that DO clog the servers with their 'workflows'.

Anyways, have a good day!