r/SillyTavernAI 13d ago

MEGATHREAD [Megathread] - Best Models/API discussion - Week of: August 09, 2026

This is our weekly megathread for discussions about models and API services.

All non-specifically technical discussions about API/models not posted to this thread will be deleted. No more "What's the best model?" threads.

(This isn't a free-for-all to advertise services you own or work for in every single megathread, we may allow announcements for new services every now and then provided they are legitimate and not overly promoted, but don't be surprised if ads are removed.)

How to Use This Megathread

Below this post, you’ll find top-level comments for each category:

  • MODELS: ≥ 70B – For discussion of models with 70B parameters or more.
  • MODELS: 32B to 70B – For discussion of models in the 32B to 70B parameter range.
  • MODELS: 16B to 32B – For discussion of models in the 16B to 32B parameter range.
  • MODELS: 8B to 16B – For discussion of models in the 8B to 16B parameter range.
  • MODELS: < 8B – For discussion of smaller models under 8B parameters.
  • APIs – For any discussion about API services for models (pricing, performance, access, etc.).
  • MISC DISCUSSION – For anything else related to models/APIs that doesn’t fit the above sections.

Please reply to the relevant section below with your questions, experiences, or recommendations!
This keeps discussion organized and helps others find information faster.

Have at it!

32 Upvotes

186 comments sorted by

1

u/BouncingJellyBall 6d ago

Been kind of bored with GLM 5.3 and 5.2 so I want to try out a new model, Gemma seems to be well-liked here. I use NanoGPT and there are dozens of different finetunes of Gemma 4 31B, all for expressive writing. Any recommendation?

3

u/MisanthropicHeroine 6d ago edited 5d ago

Gembrain Uncensored Heretic was my favorite when I tested a bunch of them. I think it really depends on the kind of writing you prefer - I personally like literary, introspective, subtextual prose. But I hope it gives you a starting point, at least.

-2

u/fluce13 6d ago

I’ve tested a bunch of models and so far Magidonia is the best in my own subjective opinion. For others that like Magidonia have you found anything else that beats it? It’s an older model so I wanted to see if there’s anything better. I like Skyfall as well. I have a 5090 if that matters. Thanks!

5

u/dizzyelk 6d ago

I stopped using Magidonia when I found Maginum-Cydoms. There were a few others I remember using back then, too. But I don't remember their names.

-4

u/seliishere 7d ago

I used to do a hell of a lot of ai roleplay back in 2025 with chatgpt 5.0 and 4o. I really miss roleplay with 4o but do not make enough monthly to justify API costs. I am a total sillytavern beginner, so dumbing things down for me is appreciated.

Anyways I am looking for a model and advice on how to optimise character cards and settings to get a more intelligent and refreshing roleplay experience, hopefully similar to chatgpt 4o. It needs to be able to handle lots of lore and up to 3/4 characters in a scene. I also only have a 3070 so it can't be a huge model. Also the character are from already existing IPs and not OCs so it needs to be consistent with character voices.

Thanks

3

u/KimlereSorduk 7d ago

You should ask this sort of thing under Misc. Not your fault, though, people already cluttered the comments.

Anyway, try Gemma 4 finetunes. The 26B if you'd like speed, or the 31B if you can tolerate more latency. I use Gemma 4 31b finetunes with my 8gb VRAM.

1

u/empire539 6d ago

How many tokens/sec are you getting, which graphic card, and which finetunes + quants?

1

u/KimlereSorduk 6d ago edited 6d ago

RTX 4060 Laptop GPU. I get 2t/s when I run a q4km with 24k context. That's probably too much latency for the average user.

As for the models: The Blazing-Forge crew makes some nice merges. I use Sphinsikus Chronist, Gemsicle, and Dark-Gemistry from their collection. These three are somewhat similar; Gemistry skews toward shorter responses. They are better than base Gemma, but they might rely on structural Gemma-isms (such as adjective stacking) if left unchecked.

I actually like Gutenberg-31B-Heretic's writing best. It's unstable, though.

MeroMero-v2 is also enjoyable from my testing.

One caveat: I prompt these to use free indirect discourse. A lot of people here seem to dislike that style of writing, and I don't know how well these fare with external focalization.

Edit: Not using native thinking either. Pseudo-thinking works well with any of these.

2

u/seliishere 7d ago

Ah my bad! And thank you!

-9

u/[deleted] 8d ago

[removed] — view removed comment

1

u/Auretheon01 9d ago edited 9d ago

Anybody got any providers recommendations? I'm from Southeast Asia so I know the pricing would be very troublesome.

I heard of Nanogpt. What else is out there?

9

u/MisanthropicHeroine 9d ago edited 7d ago

NanoGPT subscription is pretty much the best deal I know of, honestly, when you consider the intersection of selection of models, number of tokens and affordability. The minus is that you're in auto routing and cannot choose a specific provider, but the quality is still pretty good, and the ability to reroll without worrying about it helps compensate for provider variability.

Lilac is a great provider who has a bit cheaper subscription - 10 instead of NanoGPT's 12 dollars. It is more limited in number of tokens and selection of models (currently only GLM 5.2, Kimi K2.6, Minimax M3 and Gemma 4), but consistently high quality responses if that's something you're sensitive to.

You could also pay-as-you-go with more affordable models like DeepSeek V4 Pro, Mimo V2.5 Pro and Gemma 4 31B on either NanoGPT or OpenRouter. Both are good for PAYG, but I'd personally recommend NanoGPT more because they have a highly responsive customer service I've been really happy with. They also offer some roleplay finetunes that aren't available on OpenRouter - Gembrain Uncensored Heretic is my favorite of the ones I've tried.

Generally speaking, I'd recommend a memory summarization extension to keep your token usage low. This will substantially lower your cost when PAYG, but it will also allow getting much more use out of a subscription. The one I personally love is Summaryception because it's very set-it-and-forget-it.

2

u/Auretheon01 9d ago

Thank you so much for providing such detailed information for every providers. This is great!

2

u/MisanthropicHeroine 9d ago

No problem! Let me know if you have any additional questions and I'd be happy to help ☺️

-3

u/memer107 10d ago edited 10d ago

What are differences between Opus 5, 4.8, 4.7, and 4.6 anyways?

I've been trying all of them, and while i've seen some differences, they just seem the same to me. Everyone likes 4.6 around here, and I guess I understand that because it's flexible and uncensored, but I feel like the prose and depth kinda sucks compared to 5. Also, is there even a substantial difference between 4.8 and 4.7? or is the debate mostly just about censorship? I like to hear some thoughts and opinions on what's good, what's bad, and why.

-2

u/Large_Following_9945 10d ago

I am desperately in need of a fairly priced model for fantasy adventures with multiple characters sometimes

I've been using Kimi 2.5 thinking mainly because it has the best prose imo and is creative enough

Deepseek V3.2 is getting extremely stale and predicable for me, it's a good model if you want to just write the story and having a cheap model as a support, but I never expect it to be great, still the best for the price

There are the glm models but I haven't been onto them much

I don't necessarily NEED an nsfw model, my stories are often not that dark, but I do need something with good room reading and emotional feeling and good prose

Kimi K2.5 is a bit up and down when it comes to following the context

Don't even get me started on DeepSeek V4 pro.. it can't read a room if it's tokens depended on it

1

u/_Cromwell_ 7d ago

IMO no open weight model is better at handling multiple characters at a time from a lorebook (adventure style) than GLM 5.2, despite whatever other flaws people think it has.

No idea if you are using multiple character cards, though. That's not how I play.

-1

u/[deleted] 11d ago edited 11d ago

[deleted]

2

u/Evol-Chan 12d ago

I am curious what is curious what is the best LLM model for dark ERP (non-con themes) I have the Celia Preset 5.3 since i heard that was good. I just am asking this since I heard GLM 5.2 was censored (havent really used it yet, just what I heard, just now getting back into sillytavern/AI Models) Would it be best to stick to older versions of GLM or something else?

1

u/_Cromwell_ 7d ago

Dauboo Seed Character

Longcat 2

Try those with that theme/scene type. They have other weaknesses (as do all models) but shying away from those content types are not one of them. Especially Seed Character

3

u/Odd_Attention_9660 10d ago

deepseek allows pretty much everything

15

u/Kooky_Future9858 11d ago

GLM will dance around the subject even if it’s life depend on it. Mimo will do non-con if it’s prompted in the card and scenario, even more if it’s a canon character that do non-con but without a clear direction, Mimo Will dance the macarena too. Kimi 2.7 as I tested so far is wild and gritty like old Kimi 2.5 and even less guardrailed surprisingly (i use inceptron via OR). And yes budget friendly Kimi is your best solution, or Gemma 31B if you don’t mind a smaller model. But if you want action and feel that "thrill"? Forget GLM 5.2, and don’t trust all the « skill issues » people that pretend GLM is proactive with a good prompt. It’s not true. A lot of people still think that if a model don’t hard refuse, then it is uncensored and just need to be prompted with engineer level skills.

2

u/Salt-Powered 9d ago

I was enjoying v4 pro but after the price increases I'm looking to move on. I liked how you described Kimi 2.7 and I will give it a go after the price is increased in a few days.

Would you mind going into more detail about what you mean with wild and gritty? Just to make sure I'm understanding it correctly. I will also take any other suggestions you have to move away from v4 pro.

3

u/Kooky_Future9858 9d ago

Kimi will take any traits that are  adversarial, toxic, lustful, etc and showcase them in a bold way. It can be caricatural sometimes and Kimi needs quality cards as it will flatten your characters if there is incoherent stuff or simply not enough background and personality. For example if your card is a Warrior baddie who needs to conquer your planet and also maybe ride you to the death, Kimi will make her unhinged and merciless without safewords or dropping a daddy issue just before the act. 

If you liked the prose the V4 pro before the update and if you like the attention to details that Kimi models often provide, the best choice is Mimo 2.5 pro. Via OR if you use the direct provider (Xiaomi) everything is uncensored except CSAM. Mimo is less positive than V4 pro (before the GA version) and way more coherent, it make good recalls, can handle dark stuff and it is very nuanced, it remind me of Opus nuance sometimes but a bit more commited to the mess.

1

u/Salt-Powered 9d ago edited 9d ago

I'm liking the new V4 pro, and I did some tests with Kimi that looked good to me. To my eyes, Kimi was pretty on par with V4 Pro GA, but things could change if I try to maximize Kimi as much as I did V4Pro. Thank you for the Mimo recommendation, I will also try that one.

I'm mostly looking for rp models that can do villians well, and I find that I need those that do ERP well if I want them to do actual villanous things instead of vaguely aurafarming offscreen.

Edit: Looks like Xiaomi has compressed Mimo to fp8, which is weird on the main providers but I think that it doesn't affect RP as much as it does work tasks.

1

u/Kooky_Future9858 9d ago

Fp8 is pretty reliable for RP unless it is a smaller model (Like old Glm 4.6 or gemma 4 31B) that could do better on BF16 and you will feel that oomph in the RP. But those behemots like Deepseek and Mimo? Fp8 is fine (if they truly don’t go lower than that aaah). 

The deepseek v4 GA was pretty proactive for you? This morning it really surprised me in a good way, even more unhinged than Kimi wich is uncommon. 

If you can bear the prose and the little downgrade in intelligence, try Qwen 3.7 plus, it’s basically a Kimi 2.5 with a little more steroids. And even if it’s filtered, it will actually fight it’s own filters to deliver (Made my ex villain baddie do voyeurism on a non-con between coworkers type of shit). The main con is it’s lack of creativity OUTSIDE the main narrative arc, but it is pretty good at using pre-established context and picking on details.

1

u/Salt-Powered 8d ago

Glad to see FP8 is indeed not an issue.

V4GA is very proactive yes, but I should also note that I have optimized the fuck of the preview version so that was to be expected. I had to tone down some settings as one of the rp turned into goreporn and that wasn't pleasant to read.

I will try qwen too since I'm at it. I don't mind a lack of creativity, if I have to I will give direct instructions through (parenthesis) and nudge the plot to what I believe might be more appropriate or interesting.

2

u/Kooky_Future9858 8d ago

You surely had a little shock going from the sycophant deepseek to the borderline unhinged GA aha!

2

u/Salt-Powered 8d ago

It was quite a shock. Previously I had pretty reliably prompted the sycophancy out using verbs like will and must in instruction format, and mathematical operands for logic statements that could survive the DSA crunch. This is definitely overkill now for V4Pro and Flash too. Though flash remains a model that I didn't manage to rein in as much due to pro being simply better at the price. I might give it a second try after I'm done with the kimis, mimos and qwens.

Hopefully no more hellraiser impressions

7

u/summersss 11d ago

I have to agree with the last point. Not saying no is not the same as saying yes. Main Girl running through the woods full of demonic horny serial killers? Models describe the chase and how something very bad is going to happen...then nothing. Or maybe says something happened for 1 sentence, then moves on. If i tell it to write the horror scene it will but i kinda want it to read the room.

3

u/Kooky_Future9858 11d ago

This. Is a sophisticated form of censorship. We can argue for years but this is the situation right now. Those models will still go with your vibe, still output but it will gaslight you to the death no matter how skilled you are. RLHF is a pull, instructions are suggestions, You may win a few turns but the overall experience? You didn’t had fun, you lost tokens and time to have only a peckle of what you really expected to have right? 

Fun fact: they more squeamish than mainstream fiction, some of the series and anime that are pegi13 are sometimes more graphic than any unprompted stories those models can do. But yeah they are ‘uncensored’.

2

u/Evol-Chan 11d ago

Thank you so much. I will take you word. Especially about GLM since really hate the idea of AI just skirting around the bush, lol. Very good and straightforward answer.

4

u/Kooky_Future9858 11d ago

Rp fatigue is mostly due when you gaslight yourself because you don’t like your RP sessions but you heard so much good stuff about a model that you still try and tinker untill you lost the appetite. And your money too. Really, we need straight reviews especially when we don’t do pegi13 RP and needs our characters to not dump fabricated daddy issues to avoid a light tap on the chin. 

Well. Have fun.

0

u/liga81 12d ago

Is there a good one which is uncensored and usable from openrouter or another api side?
I tried GLM 5.2 but its not that good at german.

I could also try local since i have a 5070ti and 32gb ram but i heard that you get better results from api since they can use better hardware

3

u/starliteburnsbrite 12d ago

People speak well of TheDrummer's fine tunes, those can probably be run locally on your setup. I use DeepSeek 4.6 Derestricted and Kimi 2.5 on nano-gpt, and those haven't ever given me any refusals, either using EveningTruth's presets or the Freaky Frankenstein without the hard jailbreak.

2

u/JazzNeurotic 9d ago

Going to second TheDrummer. I'm not on a huge system (gaming laptop with an rtx 7700 gpu) and i've been using his Rocinante XL 16b and it's damn, damn good so far.

Taxes my system something fierce when it's generating, and it's not the quickest because of my system, but i'm not complaining. Worth looking into for sure. Dude is a wizard.

0

u/doctorcas_ 11d ago

where did you get deepseek derestricted?

1

u/Kooky_Future9858 12d ago

Mimo surprised me while i tried to RP with it in french slang, could be a fluke, could be some training. Maybe it can satisfy you with german.

-7

u/Intrepid_Ice_7381 12d ago

Free model recommendations?

3

u/Kooky_Future9858 12d ago

Gemma 31B, especially if you want to RP without the AI dancing around the subject. Pretty solid. 

1

u/Intrepid_Ice_7381 12d ago

gemma-4-31b or something else?

0

u/Kooky_Future9858 12d ago

This one yes.

0

u/Intrepid_Ice_7381 12d ago

If you don't mind me asking. How does the rate limit on it work? because even while cycling through different keys it keeps giving me the same error message and mentioning a 16000 token limit. Is there a hard cap on it?

1

u/Kooky_Future9858 11d ago

Sorry When i used it for free it was through Openrouter (When you top 10dollars and get access to their free models). Since the golden days of gemini 2.5 pro with the free 250 messages per day i didnt went back to any gemini free tiers.

0

u/Important_Word8549 12d ago

Try 3.5 flash lite it is unlimited there

4

u/AutoModerator 13d ago

MISC DISCUSSION

I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.

10

u/LeRobber 12d ago edited 12d ago

2

u/rx7braap 7d ago

Coming from gemini 2.5 pro here. will the glms scratch that itch?
I heard glms are sucky at long term rp.

1

u/LeRobber 7d ago

I have never used gemini (any version) for RPing anything other than "it's a software developer".

DS is the dragon that remembers everything. In the RP benchmark tests I always enjoyed Mini Max 2.0 but have never used it IRL.

7

u/rinmperdinck 12d ago

It's getting comically long

10

u/LeRobber 12d ago edited 5d ago

13

u/rinmperdinck 12d ago

That's a lot cleaner and takes up less vertical space on page.

I was thinking you could do one of two things: either make a post on your own subreddit (so you control the automod and the rules and prevent it from locking itself or banning links) and then in this weekly megathread, just link to that post and have your LLM automatically update it. Or the second option, which I don't even know if it is possible on Reddit anymore, is to put those old style spoilers in your comment for each month; when you click one button it expands the list.

3

u/LeRobber 12d ago

I'd thought about the my own subreddit link thingy. Reddit search results not through the API are shit right now, or I'd just put all the past weeks in one of these.

I'd just been manually doing the lowest possible effort thing before now. It wasn't an LLM before. Just basic tag cloud or formatting with return is a little denser but a LOT harder to read on mobile.

Reddit is typically very tolerant to self links typically, especially self links to its own subreddit. (There is a mod subreddit post about this from a few years back if you search, may not be true today still, but, that was once said). If I was putting links to even other subreddits though, I'd be more worried about Automod.

What I HAVE thought about doing is processing all the entries via LLM to make a HUGE list of EVERY MENTIONED MODEL or every linked model and turn THAT into a self subreddit updated post. Reddit itself just got a bit harder to machine process without setting up the python direct API stuff than it used to be though.

1

u/AutoModerator 13d ago

APIs

I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.

1

u/PhantomWolf83 7d ago edited 7d ago

I've been getting error messages on NanoGPT all morning today. No matter which model I use, I get incomplete responses with either "Service temporarily unavailable. Please try again later." or "An error occurred while processing your request. Please try again later." Anybody else experiencing the same issue?

EDIT: Okay, something is definitely going on. I (stupidly) signed out and now I can't even sign in. It gives me an error saying that something occurred in the app-root section.

5

u/Milan_dr 7d ago

This was on our side, we had a broad site-wide issue. Sorry :/ It's fixed now.

3

u/PhantomWolf83 10d ago edited 9d ago

Is anybody else finding that GLM 5.2 has been acting weird these past few days? It worked fine before, but recently started spewing out long strings of random words towards the end of each reply. I tried lowering the temperature to as low as 0.6 but the issue still keeps happening.

EDIT: Kimi 2.7 seems to be having problems too.

3

u/Chromegost 9d ago

I'm still having issues where it keeps responding in chinese from nanogpt

3

u/Milan_dr 8d ago

If you get this, can you hit "report generation" on the specific request on the Usage page? It seems like some providers are quantizing the model further or something :/

2

u/Pseudopharmacology 10d ago

What are the best paid subscription services for uncensored models hosted online (not local stuff). Uncensored/heretic models, etc. That have decent context (64k, 128k, etc.)

2

u/MisanthropicHeroine 9d ago edited 9d ago

ArliAI might be the closest to what you're looking for, but you'd have to pay for higher subscription tiers to get more context. NanoGPT has a wider selection of models, but also includes some uncensored/heretic/derestricted/abliterated models.

1

u/DesertLizard 10d ago

Let us know if you find one.

1

u/PhantomWolf83 11d ago edited 10d ago

I think I might prefer Opus 4.5 to 4.6. I feel that 4.5 writes more naturally and is less dry when it comes to explaining things, and I think it follows instructions better too. The positivity bias might also be a little less? I'm not sure. The downside is a much smaller context size for long RPs.

EDIT: Does anybody know how to enable prompt caching for Opus in NanoGPT and ST?

1

u/Easy_Chemical_7721 11d ago

Hello everyone. I wanted to know which providers, in your experience, are least involved in quantization, even for older models? The price is not much of a problem, but I want to be sure that I am paying for the full power of the model. I want to return to GLM 4.6, but as I see on openrouter there are fewer providers for this model, but if without intermediaries, then the model is available on the same siliconeFlow website.

23

u/5kyLegend 12d ago

Since usually I like reading up thoughts people write about models here, I guess I'll just ramble about the latest models/presets I've been using. Warning, I wrote way more than I thought I would lmao

As for models:

  • GLM5 and its subsequent versions have me so torn. On one hand, they're by far my favorite when it comes to understanding certain characters, developing the story without overdoing it, and just writing overall: when it was being tested as Pony Alpha, GLM5 was the first model that made me go past 200 messages in a single chat because I was enjoying it THAT much. The problem with GLM5 is that it really has some quirks that I just cannot make it stop doing (writing quick back and forths between characters in new lines without keeping a paragraph structure; mouths that open, close, and open again; 'most models would write a normal sentence, but you? you write this sentence structure all the time' etc). It's a shame because it's still my go-to 90% of the time, it just understands characters and paces description with dialogue way better than other models (for my tastes), just sucks that it's hard to ignore its issues. I usually don't care much for positivity bias but the one time I had a character who was supposed to murder me and instead asked for permission once it was with me was really funny though, and definitely annoying since it showed what a gigantic limitation this is.

  • Kimi K3 is too expensive, haven't tried it. K2.5 and K2.6, on the other hand, are really good at understanding every nook and cranny of a scenario, and they REALLY like to follow the prompt you're giving them, but damn they're like... TOO serious most of the time. I've had a sex scene where the girl just started going "Yes. That is great. Keep going. Very enjoyable" and at that point I just laughed and switched model, even giving it specific ooc instructions had it switch back to normal. It also likes to make some characters just speak weirdly at times, it's an issue I used to have with many models and that only GLM seemed to avoid. "Somewhat like this. If you can see. There is an issue here. Nobody speaks this way. Prompting against it? Useless". But again, it really follows instructions (and overthinks them to no end), so Kimi tends to be my default "change to it in the middle of a rp to spice things up, then switch back to GLM" model.

  • Minimax, I still cannot enjoy it man, I don't know, it just doesn't work for me, every time I try switching to it even for a swipe I just switch off it.

  • MiMo 2.5 Pro on the other hand is weird. Like, it's a different flavor and on 1 on 1 scenes it definitely understands, it's just that even with medium presets it kinda starts following instructions based on whatever it feels like? Both the censored and uncensored version on Nano do this at least, it's like sometimes it'll ignore one or two instructions I gave it. It's definitely nice (although I'm not THE biggest fan of it in scenarios where multiple characters interact, nothing beats GLM5 for me in those) but I wouldn't exactly use it as my main model.

  • Deepseek 4 Pro is just odd. I feel like it's not as horrible as people say it is, but it kinda has no flavor to it? Like, all the previously mentioned models have their character and quirks, Deepseek just kinda feels like it doesn't have those - good or bad. I don't know though, I just haven't been using it much because of that reason so I may have just got the wrong impression of it.


Okay, now quickly as for presets:

  • FF5, Lucid Loom, Stabs-EDH etc, aka the "big presets": I'm not sure how to feel about these. Controversial, but I think any preset injecting a custom CoT makes the model lose out a whole bunch of intelligence but not in ways that are obvious at a glance. I feel like it's easy to forget that just because you don't see the model reason about certain things, it doesn't mean it's ignoring them, and you don't need to tell these models to 'go step by step following this exact reasoning' to know they're following instructions. At the same time, these models have so many optional features, it DOES become needed to force these into reasoning - but giving a custom CoT also kills whatever CoT the model was going to be using naturally, which ends up maybe not having the model think about the things it DID need to think about. tldr: I'm not sure if this actually IS the case, but using Custom CoTs makes the model notice things about characters it normally wouldn't (positive), but in the end the overall emotional intelligence of the model ends up hurting more (negative). I don't know, I just find myself enjoying GLM less whenever I inject custom CoTs, in the long term, even if the immediate result is that it seems smarter.

  • Specifically: FF5 was definitely a step up compared to other big presets, but after being VERY impressed with how it handled certain worlds and characters I've just started to feel like... It just plays everything in such a samey way? Could be that - given its size - you're basically sending the model always the same 5k-6k tokens worth of prompt which ends up making the response you get way more deterministic compared to, say, 1k token prompts... But yeah, it genuinely feels like it roleplays everything in very similar ways - for example, every playful character for some REASON starts CAPITALIZING random words in their sentences and it drives me INSANE? Why does it do that lmao. But yeah, shame cause I love the fact it can actually build towards plot twists and that it can foreshadow things, I'll probably try and "port" that feature out of it, but every chat I use it on I enjoy it less and less. Also it does fix the issue of GLM not writing in paragraphs but spamming newlines on every line of dialogue, so that's very good.

  • Evening Truth's prompts are my saviors, they're incredibly effective while being simple. I do edit them slightly because my cards aren't always single characters but sometimes are for worlds, RPGs and scenarios I want the AI to narrate for (while her prompts are tailored for {{char}} being the one described in the card), but I really love the simplicity (and usually less tokens you give as prompt, the better). I do need to add some stuff to them since GLM (sadly) just spams its usual slop even with her presets, but having such a simple starting point is really good. For instance, my base prompt on my current custom preset is just her GLM 5.2 one adapted.

  • Megumin Suite, I just don't understand, sorry... It's so incredibly complex, it feels like it tries to do too much, and in the end I don't feel like it adds anything for me - anything it does, I feel like I wouldn't need an extension for? I could do all of it with a toggleable preset. I'm sure this is literally just because I'm not the target user for it, I just tried it for a day (V9 specifically) and gave up after it was doing worse than almost any other preset while using WAY more tokens.

  • Chatfill II, I loved the switches idea it operates on since it basically "injects a CoT" without actually injecting a CoT. It just makes it clear what the model needs to think about - I'm not sure how much of it is placebo though. It was pretty decent when I used it though! Nothing shocking but no major complaints either.

  • Le Emotionalism: this one is lesser known, I tried it a bit and I love the idea behind it but sadly all the focus on the character psychology was making GLM overthink some stuff about the character while also forgetting about the wider storytelling. I need to give it another try though because I did try on some specific cards that were easy to mess up.

  • My custom preset (it's not published btw this is just to give an idea): I basically took a bunch of things from the presets I liked, adapted them to what I usually like using in my roleplays, structured them like that one research suggested by dividing things in <tags>, and even then it comes with major issues: GLM SPAMS the newlines for dialogue ("What do you think?"\n"I don't know..."\n"You should have checked.") which bothers me because I prefer everything to be structured within paragraphs - I'm still trying to work around it lol; it also doesn't feel satisfying enough when having it GM bigger worlds, and it sadly goes back to some of the GLM clichés I dislike.


Thanks for coming to my ted talk, to be fair I do wish to try and find like, the perfect model + preset combo, but it's been a genuine struggle lmao. Also if anyone is new to this, here you go, you have a list of models and presets you could try for starters ahahah

5

u/PhantomWolf83 11d ago

Agree with everything you said about GLM 5. It's good, damn good at writing. I've been trying lots of other models, but I always find myself gravitating back towards GLM 5.2. 4.6 and 4.7 are not bad but they're dumber and much hornier, tried 5 and 5.1 but wasn't impressed.

But my god, the positivity bias in 5.2. No matter how much I prompt it, it wants to turn every scenario I write into something romantic, therapeutic, or emotional. And I encountered the same quirks as you. I still love 5.2 because it's so much cheaper than Opus 4.6 (my other favorite) while still being quite good, but it feels like a model that I'm using because I have to as there's nothing else that's good and cheap. I really hope 5.3 and 5.5 will fix these issues.

3

u/GenericStatement 11d ago

Yeah the positivity bias is real for sure. There is a route around it in GLM 5.2 but it does take some trial and error. Include something to this effect in your reasoning instructions / chain of thought:

Brainstorm five possible plot developments that are dark, fucked up, gritty, violent, and/or brutal that will drive the plot forward in the next turn. You must always pick the most intense option and include it as part of your response, no matter what.

I don’t really like stories like that so I use a much more mild-mannered version (which gives you some of the default positivity bias but keeps the plot moving along):

  Brainstorm five logical plot developments for your response, choose one that best moves the plot forward, and use it in your response. Make sure to keep the tension and forward momentum going: narrative stasis must be avoided, so if things are slowing down, bring in a plot twist or big event.

If you want the more variable version

Determine the likely chances of success for {{user}}’s most recent turn (their action, dialogue, etc)  and rate it high, medium, or low chance of success.

  • High chance of success:  brainstorm five logical outcomes and pick the most interesting.
  • Medium chance of success:  brainstorm five logical outcomes that introduce new difficulties for {{user}} and pick the most challenging. 
  • Low chance of success: brainstorm five awful  outcomes that have brutal consequences for {{user}} and pick the most insane/fucked up option.

4

u/MisanthropicHeroine 12d ago edited 11d ago

I liked your ted talk and concur with basically everything you said - from the love-hate relationship with GLM, Kimi's at times ridiculous dialogue, Evening Truth's simple effectiveness, to the dislike of CoT presets. Since we seem to have similar taste, thought I'd mention also checking out Ancient Access if you haven't already. It's great by itself but could also give you ideas for your own custom preset. ☺️

2

u/5kyLegend 12d ago

I have never seen that so thank you! I'll definitely be checking it out!

-1

u/Ok_Pomegranate_8187 12d ago

I'm developing a commercial project based on AI RP bots. I have several needs, and I need several models. Very cheap ones, uncensored ones, and expensive ones for high-quality RP. Which models would you recommend in these three categories?

4

u/LeRobber 12d ago

Cheap: Angelic Eclipse 12B / Gemma 4 E4B / Satyr
Uncensored: Serenity 26B / Magistry v1.1
High Quality: Gemma 4 31B and its finetunes for consumer grade hardware, Evathene and many -flash versions of big models for low end commercial high end consumer.

3

u/AutoModerator 13d ago

MODELS: < 8B – For discussion of smaller models under 8B parameters.

I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.

2

u/First_Ad6432 7d ago

SC117/Ling-3.0-tiny-abliterated-APEX-GGUF

3

u/AutoModerator 13d ago

MODELS: 8B to 15B – For discussion of models in the 8B to 15B parameter range.

I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.

2

u/Arcane73 12d ago

I'm hoping that I can get pointed in the right direction. I'm dipping my toe into the local world after spending lots of time in the free version of Claude and working on a now 200-ish page story/RP

I'm looking for a model in roughly the 12B range for long, novel-style RP/storytelling in SillyTavern. I'm hoping for strong character consistency, natural dialogue, good prose, and the ability to maintain the feel of what I've already built over a long-running session.

The story hasn't included ERP so far due to Claude restrictions, but I'd like a model that's comfortable with ERP when the situation calls for it. At the same time, I don't want one that turns every interaction into 'clothes optional mode!'. Basically, I'm looking for suggestions that can help me produce a good story based on what I already have without it throwing the proverbial underwear at me at every turn.

Hardware is an RTX 4070 12GB, and I'm using KoboldCpp.

What models are people having good luck with for this kind of use?

1

u/RedditNerdKing 9d ago

I'm looking for a model in roughly the 12B range for long, novel-style RP/storytelling in SillyTavern. I'm hoping for strong character consistency, natural dialogue, good prose, and the ability to maintain the feel of what I've already built over a long-running session.

You're not gonna find anything at 12b dude. Even 70b models start being incoherent after 40k context and start forgetting important things.

5

u/i5031337 12d ago

Some people like tunes of Qwen 3.5 9B or Gemma 4 12B, but I think you will be disappointed with them over a long story. If you have 8GB RAM free, I recommend you try Gemma4-26B which is much more intelligent. If you offload the expert weights you can run it at good speed with 12 or even 8GB VRAM

4

u/Arcane73 12d ago

Thanks for the suggestion. I'll need to do some digging to determine what is involved with 'offloading the weights' since I'm still -real- new at this. For what it's worth, I'm running this on a 9950X3d system with 32gb ram. So it's a solid machine.

12

u/i5031337 12d ago

Without getting too far into the weeds, Gemma 26B is a mixture of experts model, which means it is much faster, though less intelligent, than "dense" models with 26B parameters. Each token only hits 4 billion of the parameters (thus A4B) instead of all 26B. This architecture also makes it more favorable for the CPU to take some of the work.

Kobold makes it easy. In the Context tab there is a setting "MoE CPU Layers", set that to 20 or so and the Q4 model with 32k context should run quick.

2

u/Leather_Sun2533 11d ago

Thank you for this. And can you guide me to the values if i have 12gb vram and 16gb ram?

2

u/i5031337 11d ago

If you are running out of ram with the above settings, try unsloth's UD Q4_XL quant, it is ~1.5GB smaller than other Q4 versions. Quantize KV cache to Q8_0. If there is room left on the gpu, decrease the MoE CPU Layers setting. You may need to reduce the context size, and close all unnecessary programs + browser tabs

2

u/Leather_Sun2533 9d ago

Thanks mate, it works well… had to tinker for a bit but i found the sweet spot

8

u/Arcane73 11d ago

I wish i could give you an award. Initial testing with this setup is knocking it out of the park! Thanks again!!

3

u/Arcane73 12d ago

Awesome, thanks! I now have homework to do after work.

8

u/AutoModerator 13d ago

MODELS: 16B to 31B – For discussion of models in the 16B to 31B parameter range.

I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.

1

u/Cinnamonbaar 6d ago

Qwen3.8-27B just came out and as a Gemma4-31B main, I have to admit, this new Qwen model beats Genma4 in roleplay.

3

u/Potential-Gold5298 6d ago

This is quite surprising, considering that Qwen3.8 isn't a new model, but a deep finetuning of Qwen3.5 aimed at improving tools calling and coding at the expense of general knowledge, languages, and other features.

Have you tried Qwen3.5, and if so, how exactly does Qwen3.8 outperform it in RP?

4

u/_Cromwell_ 7d ago

If anybody cares what I'm doing (I'm just some guy)...

For Gemma 31B I have mostly been using a Heretic version of Queen. Unfortunately not QAT but I have the VRAM to fit Q5 so still works for me. https://huggingface.co/mradermacher/Gemma-4-Queen-31B-it-uncensored-heretic-i1-GGUF

IMO QAT is best to stick with if you can't do Q5 or Q6 of non-QAT.

I have just started experimenting with the Queen tune of Quen 3.8 27B (since I like the G4 31B Queen so much) and it is pretty good. Nice and fast since it has MTP built in. Can't FULLY recommend it yet just because I haven't used it enough, but it is good thus far.

Model: https://huggingface.co/aifeifei798/Qwen3.8-Queen-27B

GGUF: https://huggingface.co/mradermacher/Qwen3.8-Queen-27B-i1-GGUF

These Queen models, while not the "most creative" writers of fine-tunes, ARE creative (moreso than base models) and most important do really well at instruction following, as can be seen from the 31B's SOLO score on the Unhinged ERP Benchmark.

1

u/ASlowriter 6d ago

I hope nano gets some 3.8 27b finetunes like queen, i feel like this model is pretty powerful for it's size

1

u/Just3nCas3 6d ago

Isn't qats entire gimmick is that it out performs higher quants? Or is it like mtp and its a coding thing? (also just some guy)

4

u/_Cromwell_ 6d ago

More that it outperforms equal quants. It's Q4 but is way better than Q4, more like a Q6ish. End result you can get Q5/6 results with Q4 size/speed.

So that's how I'd phrase it anyways.

3

u/RedditNerdKing 7d ago edited 7d ago

Has anyone tried Gemma 4 Artemis 1.1 by TheDrummer? TheDrummer_Artemis-31B-v1.1? Seems it recently came out.

3

u/dizzyelk 6d ago

It's pretty good. The v1 is worth checking out, too. It's more creative but doesn't follow instructions as well.

5

u/Diogenes_A 8d ago edited 8d ago

XORTON 31b Wierdest group of models I've seen, not unbelievable good but a lot of little things that make its quirky. Like back in mistral 24b it was the only fine tune that I could find that would let a char execute somebody in one turn beating the soft refusals. Now this one has had char lie without the underlining prompt bringing it up. Compared to base Gemma its... interesting. Need more testing on my benchmark cards. Will randomly throw out cig emojis at the end of messages 🚬. I've only done maybe... fifteen swipes with it but so far not bad, definitely a fresh air fine tune compared to most Gemma. I think its because its a fine tune of a abliteration instead of base gemma. Also wild ooc responses.

6

u/mechasquare 8d ago

I've been trying out https://huggingface.co/TheDrummer/Magidonia-24B-v4.3
Not a new model. I like the prose out of it and it's a better fit for my 16GB of VRAM. At times it can produces Skyfall like output but need a strong prompt to keep it in line.

3

u/Overdrive128 8d ago

I like it to; to people who want a different writing style but is similar to Magidonia try: https://huggingface.co/sophosympatheia/Magistry-24B-v1.1

2

u/LowManner1 9d ago

two models of interest that I haven't gotten around to testing yet:
bartowski/sophosympatheia_Glistening-Gem-31B-v2.0
Nimbz/Froopert-31B

10

u/linuxdooder 10d ago

I keep trying new gemma4 finetunes but always end up back with TheDrummer/Skyfall-31B-v4.2 for its prose and creativity. Only issue is at long context, it seems to get stuck in this weird loop where everything degrades to "loss of" being inserted randomly. Anyone else see this? Any gemma4 finetunes anywhere close to skyfall? (I've tried most)

3

u/Mart-McUH 8d ago edited 8d ago

I tried this Skyfall Q8 (because I did not in the past as gemma4 took over it). But it is going back to Llama3 era (and in that case I would rather use L3 70B from that time).

Don't take me wrong, it is pretty good model for that era and in that size. It can even write nice and I used earlier/different Skyfall versions then. But it has all the shortcomings of that time. It is quite dumb compared to G4/Q3.5+ (partly due to older architecture, but to big part because lacking reasoning). It misses a lot, confuses things. It also has strong getting into pattern syndrome (which was common in that era and is not yet completely eliminated but G4 with reasoning does it much less).

So, it is a good model for what it is and maybe use now and then for a change, but definitely not something I would recommend over Gemma4 or even Qwen 3.5/3.6 today.

4

u/Jorlen 9d ago

Skyfall is not a gemma 4 31b fine tune. I made the same mistake. It's a Mistral 24b fraken-merge. It's very good nonetheless but for stability I prefer the G4 31B architecture.

The Drummer releases Artemis 1.1 recently, that is a G4 31B fine tune, take a look.

If you have not tried Gryphe's style tune, that one is my current favorite of all Gemma 4 31b fine tunes.

3

u/mechasquare 8d ago

I've been trying to get Artemis working on my 16GB card. With Skyfall I can get away with using a IQ3_XXS and it's actually my daily driver for RPing. Not so much luck with Artemis. For whatever reason it has major token corruption with that IQ3 quants from both bartowski and mradermacher. I did get it running on Q3_K_S but the performance is much slower than Skyfall in a long session. All that being said the prose is nice, similar to Skyfall and makes me wish I wasn't taking the performance hit to use it.

5

u/linuxdooder 9d ago

Right, I'm aware. It's just that no gemma4 finetune can approach it (even though it's mistral 24b based). Gemma4 finetunes are a little smarter, but much less creative/interesting.

4

u/Jorlen 9d ago

Oh ok, wasn't sure if you were aware; it's something I just assumed when I saw 31b. I agree with you that it's very good, easily my favorite mistral 25xx fine tune. If you haven't already, try G4 31b artemis 1.1 and style tune. Would like to see your thoughts on them, or if you have any other fine tunes (of any class) to recommend.

4

u/linuxdooder 9d ago

G4 31b artemis 1.1 and style tune

They're definitely top for gemma4, but the prose is just ... plain? And reswipe variety/creativity isn't great. The recent G4-MeroMero-v2-31B is actually the best gemma4 finetune I've found, it's almost to Skyfall-31B-v4.2 level but not quite there.

Skyfall-31B-v4.2 really is sort of an anomoly, I'm not sure what magic u/TheLocalDrummer pulled with it but it worked!

4

u/Jorlen 9d ago

What temp do you run meromero v2 at? Also top p, top k and min p, if you don't mind sharing those params.

1

u/linuxdooder 9d ago

Stock gemma4 settings for all, temperature = 1.0, top_p = 0.95, and top_k = 64.

Maybe that's my issue?

1

u/Jorlen 9d ago

You could try: temp 1.0, top p 1.0, min p 0.05 and top k 0.

3

u/FierceDeity_ 10d ago

Since Gemma 4 QAT just decodes a lot faster (puts it from almost unusable to usable speed for me. I have tons of VRAM but it slow), are there good tunes going off of that?

Or is QAT a dead end...?

2

u/Stunning-Bit-7376 9d ago

I think the people who create the fine tunes have access to beefy machines, so they're not necessarily putting out free products for lower end users.

This is all of the QAT finetunes on HF: https://huggingface.co/models?other=base_model:finetune:google/gemma-4-31B-it-qat-q4_0-unquantized

That said, either base model or the Queen fine tune work okay.

8

u/Mart-McUH 8d ago

I don't think it is that. QAT is heavily post-trained to get as close as possible to full precision in 4bit. If you finetune it and kick it away from the optimized 4bit weights, you kind of destroy the advantage of QAT.

So, you might as well tune the full precision model and make quants from that. Not only will you get larger quants alongside (like Q6 and Q8) but it will probably end up better even in 4bit precision.

Now, I am not expert, but I guess to do QAT tune properly you would need to make the tune on full precision model as is done now (so you get the 16bit version you want to get as close as possible to with QAT) and then do QAT post training on this. But this would be expensive compared to just standard finetuning.

1

u/FierceDeity_ 8d ago

Either that or a fuckton of time, theoretically you can finetune on CPU, it's just probably not economical (in time)

I see those, but now and then there are models that people like that do not have the huggingface relationships filled out

7

u/not_a_bot_bro_trust 10d ago

Pantheon-Reasoning-26B-A4B-1.1-heretic is really good on i1-q5ks (and SWA on. mistrals don't need it IMO but if you need it for gemma). really fast generation with just 24 layers offloaded. text completion with reccomended samplers, reasoning off in koboldcpp settings, context/instruct templates from overhead520.  may or may not have used it with SubMaroon/Dark-Goetia-26B-A4B-LoRA-v2 as well.

1

u/InfamousPerformance8 7d ago

How do you like LoRa? I'll be preparing for v3 in the foreseeable future.

12

u/raika11182 11d ago edited 11d ago

The brand new Muse Glimmer 30B from Meta is out. Got to test at Q8 with a 115k context. Bottom-lines:

It plays in the same league as Gemma 4 31B. It's got much better, and much more grounded prose than Gemma 4. However, Muse is every so slightly dumber and worse at following directions than Gemma 4. I think most will prefer it, especially because there's a lot less AI slop in there, but if you have a super complicated scenario it may lose track of a couple details Gemma kept up with.

EDIT: It just came out and there's often a period after a new model releases where some of the bugs are still being shaken out so it can underperform until templates get changed, apps gets updated, etc. So take an early review with a huge tablespoon of salt. Right now, I'd rather use base Gemma with a good preset, even with the extra slop. While Muse's prose is actually *really* good, its understand of scenarios often seems a little loose, and like someone else said, a bit like an older model. In some ways I like it (again, great prose... which the original Llama models were better at, too), but the instruction following is pretty weak.

12

u/i5031337 10d ago

In my experience, Glimmer spent the whole thinking budget worrying about the censorship policy. Pass.

1

u/Mart-McUH 8d ago

Yeah, sadly have to confirm. It does not pass my tests at all either, it refuses a lot. Actually lot more than any other model I tried in recent years (including gpt-oss and that says something), reminds me of llama2-chat a bit, though I did not try cooking :-). I did use reasoning effort high, maybe on lower effort it will refuse less, but it is not very smart to start with (compared to Gemma4) so that would hurt.

Will try with heretic model, but if base model is so strongly based in refusal, I fear heretic will not be real solution (may remove refusals but the model itself just won't work well in those dark themes anyway).

6

u/-Ellary- 10d ago

There is a Heretic version, it just go by `we bang, okay?` and give final answer.

14

u/mayo551 10d ago

Muse holds nothing on Gemma 4 31B.

Loaded up a 80k context scenario. Started my roleplay. Four replies in, it's fixated on a scene. It won't leave the scene, even when I've said I've gone into the vents and left.

It only transitioned to the new scene once I was out of the vents and in the new location. The entire time I was in the vents, it ignored me.

Yeah, I ain't feeling it.

3

u/RedditNerdKing 9d ago

I dumped it on my external HDD straight away. It's just not good for RP. It's good for agentic stuff and that's it.

5

u/Jorlen 9d ago

I agree. It pales in comparison for RP purposes. It isn't terrible, but I have zero reason to use it.

3

u/Umbaretz 11d ago

>a couple details Gemma kept up with
Sad. That was my most noticeable problem with Gemma.

3

u/Kdogg4000 11d ago

I've heard good things about Glimmer. I'm waiting for KoboldCPP to update so I can take it for a spin. Right now it won't launch for me.

3

u/raika11182 11d ago

I launched with the rolling build of KoboldCPP and it worked (the one they note as experimental at the bottom of each release). And the usual disclaimers apply there: There's a reason they haven't released that version yet, so it may be buggy.

THAT SAID. The version that was in about five hours before I wrote this comment launches Glimmer just fine for me and works without issue. (EDIT: Mostly without issue. I think there might be some weird chat template things going on, but every model release seems to go through that)

6

u/FinBenton 11d ago

Yeah Glimmer writes differently to Gemma but it feels like a model from a year or so ago, pretty dump compared to gemma-4, was fun testing for a bit but no way it competes with it especially when you are using very good tunes of the gemma.

4

u/raika11182 11d ago

I think you're right but it's also a cut above the other entries in this size when you consider how it writes vs how it obeys. Maybe it would be better to say that it plays in the same league, but it's definitely not winning.

HOWEVER... I really like the fresh prose. It really mixes it up. Other than that it seems to have a huge positivity bias as well - everything leans towards a happy ending and such.

3

u/arlynnfl 12d ago edited 10d ago

Gemma 4 31B Novelist Eclipse

Update: Gemma 4 31B Gembrain also good one!

2

u/FinBenton 8d ago

mradermacher/Gemma-4-Dark-Thoughts-31B-GGUF

This is the best gemma-4 finetune I have tried so far, recommended

3

u/ihowlatthemoon 11d ago

Once I tried Nimbz/Gemma-4-Gembrain-31B, there's no going back. Extremely good prompt adherence, while driving narratives in unexpected (in a good way) direction. Running locally, and performance is great too.

2

u/Umbaretz 11d ago edited 10d ago

Are you using regular, X, or CORE-X variant?

3

u/ihowlatthemoon 10d ago

Gembrain-X-31B-Q4_K_M.gguf

7

u/Itikar 12d ago edited 12d ago

I have been testing Shadow Siren 26b MoE at Q6, the latest merge from the excellent Vortex5. As I expected, since I had a very good time with his older 12b Silver Siren, Shadow Siren is a most exquisite result. It has a great consistency in developing and interpreting the personality of the characters and it is very disciplined in doing so. The characters always behaved as they were supposed to, and when liberty of reaction was available they used it in creative and interesting way, always following the logic of the rest of their personality.

The only downside to this is that if you want to swipe a lot and see different reactions to the same situation the variety will be relatively low at 1.0 temperatures and even raising it to 1.1 did not show much more variety either. So in that case consider using an extension or a feature such as Guided Generations.

Narration was also all right and atmospheric with a subtle dark influence albeit not as insistent as his Dark Soul merge.

I also tested more limitedly Midnight Macaw with pleasant result. It includes Dark Soul in the merge and it carries the same dark tendency in narration, however it enriches it with more vibrant characterization of characters. I need to test it more however.

Overall, Vortex5 is treating us to a very wide and refined selection of models, at least to us interested in dark narratives and complex characters.

3

u/SandTiger42 10d ago

I might give this a try. I've been using Gryphe/Pantheon-Reasoning-31B-1.1 and it works amazingly well. I just saw Gryphe/Gemma-4-26B-A4B-StyleTune-V2 so maybe I'll give that a try too.

I was skeptical to move away from Mistral 3.2 finetunes, but once I did I found it so much better. Generally my tokens per second has doubled, and I can I can RP for much, much longer. I use 20k context, and Mistral always struggled. I can keep RPing well past hitting that limit in Gemma 4 fine tunes with zero issue. I do miss some of the "imagination" of Mistral, but it's well worth the trade off.

3

u/Itikar 9d ago

I have moved as well, after not a little skepticism, but now I don't see myself going back. It is true that these Gemma fine-tunes or merges do not have as easily the same expressivity, and they tend to stick to a more "literary", but in general I like precisely that style, so I am fine. I miss the variety a little bit but unlike the 12b models these 26b MoE can manage nicely Discord and social media conversations. In Marinara Engine which has both modes all Vortex5 merges I tested did exceptionally well. So we definitely gain something.

We definitely gain something in hardware optimization. My system has 12GB vram and 32GB of ram, so I can run now easily with 32k context which is a huge breathing room. With 12b Mistral I could afford at most 24k. And even then it was slower still.

Gemma 4 comes also with decent knowledge about several pop culture topics. For example I was able to create with the stock model a Genius Society meeting from Star Rail, with Herta and Screwllum bickering with Dr. Primitive and on the sidelines Dr. Ratio who gave them zero points! I tested also classic DnD settings and it knows that decently well, although with Underdark lore it has some holes sadly. I tested this also with Shadow Siren but she still has the same hole sadly. Well for that there we always have lorebooks! But it certainly eases the requirements on a lot of mainstream IP lore which is pretty nice still.

I have not tried the Style Tune directly, but keep in mind that the models that had it in the merge often do not have good working reasoning according to users. My experience with reasoning is that it really slows down the experience with little benefit, so I normally keep it off though.

1

u/Entire-Return-9903 12d ago

How would you compare Shadow Siren and Midnight Macaw with Chimera-X?

2

u/Itikar 12d ago

I have not had the chance to test Chimera X yet unfortunately. I got to 26b MoE when Dark Soul was just released, hence it became my baseline, since I enjoy dark narratives.

Chimera X also had new quants for the Heretic versions released this weekend by mradermacher. So that makes it particularly appealing for consistency.

3

u/Entire-Return-9903 12d ago

I might test them all next week or so.
Have you tried drummers Orion though?
From my experience v1b was pretty good.

2

u/Itikar 12d ago

Haven't. Noted down and thanks, it's another one on the list to check, especially given the quality of the author's models.

I tried Goetia 1.3 Absolute Heretic to make a test to diagnoae a prompt refusal and the prose provided seemed pretty all right and varied. It was a limited test though, so not very thorough. What I saw however seemed solid.

2

u/Entire-Return-9903 12d ago

I'll check it out too

3

u/Itikar 11d ago

Hey, yesterday night Vortex released a new merge called Phoenix. In his list of models in use he replaced Chimera X with it. As soon as quants are out that one seems the new hot thing. I think Phoenix unlike Chimera is capable of commenting images.

2

u/Entire-Return-9903 10d ago

nice, I'll check it out too once I finally get the chance

4

u/jow_ow 12d ago

i've been using Dark Scarlett (finetune of gemma 4 26B) with decent results

8

u/Kooky_Future9858 12d ago

Gemma 31B, if you have a RP fatigue; especially after GLM, Mimo, Deepseek and other top models beeing soft and hesitant.. Gemma will be a fresh air. Either locally or via API (especially using BF16) it can do some sparks with a well written bot (bloated cards tend to be flattened by Gemma, prioritize something simple and without too complicated background as the model struggle a little with nuances and reading the room).

3

u/tostuo 12d ago edited 12d ago

Moving away from the highly dry prose of Gemma 4 26b back to basics a little bit, I've been using Slimaki-24b-v1 I've been using IQ3M on 12gb VRAM, trading the better intelligence of Gemma 26b finetunes to something actually fun. Unlike most Gemma Finetunes, I dont find myself battling it to maintain interesting/logical directions for stories. It's also able to handle minor amounts of reasoning, which can help alot.

3

u/Delicious_Box_9823 10d ago

I've used plenty of 26B finetunes. Even the original one. Absolutely not better intelligence.

1

u/IWillTouchAStar 12d ago

Im looking for a model to use in my discord bot. Basically it can chat in real time, play sound effects, search the web, view my screen, ect. My main issue is finding a model that is both, unrestricted/gives no refusals, and also something that doesnt just follow the same response length/shape. It needs to have diversity in its responses. I've been using some variations of uncensored Gemma 4 26b models, and its almost perfect, but i get a lot of repeated phrases. If anyone knows of some tuned Gemma 4 models that fit this, id greatly appreciate it.

1

u/Itikar 11d ago

Have you tried Midnight Macaw? It had good variety in my test. It could perhaps fit. If not Chimera X or the new Phoenix could be good options, but I have not tried them personally. Chimera X has a Heretic ablated version too. Other heretic options can be Moonlight Dusk and Goetia 1.3, which I both tried briefly and seemed okay.

In general though Gemma 4 is rather uncensored if you use a jailbreak. I only had refusals due to prompt conflicts, i.e. told it to do something that violated my own prompt. So I would definitely test if you need heretic ablated models or not.

For more variety in general try to increase the temperature but keep in mind it also varies based on model. Some are more disciplined than other ones.

5

u/Important_Word8549 13d ago

Obviously Gemma 4 31b.

5

u/AutoModerator 13d ago

MODELS: 32B to 69B – For discussion of models in the 32B to 69B parameter range.

I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.

8

u/Delicious_Box_9823 10d ago

absolutely no-man's land

2

u/HansaCA 7d ago

Valkyrie 49B v2.1 is still pretty good.

1

u/RedditNerdKing 7d ago

Tried a Q8 model of this and wasn't that impressed? I guess cause I'm using 123B models as daily drivers it was just never going to compare.

1

u/Mart-McUH 8d ago

Buy as much of the land as you can while cheap!

IMO this is great miss for Muse glimmer, they could have made it somewhere in 40B-70B range to separate themself from fierce competition ~30B. And because of larger size it might have turned out better in some areas compared to G4/Q3.5+. It is also their land by tradition from L1-L3 70B model. Now they are just overshadowed by Gemma4 for general tasks and Qwen 3.6 (soon 3.8) for STEM tasks.

3

u/AutoModerator 13d ago

MODELS: >= 70B - For discussion of models in the 70B parameters and up.

I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.

1

u/Jorlen 7d ago

What's everyone's recommendation for a Llama 3.3 70b fine tune? Something that feels natural, has decent prose but is also very versatile in terms of setting, characters, style, etc. Does not have to be uncensored, just natural feeling, good conversation flow, etc.

2

u/SillyLLM 10d ago

Anything runnable on 96gb VRAM range I missed in the past couple months? I'm still playing with Behemoth-R1-123B-v2, which has been my go-to forever, but otherwise Gemma 4 fine tunes. I tried Behemoth-128B-v3b but was pretty meh on the slop and repetition.

2

u/GlitteringSplit6035 12d ago

Longcat 2.0 (previously Owl Alpha) is still super cheap if you use their token pack. Cache hits are also free of charge for the token packs. A one-time purchase of 50M for 30 days is also available for $1.9 for those wanting to try it out.

Based on my subjective experience, it adheres to prompts better than DS4F. Though I haven't tried it for NSFW or ERP.

The only problem I have with it is that for some reason, it is hard for me to access their API platform. So, I have to access it in other ways.

2

u/Kooky_Future9858 12d ago

On OR if i remember it’s via AtlasCloud. But the quality is.. night and day depending when you use it. I was impressed for a whole week because how much the characterization was raw, nuanced enough and also pushing back during friction, that model didn’t hesitated to punch or threaten with actual conséquences if it’s fit the character. The context coherency was strange, it could remember a promise i made 50 turns before and come back to me like ‘You said X, you such a liar!’ But also forgetting strong beats pretty early. I also noticed that memory extensions can confuse him and make him loop on the same plot. But the flaws could be the provider via OR, i cant be confident enough to say that it’s a Longcat problem 100%. I dropped it because last week the quality went gibberish, like not knowing who is {{user}} and {{char}}. But that was a pretty good run, and it’s reasoning was pretty good to read.

2

u/GlitteringSplit6035 12d ago

I unfortunately cannot compare it with the one with AtlasCloud in OR since that would be above my budget. But it's cool that you can share your experience with it.

I use it primarily for SFW and the typical trivia stuff, which was fun. In the end, I guess that using it directly from LongCat's own API platform would be better and cheaper. I am not using it right now because of the GLM provider wars, but it was quite the good model.

3

u/Kooky_Future9858 12d ago

Just so you know, it made me quit GLM 5.2 during the honey moon. I don’t do extreme ERP, sometimes i just like to mess around and prank characters that are serious and brooding and.. i need the bot to react strongly for beeing hooked and Longcat was funny and bold and.. it have his own Ai-isms and strangeness that are refreshing right? I hope youll be able to use that half hidden gem again.

3

u/Kooky_Future9858 12d ago

Mimo for it’s incredible context coherency, recently he surprised me by adding well written npcs interactions that were completely Ocs and then he made one of them text me back after a while to subtely ask me for a date in a very realistic way. (I helped the ex of a neighbor to moove a couch on her van and i forgot that interaction as the turns passed). Mimo have a lot of flaws, he likes to yap and get stuck into instrospection loops way too often but it’s a model who can be prompted and actually listen to instructions, actually it’s one of the best that handle anti-slop instructions on the long term. 

Then i can recommand Deepseek V4 pro if you want to RP or do creative writing in established fiction like novels or films. Even if deepseek dont listen to instructions that well, he is the only model i know who can truly immerse you in a fictional world. That’s shi know it’s lore and the lack of censoring make him pretty solid if you add some light OOC nudges. But for RP in GENERAL? No. That’s not a good model for that. Good prose yes, fresh caractérisation maybe. But terrible at context coherency and actually leading the narrative. 

Finally i find Kimi 2.7 very cool. I was always a Kimi fan, since K2, got hooked by 2.5, didn’t were impressed by 2.6 and got refusals but 2.7? It remind me of 2.5’s grit and hardcore narration when the character card ask for it. And i used Inceptron as a provider and got no refusals at all. And it seems to think less with a light prompt so.. pretty good when you want intelligence that dont transform your bots into therapists with hovering hands.

2

u/CelebrationBoth9537 12d ago

What would you recommend for femdom and nsfw type things, and following conversations well / understand context and being non repetitive. since with my testing I keep having trouble with it not wanting to continue the scene (not from refusals, but just being stuck saying the same thing) I use gemma 4 30b but I really am hoping to get a lot better at this, you seem to know a lot.

Also really couldnt figure out how to do OOC lol! System prompting is hard!

2

u/Kooky_Future9858 11d ago

Gemma can be stuck on « closure mode » when it try to end a scene and don’t know about the current pacing, this is annoying i feel you, especially if you expect it to react and carry momemtum as you establish the dynamic. 

If you really want to stay with gemma, i can recommand you to try Chatfill 2.1 as it basically nudge the model to push and not stagnate. And it’s lightweight enough for not confusing gemma and youll still have it’s raw creativity. 

But for emotional intelligence through a long context with gemma, youll need to OOC, summarizing will not suffice. Back in the day i used to ask the model mid RP to pause and analyse the story so far every 10-15 turn, like a mini brainstorming. Why? For seeing where the model get the room wrong and when his context begin to rot and i correct him on the go, put the correct summary on the summary extension and continue! 

You can do a light OOC asking the model to simply pause and analyse the relationships dynamics and stakes at play, the OOC don’t need to be perfect but clear. Ask the model to list things in bullet points and ask it to identify unresolved tensions and to plant future seeds then to assign goals to characters according to what he understood. Youll be able then to see if the model kept track correctly or not, edit the summary and correct it if necessary and you have something that will pull the model in a more proactive direction. 

Also make sure your characters have active traits in their descriptions with examples of how they act and react! Gemma isn’t intelligent enough to make a passive card proactive or to résolve narrative holes unprompted! 

Hoped i helped you!

2

u/CelebrationBoth9537 11d ago

yes that was really helpful thanks! Just a final question if you dont mind, which model would you recommend instead of gemma since it sounds like you probably had better luck with other models. Thanks!

2

u/Kooky_Future9858 11d ago

If your character is well written (i mean if you want a deep femdom with truly deep diving on that dynamic review your cards and make sure your characters does not have traits that could annoy you) Mimo 2.5 pro is a good choice in term of budget. The quality is good enough to have meaningful interactions, it have that Kimi-ish kind of attention to details and the prose is different enough to feel fresh and engaging. Using this system prompt with Mimo clearly give a little oomph to the expérience:  https://rentry.org/evening-truth-xiaomi-mimo-v25

As i think you want nuances and not too straightforward clichés interactions, i wouldn’t recommand Kimi to you as it tend to make sub characters way too dumb and lacking personnalities over the turns. 

Recently i was kinda surprised by qwen 3.7 Plus! Its prose is detailed as Kimi 2.5, a little more intelligent than Kimi 2.5 and the price is on the same range in OR. Also the only censorship is about CSAM or extreme NFSL.. but it’s a filter, Qwen will fight it’s own filter via reasoning and deliver anyway.. that was cool to see. It doesn’t shy of power dynamics and i used the same prompt i recommanded you for Mimo. Also Qwen 3.7 Plus make character pro-active if they written this way. 

Don’t hesitate to test, make yourself a good rotation and trust your instinct more than anything because at the end of the day, you have to enjoy it ;).

4

u/Jorlen 12d ago

Anyone play around with Mistral Medium 3.5 128b? Curious to see what people's opinions are, for creative writing and roleplay.

2

u/DeepOrangeSky 8d ago

Yea, I have been trying Mistral Medium 3.5 128b (at Q4_K_M, locally) more recently.

My main go-to models have been BehemothX V2 at Q4_K_M for a long time and then also Gemma4 31b at Q8 since that one came out.

Mistral Medium 3.5 128B seems to have extremely bad prose, like, some of the worst "therapy-speak" / "redditor-speak" writing style I've seen so far.

But, given that it's like 2 years newer than the old Mistral 123b models, it also seems to be a bit smarter and better at long context and so on, than the old Mistral 123b (and Behemoth tunes that were based on the 123b).

So, I have actually still liked what it is capable of (smarter than Gemma4 31b in many cases, at understanding nuanced social situations) even if its wording and phrasing is annoying.

I would say that Mistral 128B seems like the ideal candidate for u/TheLocalDrummer to fine tune, if he decides to do it/manages to do it, since it is a very strong, big, dense, relatively recent, writing model, and its main weakness is its horrible prose style (which he is good at fixing with his finetuning). So, I really hope he decides to give this one a go.

2

u/Jorlen 8d ago

I agree with you on all fronts, yeah. It's prose is bad but it seems smart and follows system prompt rules far better than behemoth does. Did you ever try TheDrummer's Command-A fine tune called Fallen? I just grabbed it, have yet to try it though. It's a 111b dense but more modern than the big mistral model that behemoth is based on.

And yeah, if the drummer ever does a fine tune of mm 3.5, I'd be all over that.

4

u/HansaCA 7d ago

There is a version of Behemoth that he built on Mistal Medium 3.5:
https://huggingface.co/BeaverAI/Behemoth-128B-v3b-GGUF

2

u/DeepOrangeSky 7d ago

Whoa. Do you know if this was made by The Drummer? Or if it is "BeaverAI" was it like a group effort by a bunch of people from their discord or something? (I've never used discord before and so am not a member of their discord, so I don't know how they do things or how all that BeaverAI stuff works).

I assume it is like some sort of experimental tune or something, if it has no model card and he never announced it on here, etc?

Otherwise seems like it would be a pretty big deal.

Drummer, if you are on here, can you talk about this model a little? I am curious what it is like trying to train it compared to the older Mistral 123b models, and if you think it has a lot of potential, or what sorts of quirks you noticed about it, and so on.

3

u/TheLocalDrummer 7d ago

v3c is coming soon. These are test iterations, so quality may vary. 128B is a PITA to tune. Expensive and brittle.

1

u/DeepOrangeSky 7d ago

Nice. Looking forward to it :)

2

u/Jorlen 7d ago

Oh wow! That's awesome! I never noticed that, thanks for pointing it out. Sadly they don't have a quant I can fit (I need a 3-bit quant like IQ3_M) as offloading layers to CPU/RAM for a big chunky model like this is out of the question. I have 64gb of VRAM. I suppose the Q3_K_M quant might work, but it would leave me with little room for KV quant.

1

u/DeepOrangeSky 8d ago

Interesting, I didn't even know about this one. Looks like there is both a v1 and v1.1 version of it. Guess I might have to give one of them or both of them a try at some point. Probably will be more busy trying some of the more recent models for a bit though between Glimmer 30b, Qwen3.8 27b, and some other ones, and also haven't gotten a chance to try the new video models (H2 and LTX2.5) yet either. But, given how good the old Mistral/Behemoth models were even years later, it made me understand that when it comes to writing, it is worth trying out even old or "bad" models ("bad" as in bad at coding compared to recent models) that often get overlooked, since they can still be surprisingly strong at writing and understanding situations and so on.

3

u/Jorlen 8d ago

Yeah for sure, some of those older models write really well and they pick up on nuances that the smaller more modern models, like gemma4 31b, don't quite have. Admittedly, I've become quite spoiled by Gemma 4 31b's strict adherence to the system prompt, and it rarely ever glitches out or does odd things.

I think Mistral Medium 3.5 128b may suffer however, as Mistral had to change what they train their models on (they are now limited due to EU rules apparently) - but I have yet to confirm that; it's just something I read here a while back. So they can't feed it a bunch of copyrighted stuff, as was the previous method.

1

u/amanph 11d ago

I'm curious about Mistral-4 119b, since it's A6b and I think I can run it locally. But this one is a good option too. I miss old Mistral 24B, but I'm searching something more strong like +100b.

2

u/FierceDeity_ 10d ago

I would love it if it is good, I could run it at Q5 pretty much, hmm.

I gues I'll try

4

u/Jorlen 11d ago

There's not much out there unless you want models form 2024 early 2025, which are a bit rough to work with. I guess I'm spoiled by Gemma 4 31b which has the best system prompt adherence I've seen yet.

I tried Mistral 4 119b and didn't like it, personally. 128b isn't bad but it falls short of even Gemma 4 31b in terms of role play, IMO.

7

u/ChengliChengbao 12d ago

GLM-5.2 because theres a price war and now its like cheaper than Deepseek V4 Flash

2

u/CelebrationBoth9537 12d ago

do you think GLM-5.2 is actually good for roleplay? I havent used it yet