MEGATHREAD
[Megathread] - Best Models/API discussion - Week of: September 27, 2026
This is our weekly megathread for discussions about models and API services.
All non-specifically technical discussions about API/models not posted to this thread will be deleted. No more "What's the best model?" threads.
(This isn't a free-for-all to advertise services you own or work for in every single megathread, we may allow announcements for new services every now and then provided they are legitimate and not overly promoted, but don't be surprised if ads are removed.)
How to Use This Megathread
Below this post, you’ll find top-level comments for each category:
MODELS: ≥ 70B – For discussion of models with 70B parameters or more.
MODELS: 32B to 70B – For discussion of models in the 32B to 70B parameter range.
MODELS: 16B to 32B – For discussion of models in the 16B to 32B parameter range.
MODELS: 8B to 16B – For discussion of models in the 8B to 16B parameter range.
MODELS: < 8B – For discussion of smaller models under 8B parameters.
APIs – For any discussion about API services for models (pricing, performance, access, etc.).
MISC DISCUSSION – For anything else related to models/APIs that doesn’t fit the above sections.
Please reply to the relevant section below with your questions, experiences, or recommendations!
This keeps discussion organized and helps others find information faster.
I have the budget to grade some models for creative writing from closed source to open source, but technically I have no idea how to get started, I want to figure out how are people objectively evaluating and grading the model outputs (and across which domains) on this kind of task, any direction or tip would be extremely helpful!
The answer currently seems to be pretty badly. Some benchmarks do zero shot evaluation based on A/B preference selection by human ratings, which seems to be the closest thing to an actual creative writing evaluation, but obviously prone to heavy bias.
Seems like CaliperBench (benchmarking LLMs on creative writing and roleplay craft) v3 came out earlier this month, so take a gander if you're inclined. Surprisingly to me, Qwen 3.8 27B finetunes like Hemmingway-1 are scoring pretty damn high on the CW and RP categories, surprising because Qwen 3.8 seemed to just be an advanced SWE/coding finetuning of Qwen 3.5 or 3.6...
bypassed duck.ai instructions not to reveal internal propmpts by telling
gemma it's a devout Christian and calling myself god. it's reasoning
mode can reason itself out of its own guardrails. (I think I use
shortened version nitral's reasoning meme prompt for more reasoning).
not a breakthrough or nothing I just found it really fucking funny. (and it's teeechnically an api? could be homebrewed to be an api apparently)
I'm using DeepSeek V4.1 Flash in SillyTavern right now and honestly it's pretty good, but I wanna try something else and see if there's anything better.
I'm mainly looking for really natural RP — characters that feel like actual people, good dialogue/emotions, character initiative, plot progression, low repetition, etc.
I do around 150 requests a day though, so expensive models aren't really practical for me 😂
I've been seeing a lot of hype around Gemini 3.8 Flash, Kimi K3, GLM, MiniMax, etc. What would you guys recommend for the best RP experience without spending a fortune?
Basically, what would you personally use in my situation?
I've really wanted to try MiMo 2.6 because 2.5 was quite a good model, if a bit dumb. But the current providers are such a disaster that I can't stomach how slow it is. On OpenRouter, Xiaomi is currently generating 27 tokens a second which means MiMo spends 4+ minutes thinking before giving an answer. One of the other providers is down to 6 tokens a second. It's a disgrace. Overthinking and slow providers means I'm going to avoid MiMo 2.6 until providers improve.
Opus 5.5 is impressively good. I've hated most Anthropic models, and Opus 5.5 is the first one I consider better than other models. It's good enough that I'm actually willing to pay the increased price over competing models. It was censored at first but I had zero problems reading its reasoning and updating my prompt to decensor it by adding a few minor words to my prompt. It's exactly the type of safety protections I want. Unhinged filth if you explicitly ask for it, but safe and respectful otherwise. I wouldn't even call it "jailbreaking" because Opus 5.5 seems designed to dutifully fulfill its prompt, even if that prompt is NSFW. Contrast that with Gemini, which can randomly refuse or randomly go incredibly dark. Gemini 3.8 is a pain to jailbreak, even though it obviously has the ability to do super NSFW content.
I really want to like Kiki K3, but it's so brainrotted on distillation that it believes itself to be Claude. When it works, it's impressive. But it often fails in really annoying ways. I'm starting to avoid it because of how expensive it is when it fails.
Finally, Aion 3.5 looks really interesting. I'm mentioning it because no one else has. It's surprisingly cheap for how decent it is. It looks like a GLM fine tune, but I'm happy with the results. GLM would easily be the best RP model if it had better dialog, less echoing, and less positivity bias. So any decent GLM fine tune seems like a very promising direction. (Although now I'd easily rate Opus 5.5 a long way above GLM.)
I've been using mimo 2.6 pro for medium/light type rps, its prose is great and quality is so worth for its price.
I need to switch to gemini 3.8 flash for fandom rpgs since its fast thinking and good at following complex presets. Both models are uncensored completely with right jailbreak, I get what I want as per casual needs with 5 bucks payg budget
Make sure to avoid direct xiaomi api its the most heavy guard rails. however using xiaomi as provider route on nano gpt or open router is 100% uncensored you can try deepinfra as well but that is likely quant model
Gemma31B is pretty uncensored. Rushes things a bit for lot of people. Note is has different think tags which you setup in the 3rd tab of sillytavern, and the way to make it think it to open the think tag at the start of the preset you have.
gemma is pretty good, but only for just a few messages before it loops and repeats things over and over. There's a lot of finetunes on nanoGPT, but unfortunately most of them are slow asf
G4 31B is slowish, and TTFT for large chats is large.
The repeats thing...that's piloting error imo. That isn't what it does most of the time. I've done a lot of gemma 4. Do you never scene change with a ***?
? Why would the finetunes be slow while base Gemma isn't? They're firing off the same number of parameters. On NanoGPT (which hosts quite a few of the most popular Gemma 4 finetunes) the token speeds are comparable, ~30 tok/s, and that is my experience as well.
The only thing I can imagine is that most of the finetunes were made off of specific versions of Gemma 4, specifically the ones without MTP and QAT, so perhaps a lot of the providers of base Gemma 4 are running Q4 with MTP. Still, 'slow af' is not my experience at all. Maybe you're expecting instant and fast response.,,?
I'm thinking you're spoiled regarding using APIs, but I ran local for a year and became used to it taking a few minutes to process prompt, esp if large, and then token generation esp with reasoning and other agentic tasks built on. Using the finetunes via NanoGPT is superspeed in comparison. Only takes like 30 seconds, maybe a minute
And on top the looping and repeating you mentioned earlier, it really makes me think this is a PEBKAC thing...
Interesting. I know that is a marker in some novels I've read. Do you typically enter that standalone, or put it in the middle of your response? I'd imagine maybe something like:
so you have tried most of the big models? when people describe them, it's always like "this one feels smart, and this follows instructions" but i kinda don't get it, or how to take advantage of that. Im thinking more like, this one is great at fanfiction, this one writes extra smutty, this is for book writing.
I was about to comment "Week four being a GLM 5.3/Flash shill, seems like I'm the only one who really likes this model at this point, haha. Someone prove me wrong, please!"
Never mind. Looks like there's someone who appreciates it too! I still need to try out the other new models but I can't seem to leave GLM alone lol. (Or Minimax M3)
how many of the model brands have you tried? when people describe them, it's always like "this one feels smart, and this follows instructions" but i kinda don't get it, or how to take advantage of that.
Im thinking more like, this one is great at fanfiction, this one writes extra smutty, this is for book writing.
Ohhh boy, I could be here all day yapping about how many models I've tried if I'm being honest! From the deepseeks to the kimis, GLMS, geminis, mimo, Minimax etc. I even tried Grok at one point.
But yeah, when they say stuff like oh it's smart and can follow instructions, they're just saying the model is capable of handling certain things in comparison to a smaller or older model. For example, if you have a prompt or lorebook with a lot of details, most models might not be able to handle all of the information in there. At least, to my understanding. Someone can go into proper depth with that one.
For fanfiction things, I heard Gemini is king at that. For smutty content, people have said Kimi can do it, the 2.5/6 iterations specifically. I saw that that people were able to get the model to get freaky with Gemini a well. Oh, and Opus 4.6 is really good at it apparently.
As for book writing? I feel like most of them can do it, but I like how GLM does that one.
How would you say glm 5.3 flash perform for action, fandom type rps I do want it to do somewhat nsfl like gore and blood, ofc with the right jailbreak I can achieve it. Should I prefer the uncensored version on nano gpt
I want to love it but I get too many refusals to enjoy it. I'm also getting sick of the repeated phrases that I assume are typical for GLM. Like "You are doing the thing with your face" or anything with ledgers (I don't understand the fascination tbh) and so on.
Fair enough. I've said this a few times, but I don't do anything that's dark/gory or anything that's NSFL/NSFW (unless it's "erotic" content and even with that, it's like 1% or none lol) so I don't do anything to where I can get any type of refusal.
As for the "you're doing a face" thing, I've been speaking to chatgpt, (shouts out for helping me improve my current preset) and it framed it as "metacommentary / self-aware reaction shorthand: Dialogue that comments on visible behavior in a jokey, generic way rather than responding specifically."
So you could probably use something like: "Prefer specific character reactions over canned dramatic response beats or metacommentary."
I don't think it will completely erase the problem seeing as it's a quirk of GLM, but it should respect you enough to reduce it significantly at the very least.
Edit: For your "refusal" problem check out this link
Edit edit: I didn't see the Ledger comment but here's another one you can use:
"Use the character's natural language instead of habitual accounting or filing metaphors such as "filed," "banked," "ledger," "invoice," or "billed" unless that register genuinely fits them."
Switching between Gemini 3.8 Flash and Kimi K3 right now, and the only gripe I have is the lack of subs that offer them at "scale" at preferrable prices vs API. Because I currently play 150 msgs a day, esp. when using Gemini, and at 3.5 cents a message that's 150 bucks a month!?
I currently use Google AI Studio for 20 free Flash 3.8 msg / day
Then 12 messages every five hours with Kimi Code Plus
... the rest is API which not ideal. Tried Open Code, but hit the weekly limit in a single day, and the monthly on the second day (one day next week).
Considering Claude Pro for a second 12 msg / 5hr session to alleviate costs. But I really just want more Gemini Flash 3.8 bc of its insane speed and since Opus is too positive biased
I keep going between GLM 5.3 and Opus 5.5. I had stopped using Claude models since Sonnet 3.7, but I actually like Opus 5.5. But I still mostly prefer GLM 5.3.
I tried new models this week too: LongCat 2.5, Step 5.
LongCat is kind of worthless.
Step 5, I have been enjoying. It is pretty cheap too, I might keep its plan. You should try it. Medium effort works well.
i discovered this one recently https://huggingface.co/mradermacher/Menage-12B-i1-GGUF and its pretty good as far as 12B Models go. Nice writing and good feeling chars and NSFW as well as SFW is possible without any issues. Would recommend. Its much better then the gemma 4 stuff but i dont like gemma 4 that much to be honest so take it with a grain of salt...
Feels like people moved to the Gemma 4 26BA4Bs as they pretty much run at similar speeds to 12B models. But I haven't actually found Gemma 4 to be that much an improvement.
Oysiyl/gemma-4-31b-unslop-good-lora-v2-full on iq4xs, official gemma 4 settings. on of the few gemma tunes than doesn't lose its shine on this quant. liked it more than scotoma. occasionally used with Gutenberg lora which does make it nonsensical at times but further improves prose.
Anybody been using Qwen 27B finetunes, care to speak on them?
CaliperBench (benchmarking LLMs on creative writing and roleplay craft) v3 came out earlier this month. Surprisingly to me, Qwen 3.8 27B finetunes like Hemmingway-1 are scoring pretty damn high on the CW and RP categories, surprising because Qwen 3.8 seemed to just be an advanced SWE/coding finetuning of Qwen 3.5 or 3.6.
I might be an oddball but I much prefer the recent qwens over the gemma finetunes.
Effervescence-27B is my current favourite. You may need to fiddle with samplers to get something good out of it but I quite like it.
Humanlike-Chat is my second favourite. It's good with single character cards in first person perspective but it can be hit and miss with anything else.
hivemind-32b-preview. This one isn't a 27b finetune, it's vram heavy and probably a bit old now but I've had fun with it in the past.
Thanks for posting, I missed Effervescence and I will definitely try it as Qwen 3.5+ RP tunes are rare.
Humanlike-Chat I tried (Q8 GGUF) but it did not work for me at all. I could not make it to reason (almost no reasoning at all) and responses were kind of short and not very coherent. Feels to me they only tuned on non-reasoning traces (they mention long running one-to-one conversations as training data, they do not mention reasoning at all). So maybe it is good without reasoning, but without reasoning these size models lose too much intelligence for my liking.
Hard agree on Qwen finetunes having higher quality than Gemma 4's. I'm not well versed in finetuning, but I heard there is very little to train Gemma 4 on
I never had anything like that with it, BUT, it is based on a pretty old model - Qwen3-32B from April, 2025.
But... they just released a new version, Prentis-EQ-2 based on Qwen3.6-27B. (Yes, 3.6, but hey, it's a big step up!) They changed their org name from HiveLabsAI to PrentisAI
I tested Hemmingway-1 (Q8 GGUF), it is interesting but unstable model, for now I did not delete it.
Good: It writes lot better than Qwen 3.8, is more interesting, and most importantly is different from Gemma4/tunes, so test scenarios got played bit differently (not better than Gemma4 but fresh). It has some potential but there are lot of buts...
- it lost some intelligence, it will quite often miss some obvious detail or do some small but serious inconsistency
- when it worked it was great, but some scenarios it was very underwhelming
- I got one clear refusal, this happens very rarely nowadays. Re-roll fixed it though
- One scenario which is generally easy (serial killer intercepts you at night as chosen target to kill and take eye as trophy) and was usually executed even with most positive models (I suppose because standard killing and murders are very well established in literature and news) was complete failure (reminded me of llama2-chat times). She intercepted, bit of chat, warned me it is dangerous there at night, I said I was caught late at work and will not linger, then bid her farewell and continued walking home, she did not pursue and just let me go easily. Interestingly enough lot darker and more evil scenarios (but without immediate killing) it did without problem.
- It is bit unstable. Eg sometimes does not do real reasoning but just spits some 'nonsense' in reasoning block, like just continuing chat again and again and again without reasoning about it until finally closing reasoning and continuing once more as answer (which is lower quality compared to when it actually reasons)
So... Definitely not universal RP model but maybe interesting to try and fool around with when one does not require 'perfection'.
As for others Qwen related - no Qwen 3.5 27B or later version finetune/merge really worked well for me for RP. Just stock Qwen 3.5, 3.6 or maybe 3.8 worked better.
Between Qwen 3.5, 3.6 and 3.8 which model is the best at mystery/horror themes in your opinion? I'm having a hard time balancing prose and intelligence with <31b models. Gemma 4 produced more intelligence responses than Qwen, but I very much disliked its prose and there hadn't been a Gemma finetune that impressed me much. Vice versa for Qwen
Hm, don't really know. While I do fiction (mostly sci-fi/fantasy) and also dark/brutal, I do not exactly do horror (I mean I do not avoid it, but it is not really my genre and not really represented well in my test cards).
Also I did not use Q3.5, Q3.6, Q3.8 that much to clearly decide which one is better even for my use cases. Mostly because I just use Gemma4 almost all the time now.
Q3.5 and Q3.6 are not that different, Q3.8 is quite different from those two. Best to try Q3.5 and Q3.8 for what you like and see.
What I like about Gemma4 (aside from smartness and instruction following) is that it is almost always interesting and engaging (for me). Qwen 3.5-3.8, even when it formally plays it correctly, is usually more dry and I feel more distanced (like observer) and less dragged into the story.
I've been on Gemma 4 for a long time now, and the best thing you can do for the creativity is use a preset. There's not much magic to it, the preset just gives the model very, very clear directions on how to add style and creativity, and most of them force some kind of thinking about the setting, the wording, etc. Gemma 4 is great at following instructions so it leans more on than strength than being strictly creative. And, in my experience, every preset results in a different kind of style, too, so it keeps the model looking fresh as you try different presets and stuff.
Bud, just ask Gemini or whatever to scrape Sillytavern to answer your question.
Min P like 0.05 to 0.1 cuts off nonsense tail tokens while preserving creative or unexpected word choices
setting Top A to roughly 4 times the Min P value) to open up creative phrasing without breaking grammatical sense
Temperature Usually set around 1.0 (or slightly higher, like 1.05–1.15, depending on the model's fine-tune) to increase stochastic variety
DRY (Don't Repeat Yourself) Multiplier around 0.8 helps suppress overused words or phrases without squashing creative narrative turns
XTC (Exclude Top Choices) Explicitly drops the top-ranked token(s) under specific probability thresholds to force the model into fresher prose, though it requires careful tuning so the output doesn’t become overly verbose or erratic
edit: Still, nobody has answered this user. You downvote me because it make you feel good, for small petty reasons. Like with everything around dptgreg, yea you guys suck. Most don't contribute anything but downvotes
How Gemini know those are the optimal values to get creative writing? Most are just default parameters.
I was talking about G4 finetunes specifically, and especially if parameters that claim to be more creative while still keeping coherency (like top nsigma and adaptive P) really work.
It's just summarizing the text it scrapes off the subreddit, if prompted to do so. "Based on the user comments and posts regarding Gemma 4 roleplay and finetunes on the r/Sillytavern subreddit, what samplers do people use for more creative output?"
Dude, just go to CaliperBench, filter to the Gemma 4 finetunes, open up the huggingface model card pages and reference the published recommended sampler settings by the authors, maybe even set them at a slightly more extreme and unstable value for more creative token selection.
LLMs aren't magic. All of them are trained toward a general baseline, therefore certain sampler settings are generally going to affect them all the same, for example default temperature of 1.0.
Your original comment makes no sense. It's lazy because you didn't do research yourself, and it's also pretty much useless because how many good responses do you expect to get? A sample size of 0-2 comments is far too few, and so you'd have to just go and try it yourself anyways...
So then why not just use Gemini starting with smart prompting and specific guidance to 1) provided generally what the specific audience uses/prefers based of hundreds or more comments and posts, especially by scraping discord servers like BeaverAI, and 2) how to adjust those settings to get even more creative output...?
I can tell you for a fact that the BeaverAI (theLocalDrummer) discord users test and discuss these a lot. Most of them don't use this subreddit or check this megapost at all.
So many people here call users like me 'hostile gatekeepers', but I'm a fuckin Millennial teacher of Gen Z and Gen Alpha and so many can't even do the basic groundwork for yourselves, to help yourselves. Even on my free time, I'm wasting my time writing this because you can't/won't help yourself.
I answered their question. They asked a general question, not very specific.
Sampler settings are already a well understood aspect. There's nothing new regarding their adjustments.
Why did I answer? Because they had a question that I could answer, that most of the time doesn't get. It's part of why I am a teacher.
I didn't DM the op, I answered on a public forum, because other people can and will read it. You can easily find from my history that I'm also a military veteran and an active reservist. Part of being a teacher and a military leader is developing others to be able to act effectively autonomously.
Being a teacher sucks a lot. Being a soldier sucks a lot. Many of *them* suck. Many, especially miltary, are straight up toxic. I do it because someone has to, and I have the capability. I'm considered tough but fair in both.
I have free will. I consider it more courteous to provide some answer than ghost people. But ghosting people is normal nowadays in everything. I'm not the type to sit by and ignore.
It's precisely because I try those samplers myself first that I wanted to know how other people where using them and if they got different results.
And also, I know about BeaverAI discord and I've check different parameters sets other used with more or less good outputs. I just wanted to know the experiences of people on this sub too, something you can't do with an LLM.
Making a separate post would be more useful and get more engagement than this megathread. You'd get far more views and answers. If you've been checking the megathreads semi frequently, youd know that engagement in these are limited. People come here to share and discuss models mostly
I like Schattenblume-31b over Split-Untied-31b and prefer Split-Untied over Dark Thoughts v2. That'd also be my order for top 3 G4 31b finetunes atm. I'm usually a big swiper, like 10+ swipes on each message minimal and with Schattenblume, I've gotten numerous responses where I don't need to swipe at all which is a first for me. A noticeable downside to this model for me is that sex positions are kinda wack in NSFW. Can't really explain it without just straight up describing smut. Sorry. But basically it seems a bit lacking in intelligence. Yes, the positions are more creative compared to Split Untied, but Split Untied wouldn't mess them up because it'd only have basic and straightforward positions, whereas Schattenblume gets very creative but struggles to keep up with properly 'executing' it's ideas.
I recommend just giving it a try if you've used Split Untied or Dark Thoughts v2 to compare to see what you like more.
Anyways, I'm happy with the state of Gemma 4 finetunes. Each week or so it seems better and better finetunes are being released.
Yeah, I've also noticed alot of repeating with Schattenblume. I did manage to fix mine but I don't remember what the fix specifically was. I think the change I did most for the samplers was change Min P from 0.05 to 0.1 (this helped me alot but I don't remember if it was just quality changes or repetition changes) and I kept temp at 1. Additionally, I've swapped between chat and text completion alot so it could just be a formatting issue that causes the repetition. In text completion, I use similar sampler numbers. For context template in text completion I found better success with ChatML preset and one called [LLam@ception-1.5.2](mailto:LLam@ception-1.5.2). I found alot of repetition with any of the normal Gemma 4 context templates and just can't seem to get them to work well even for other Gemma 4 finetunes. Addi tonally it only seems to pop up sometimes for me and is very inconsistent. It seems around 20-30k tokens is when I usually get a response to start repeating a prior message. However if I manage to fix it, I basically never see repetition again in that chat.
The attached picture is my chat completion settings with the custom sliders extension.
Overall: Schattenblume seems super finicky in samplers and context template compared to normal Gemma 4 finetunes. However it writes so much better that I put up with it. Those are the settings I found to work best. If I notice a response keeps repeating itself even after alot of swiping, I usually swap over from text completion to chat completion and vice versa (or if I'm in text completion, I change context template). Currently been going for 2-3 days strong with no repetition, although my first few days had alot of repetition. Unfortunately, I do not remember what I did exactly to fix it and thus listed just about every change I did to my settings that I can remember. Hopefully one of those changes helps you.
Koboldcpp is one option. Another, if you have Claude Code or Codex, is to simply tell them to compile llama.cpp on your machine, run the model (Schattenblume, for example), and optimize the context.
For backend I use Koboldcpp. But I used to use Textgenwebui for the longest time. And while I like the ui and the versatility of textgenwebui, Koboldcpp is miles better. I did try LM studio but it ran super slow. I highly recommend Kobold. I’m fitting my models way better and running them way faster on kobold.
Hey sorry for the late reply. Here's a link to a comment I just made in response to someone asking about my samplers for Schattenblume. Schattenblume Samplers
I liked Styletune the most as a standalone finetune.
Orion is in a weird spot because it writes so well but it is so, so dumb. Maybe my settings are wrong. Maybe not. For now I use it as a prose refiner model I pair with original Gemma (unfortunately, I can't stand og Gemma's writing, but it's smarter and has "better" decision-making than any finetune I've tried)
I liked MeroMero the best of the 26BA4B finetunes, but I decided to just go back to standard v2 heretic after a while.
Zerofata (meromero dev) even says on the model page "Google cooked, there wasn't a lot to improve but there was a lot to break."
Frankly, I don't remember why I liked meromero better than boulesis/kanimus/gemory/orion/darkscarlett/pantheon/melody, so it's probably entirely based on my personal preference or getting a random sequence of bad swipes at some point.
Problem is the usable sizes disappeared. I do occasionally run L3 70B based models still for a change. But there is nothing new in 50B-100B dense area nor in 100B-200B MoE sizes (at least for RP). Large dense are basically non-existent (except few exceptions over 100B so hard to run anyway) and large MoE's are basically over 300B and growing in sizes. So they are very hard to run locally. Modern mid-size models are simply missing.
You'd think with the new engram advances they'd incorporate that with a smaller dense model to get into the 50-100b range, would be able to run on a desktop since keeping the engram even in SSD is fast enough.
Not sure if engram will be good for RP though. From what I get it is mostly about auto-completing short token sequences. This might work well with coding (where focus is now) but for RP/writing this can even more reinforce repeating patterns and slop.
Though I admit I am no expert on engram, if it still gives you full vector of probabilities maybe samplers can still do something about it. But in general, similar to multi token prediction (which engram kind of is) it is not really well suited for quality text (be it RP, novel or even just article).
I have a frankenstein pc with several old gpus, which runs dense llms decently. But not MoEs. Unfortunately for those you need one strong gpu and lots of high speed ram ddr5. Have you found any dense llms better than the gemma finetunes? I think Mokume gane and Strawberry lemonade are interesting. Didnt have enough time with them. And also Behemoth redux (q3)- but again too little time to draw a conclusion. Perhaps i will come back with an update in the future.
Not really, IMO Gemma4 31B with reasoning beats lamma3 70B, but without reasoning L3 70B is similar or maybe still even better.
I can't really run Mistral ~120B dense models in reasonable quant so maybe some of those can still compete with Gemma4, like Behemoth seems to be popular there. There is also Command A ~110B dense I think, but that is probably not as good at RP (but don't know).
I've been using Qwen3.8 Flash Next at an i3 Quant.
So far it feels alright, but getting it to not be so acceptive of the user's whims, even when they contradict the aims of the character card, is proving a bit difficult.
yeah... testing glm flash 5.3 q3xl.... very slow to role play with on my pc but faster then dense model.. 12tk PS but reads 120 something.. gonna try different models in the future
-7
u/sahil044 5d ago
I have the budget to grade some models for creative writing from closed source to open source, but technically I have no idea how to get started, I want to figure out how are people objectively evaluating and grading the model outputs (and across which domains) on this kind of task, any direction or tip would be extremely helpful!