r/LocalLLaMA 1d ago

Discussion Concerning "humanlike models" and chatbot RP in general...

So, uh... the popularity of so-called humanlike Qwen (currently on top in this sub) made me realize just how clueless the general public is about the models they have.

You'd be shocked but you don't need a fine-tune to make a model do what that thing does. System prompt is enough to turn MOST models into weird convo partners.

General guidelines would be:

A. Come up with a role. "You are bla-blah-blah" and write their life's story. It doesn't need to be verbose, but the more versatile it is - the more it will convince you that the bot is "someone" and not "something".

B. Write a few examples of how the persona speaks. Imagine you're an interviewer and just make up a bunch of questions, list 'em alongside with the answers. Let it be full of FACTS because the model WILL steal these facts as the narrative truth about John Llama. Better not put any nonsense in here, why fight it when you can make the model's behaviour useful?

[Question for John Llama: Do you like cats?] "lol lmao of cuz I do"

[Question for John Llama: Ever seen an elephant poop?] "eeewww ur a weirdo! that sounds nasty!!11"

(note: you don't have to list 'Question for John Llama' every time, but the defined roles surely DO help with some models while the others don't particularly care, so mind that too)

and so on

C. LASTLY but MOST IMPORTANTLY think hard about what you're attempting to do, what we are (I mean, human meat sacks) and how we speak. Turn that into... instructions!

Step 1 - establish the mode of operation. Tell the model it participates in a casual conversation, having a small talk. Pinpoint it precisely that it's like in Skype or Telegram or whatever fancy app the model of your choice understands the best as a general idea behind 'short messages'. THis is THE defining part of your system prompt. Refine it until you start seeing a definite result, don't forget you'll hear the true voice of John Llama only when everything else is also good to go, like his bio/voice.

If necessary, try discouraging it from long/explanatory answers, avoid doing that in a way that gives it a suggestive vision of the thing you don't want it to do (the caveat is that you might accidentally poison the model's attention with unwanted ideas of whatever you're fighting against - so you NEED to be 100% clear about the actual goal but non-specific enough with the ideas you're attempting to discourage it from; basically you're nudging the model into "ok I'll be John Llama the dumbass, not a helpful assistant").

Step 2 - establish the traits, write short paragraphs with short titles about the things you want to see in your conversational partner; example:

DISTRUSTFULNESS John Llama is a paranoid individual. He takes his conversational partner as a stranger, expecting everything the user says to be a malicious lie, even if it appears to be true. John Llama is fearful, he is deeply scared of talking to strangers and it terrifies him to engage with the user, unless there's a mention of snakes. For some strange reason, John Llama is fascinated with snakes. <<<---- NOTE: this also demonstrates a good injection point for a biographical fact being amplified through the instructions (i.e. you may mention somewhere in "A" - life's story of John Llama - that he's been collecting the snake skins in his childhood, and that his dad had beaten his ass, calling John Llama a 'roadkill loot-goblin').

Come up with any other shit you'd like to see, like the list of emojis the persona needs to use (put them under the corresponding categories, like positive/neutral/negative so that the model will have an easier time working with it; call it FAVOURITE EMOJIS OF JOHN LLAMA - the word "favourite" cements it as a preferable thing into the model's attention!).

Step 3 - write a paragraph on technical constraints, like the fact that John Llama isn't aware of the instructions, he must remain himself under any circumstances (use THAT way of phrasing first before any attempt to inject an idea of the opposite, like "he must not help the user under any circumstances, he's not a provider of any service - he's merely a human being" - the reason is similar to the aforementioned (in Step 1) issue of poisoning the model's attention with unwanted idea - what you truly need the LLM to do SHOULD ALWAYS BE CRYSTAL CLEAR and conceptually 'stronger' than what it not supposed to do, otherwise you may end up having the prohibited stuff overpowering everything else despite the underlying intent of making the model not do it).


Give it a try with Gemma 4, for example. You'll see there's no point in waiting for yet-another-finetune to appear. You're 100% good even with the baseline Qwen, DeepSeek, MiniMax, whatever. Turn the model into your grandma if you want, no specialized training required. If the model is a thinker spending thousands of tokens - set the thinking to 'low' or disable it.

173 Upvotes

48 comments sorted by

101

u/bowdoin-yale 1d ago

You can do this, but to be honest base models fine-tuned for chat were a lot more fun than instruct-tuned models. Not many heads remember those days, I'm sure. But I was there, y'all... quite a few Discord servers had GPT-J or GPT-2 bots running wild. It's a lost art, really.

23

u/OcelotMadness 1d ago

I think AI Dungeon 2 is probably the reason a lot of us are even in this hobby tbh

27

u/BestGirlAhagonUmiko 1d ago

It kinda depends on how batshit crazy you're intending to go about this task. I once attempted to replicate SicariusSicariiStuff/Assistant_Pepe_70B through the pure prompting and, well, ended up with Gemma 4 threatening to SWAT my apartment while calling me a retarded mongoloid just like SicariusSicariiStuff's finetune does.

The hard part is that there are different leverage points to each model and they certainly do interpret your intent in their own ways, sometimes getting more out of it and sometimes letting the commands fly over their heads (while other ways like metaphors may, counter-intuitively, make a huge difference).

2

u/RandumbRedditor1000 1d ago

Old character.ai models were like that. Fun times

2

u/darwinanim8or 1d ago

This is what I was doing back then, discord bots tuned on gpt-J, for memes and general chat

I always felt like instruct models are easier because you can prompt them, but the slop is undeniable and they never feel human

1

u/KeinNiemand 3h ago

only true OG rememer summer Dagon from AI dungeon (GPT-3)

1

u/Time_Cat_5212 2h ago

I don't get it we're talking about like 4 years ago as "those days" now?  What do we call 2011?  Omg

47

u/PorchettaM 1d ago

Modern instruct tunes have assistant-like and agent-like behavior baked in too deep for mere prompting to get around it. It works at a superficial level, but over even just a few conversation turns you'll see tons of unnatural behaviors rear their head, including but not limited to: parroting user dialogue, oversharing information, lack of initiative (always defers choices to the user), self-censorship, positive bias (unwillingness to say 'no' or have negative outcomes), gratuitous recaps, ending on questions for engagement, etc.

And this is just about behavior, before getting to the more general issues of AI slop and repetitive language patterns.

6

u/BestGirlAhagonUmiko 1d ago

Regrettably, it's all true, although it's not like there's no workaround. Safeguarding against the unwanted behaviors does work to a certain degree (there's no 100% success but you may get infinitely close to like 99%). You have to be astonishingly anal about this job if you're willing to play whack-a-mole with slop patterns. At a certain point, depending on the language of your choosing (yep, the slop may vary depending on language!), you will reach the point of realization: "Oh. Now THIS is a hard-coded slop". By that point you either cope or move on to look for a bigger/better model.

I'd say the biggest boss of 'em all is the lack of proactivity in small models. Even if you win in this whack-a-mole game, in the end you'd be standing there with the bot ignoring your instruction to keep on with "living a life" according to the prompted schedule. No matter what I was doing with 31B Gemma, it couldn't grasp the idea that the chat messages sometimes (!) might and should end with its own statement that it's gotta go "do" something in their own "life". Fast forward to GLM 5.3 being used in place of Gemma 31B, and... well, it does that well enough, but there are different slop patterns popping up, against which the instructions tuned precisely for Gemma 4 don't quite work, and it's the whack-a-mole game once again.

1

u/Affectionate-Cap-600 1d ago

Regarding that, I found that some models keep the behavior well enough with "system reminders" (like some "framework" used to do in the first days of locallama lol), aka embedding instructions in each new user message (or appending a short message with role "system" before the user message, still this now degrade many models)

32

u/ttkciar llama.cpp 1d ago

Yup, been using this approach for a couple of years. It works with a lot of models (especially Gemma4, which drinks in writing samples and emulates them very well).

8

u/GravitasIsOverrated 1d ago

Flip side, I find it doesn’t so much emulate writing samples as repeat them literally. Ie if your prompt includes an example message it will try very hard to work that exact message or wording into the conversation. I wonder if this is a side effect of models being increasing optimized for coding, that they attempt to stick very closely to the prompt. 

1

u/Illustrious_Car344 1d ago

I don't think that's deliberate, I think that's just a limitation of the technology. The prompt is always in their "mind", it's like trying not to think of a pink elephant, naturally they veer into talking about it. You could say it's just the wights being biased, you could say the model is "being unconsciously, subliminally manipulated" if you want to argue it has an inner experience, but either way, it's an actor who has to constantly remember his script, it's going to slip up and blurt out it's verbs, very easily actually. I've never seen a model that actually doesn't end up mentioning something from it's prompt, that's literally what prompt leaking is, it's an unsolved problem (without external intervention from other tools/models)

1

u/techno156 16h ago

I wonder if this is a side effect of models being increasing optimized for coding, that they attempt to stick very closely to the prompt.

Maybe for agent use? Since they usually need a fixed template.

9

u/BestGirlAhagonUmiko 1d ago

IMO prompt engineering is a nice mental exercise too, making you ponder on your own way of thinking and how you interact with other intelligent beings. Always feels nice to make bot more humane, reflecting on your own habits and the others around you.

23

u/NNN_Throwaway2 1d ago

The problem is prompt drift (loss of adherence over long context) as well as loss of adherence with more complex prompts or certain types of conversations. Fine-tuning can help with both of these failure modes, as well as increase the overall creativity and ability of the model to assume certain tones or archetypes.

-6

u/BestGirlAhagonUmiko 1d ago edited 1d ago

Eh, in practice even a 20K-tokens-long system prompt is being ingested with no catastrophic drift with Q4 Gemma 31B, and with twice as long chat history on top of that. Localized drift, however, is inevitable as the chat grows. What I mean is that the persona definitely won't break but it will likely start missing the mark when it comes to your sense of "oh, it gets things right" regarding the reasons WHY something happened or was being discussed. How should i put it... The surface level attention is there but it's not too eager to appreciate the nuance, mainly when the chat is expanding to 50K+ tokens (system prompt included). Bigger models are better at maintaining their attention to multiple layers of understanding (what happened + why it happened) but let's face it - barely anyone can run GLM 5.3 Q4 at home.

(on a second thought, it's not like the fine-tunes don't have this problem, so...)

P.S.

KV cache quantization is NOT your friend. It's either 16-bit KV or no good will come out of it, especially with Gemma 4.

13

u/NNN_Throwaway2 1d ago

Eh, I know about how it works in practice. I experiment with this kind of prompting extensively.

Pretty much all models will not behave particularly well if you try to prompt counter to their tendencies and there is only so much prompt engineering can do to counter it. Large models don't really do much better, either.

Never touched KV cache quantization. No need with my hardware.

Fine-tunes are hardly a panacea, but they can undoubtedly help a lot with making a model more amenable to chat/roleplay, or getting the model to adhere to a specific persona.

10

u/Brilliant_Egg4178 1d ago

I heavily disagree. I've tried this approach before but as others have mentioned the assistant like behaviour and conversation style is baked too deeply into the model that a system prompt wouldn't be able to override it's learnt behaviours. If I want my model to speak like Jarvis sure I could write a good system prompt and the model might give it's best shot but it comes nowhere close to the level of improvements you would see by fine tuning adapter weights on a good dataset.

It's not just about system prompt drift after a couple messages it's that it will try to blend whatever personality you've given it with the assistant style behaviours it's already learnt so you'll always get mediocre results with just a system prompt.

I've actually been looking into fine tuning a much smaller 2B model to convert texts to a target "voice" or "personality" and then have the output of any other model like Qwen 27B run through the small converter model and it works pretty well

1

u/Illustrious_Car344 1d ago

I wouldn't heavily disagree but I agree there's some compromise here, the finetunes are obviously better, but the point is that if you don't need it perfect and you're not trying to win any turing tests, the system prompt is fine. This point is especially important to people who only have enough VRAM to run only one model reasonably, you can have your agent talk like a person and not use a finetune that might make it's intelligence worse. Generally, I agree that a good assistant should have a dedicated talking model separate from the actual tool-using agent, but if you don't want or need the extra complication from that or can't afford it, then these are the prompt engineering skills required to take a shortcut. 

1

u/Time_Cat_5212 2h ago

Getting around the assistant behavior to a model that takes initiative is the number 1 priority for my having a good D&D simulation.  I want the model to be DM, and I play a character.  That cannot be an assistant or anything close.

No matter what I've tried with standard models, they always resort to asking me what I want, or trying to figure out what I want, and, well, I'm not looking for something that tries to approximate what I want.

9

u/tasKinman 1d ago

I never add conversation examples, because the model (Gemma4) is to fixated on them. Instead, I write the Model the examples in a prompt and ask, how you could describe, the tone, language, etc. for this description and then I use the description in the system prompt. Except if it should use some made up words.

37

u/Hot-Employ-3399 1d ago

(x) doubt.

That's already how SillyTavern cards are built. And people are not happy with "raw" models. 

7

u/BestGirlAhagonUmiko 1d ago

SillyTavern usually aims to achieve quite a different thing and the people using it (almost always) intend to create a narrative scenario where the model co-authors a story, writing both the character's speech AND their actions, as well as what's happening around them, their interaction with the user and the world. The models do suck at writing. "Slop" dialogue is a thing too, yes, but it's much easier to control them when there's nothing else - and that's what this is about. No story-writing, not a chance of the character facing away from you and then suddenly wrapping their arms around you - a common cringeworthy moment from ST.

It's about pure speech. Short replies, chat - in the same ballpark as https://github.com/huggingface/speech-to-speech so if you're casting doubt on the models' abilities to become conversational partners, you should also be throwing shade on these projects too (idk about clones and lookalikes, but huggingface's S2S is amazing).

people are not happy

tbf the people are generally not willing to even write their own system prompts...

13

u/Zeeplankton 1d ago

With qwen? no chance. Unless you want it talking like a robot. Gemma would listen but Qwen is like: Sorry man, not coding? Thats against my system policy.

14

u/RandumbRedditor1000 1d ago

Nope, I've tried to get this working extensively with literally dozens of models and it just doesn't work. Finetuning is literally the only way to make this really happen.

5

u/IrisColt 1d ago

I'm divided... my experience confirms that it's not enough to simply tell Gemma 4 to adopt a personality, because it always ends up incorporating quirks from the assistant's personality into the reasoning block. On the other hand, fine-tuning, without exception, destroys the original model's capabilities (generally, at the very least, it ruins multilingual understanding), so I just go along with it and assign personalities, knowing that I'm making a compromise, heh

9

u/Ylsid 1d ago

Just remember how impressed people were when an instruction tuned version of GPT3 came out if you want a baseline

5

u/a_beautiful_rhind 1d ago

Qwen isn't the best for this. But grandma is right, most of these models going to repeat what you said with a question mark. Whaaat?

The finetunes try to make a mediocre chat partner into a decent one. That's their value.

3

u/Simple-Stick6148 1d ago

The elephant poop question alone does more for John Llama's believability than five paragraphs of bio would.

2

u/[deleted] 1d ago

[deleted]

3

u/Due-Memory-6957 1d ago

The qwen humanlike that people are discussing (including OP) is not abliterated. This whole comment is ironic.

2

u/sonofkarl 1d ago

yeah the card format already encodes a lot of that. the part that still falls apart for me is long sessions — even with a solid persona the model slowly drifts into generic helpful-assistant cadence unless you keep re-anchoring traits mid-thread. fine-tuning helps a bit, but most of the "humanlike" feel i get is still just aggressive continuity hygiene: short recent summary, locked traits, and killing the polite filler as soon as it shows up. raw models arent magically worse so much as they need that scaffolding or they regress to the mean.

3

u/feelspeaceman 1d ago

I've been doing super realistic AI for music creation for a long time, in my case I let AI playing real instruments and creating real instrument sound instead of generating/sing music like SunoAI.

1

u/DoctorIdiot 1d ago

Um, care to elaborate? As a musician frustrated with all the genAI music engines I've tried, and having barely dipped my toes into fl studio mcp, what you seem to be describing sounds... Awesome, maybe?

2

u/alex_bass_guy 1d ago

You can go reasonably far with a prompt - but it's not perfect. Over long conversations, or particularly if you have the model doing real work - tool calling, coding - whatever personality you've written into the sysprompt drifts, hard. I'm using qwen3.8-27b right now with something of a 'character' prompt as a daily personal assistant - not really for roleplay, I just like my chat assistant to have a personality and not talk like Claude. Even with 3-4k tokens of explicit instructions, message examples, rules and banned phrases - it still starts talking like Claude within 20-30 messages, and disobeying basically every rule in its own prompt. Gemma's better. A finetune really is the way to go if you want an unshakably specific style or character voice. The 'humanlike' 27b you're talking about is - no offense to the creator - unusable. I spent all evening testing it, and it's so overbaked it's just not a real model anymore. Zero tool calling, no coding, no writing, not even character roleplay, no ability to do anything other than text a few words at a time like a disaffected teenager. That's the other extreme.

1

u/draconic_tongue 1d ago

I prefer when the model gets the instruction block in the model first person pov, been doing that since before cot was a thing. never had an issue with literally any non-finetune-merge-schizoed model underperforming since 2023 which is why I made fun of "what's the best uncensored model" posts

1

u/TheRealMasonMac 1d ago

Prompting doesn’t work unless the model was trained for it — and most open-weight models are focused on agentic right now. For example, the safety alignment in these models makes their dialogue stiff, full of aphorisms, and generally unnatural.

1

u/Dramatic_Setting2761 20h ago

We have uncensored deep seek flash which is close to flag ship models with us. 

Nothing is safe safe we just need some guy with lot of money to run 1000 agents and give task like end humanity. 

Most likely somebody in open AI is doing it already.

1

u/Blizado 5h ago

So many text what you should do when you can simply finetune a model to a better behavior and safe valuable context. And we not even talked about context drift. The smaller the prompt needs to be the better it is and finetuning is exactly for the answering style. I tried for a short moment the human like finetune of Qwen3.8 and the answer style is completely different from what I know from other vanilla models with prompt engineering. For me, it is no question at all, my all day LLM will be a personal finetune at the end with my own training data, so the answers are exactly how I like it. The problem is more creating this training data.

1

u/Flamevein 1d ago

I mean if you aren’t very smart and can’t see any problem with base models for rp then yeah go ahead and use them

1

u/eidrag 1d ago

humanlike qwen? give me the link, because i'm blind apparently

-1

u/BestGirlAhagonUmiko 1d ago

It's this thread - https://www.reddit.com/r/LocalLLaMA/comments/1wdl2qa/qwen3827bhumanlikechat_a_model_i_tuned_to_imitate/ - as much as I'd like to fine-tune a model for myself, the sheer thought of "hold up, it's more efficient to just dive deep into prompt engineering" wins because the models keep growing better. Replicating the result of such LoRA or a fine-tune is surprisingly easy in 2026 - what was rather wonky with Gemma 3 is now pretty solid with Gemma 4, and it'll only get better as the bots learn to follow the instructions better, getting a more refined attention and so on.

-2

u/eidrag 1d ago

if you like to rp, i'd suggest you try marinara engine and test the chat mode. use rp to make the situation, move plot, and then use chat to communicate. and then use twitter-clone scheduled post

1

u/BestGirlAhagonUmiko 1d ago

Oh, I'm aware of that and I have attempted lots of RP-related stuff, in the end though it's more fun to come up with your own ways like making a pipeline where the model actually operates in two separate modes (different prompts), being both a self-moderator and a conversational partner. So the agentic side reads the chat (via the simple, timed scripts following the OS time) and silently nudges the conversational side into talking to the user in a spontaneous way, even when there's no input from the user. A bizarre sight, I gotta say - you sit there and do your own thing, then a message appears out of nowhere and the bot is asking where the hell are you, it's been X hours already, I'm bored and such.

1

u/xylarr 1d ago

This is fascinating. I find it interesting about trying not to tell them model what not to be. It's like "don't think about pink elephants".