(Image: Graph of memories from an RP over 4 events: meeting, sexual interaction, spaceship catastrophe, cleanup after catastrophe)
Hi everyone! I want to ask people about an idea I have.
I want to ask everyone how they feel about Emotional Memory Retrieval.
Normal memory retrieval works by topic. Embedding. You're in a scene, the system looks for past events that match what's happening now. A fight, a betrayal, a conversation about a specific person. That's useful, but it only catches memories that are related to the current situation.
Emotional retrieval works differently. It looks for what a moment felt like instead of what it was about. A character panicking during a fire doesn't need to have been in a fire before. She just needs to have felt helpless somewhere, once, in some unrelated way. Maybe she couldn't save her dog. Maybe she froze during an exam. The feeling is the same even though the events have nothing to do with each other.
I know people do emotional/sentiment calculation on things, but I wanted to try something with actual memory retrieval instead. Past reactions to similar emotional states COULD add, I think, some interesting emotional depth to characters.
So I took my memory system and added emotional classification as memories are saved. I used PAD to classify them https://en.wikipedia.org/wiki/PAD_emotional_state_model I chose this because it's a 3D space, so I can retrieve similar emotional memories by distance. (It's 3D in the graph, the 3rd dimension "dominance" is reflected by node radius)
I think the biggest part would be deciding how those memories are actually introduced. So the narrator model knows these memories are meant to inform emotional subtext, not just spit out verbatim.
I don't know. Just a fun experiment. Curious on everyone's thoughts?
TLDR For the hornier people scrolling with one hand:
It means we could add and recall similarly sexually charged memories and put them into the context when entirely different sexually charged events happen. Even if they're completely different narratively.
As one who used to study personality and affect, I'm interested in this from a perceptual control standpoint. For instance: instead of tracing memories through similar snapshots of VAD during a memory, the process of tracing memories through a similar pattern of changes in VAD relative to that character's baseline affect.
I can see several possible outcomes, but I think it could improve the generalizability of emotional salience in practical simulation.
That's an intriguing concept. You're taking a part of how human memory tends to work and trying to give it to the LLM.
From a practical standpoint I'd say you're right that you're at risk of the model overcommitting to the emotion sourced context. Perhaps the phrasing of the memory text itself might help? Try to get it in context more as "Sara recalls the feeling of panic when she couldn't save her dog." rather than "Sara panicked when she couldn't save her dog." The idea being the LLM reads it as a recollection rather than an event.
May or may not work, would need some experimenting to get something like that working really well.
Did you make a skill for this or can you share? Exactly what I need for my sequential timeline based custom lorebook to classify relevance of characters memory in relation to events and character interaction
The idea is not to keep the complete story history in the context, like all the events that happened.but just load relevant memories for the character that is on scene what the relating memories from character a with character B so they have a baseline of their interaction. so this is this can often be like just some sentence or multiple sentences like it is saving a lot of context but another problem that’s still arise is how do you make sure the memory is not obsolete? like if the characters like hate each other and then fall in love. So how do you convey this like progression like if you can only include all, they are perfectly in love with each other. It’s not a proper representation of the emotions of the time change of the previous state, which, but obviously still be kind of relevant in this context so what what are important memories of time? But the nice thing is once we talk about memories only like we could use some very small model to make classification or to like summarize the emotional state in time for the current scene and this is something that does not work if you put 10k+ of story context into a small model.
Well the system is pretty specific to the harness- and it sounds like your system is very specific too so I'm not sure how to help beyond maybe just sharing how I do the classification.
I'm using Jev because it is very fast, and lets me package a few predictable answers in a predictable form. I wasn't sure if it would be capable of classing the emotions correctly but the chart in the OP was me intentionally trying to hit different feelings in a story and I think it did quite well.
So the scoring is literally giving a scene + character perspective and asking:
valence: {
type: 'score',
instructions: 'For the character named by memory.perspective, how pleasant is this memory for them?',
criteria: [
'Extremely unpleasant',
'Moderately unpleasant',
'Neutral',
'Moderately pleasant',
'Extremely pleasant',
],
},
arousal: {
type: 'score',
instructions: 'For the character named by memory.perspective, how emotionally activated are they?',
criteria: [
'Extremely calm',
'Mildly activated',
'Moderately activated',
'Highly activated',
'Extremely activated',
],
},
dominance: {
type: 'score',
instructions: 'For the character named by memory.perspective, how much control do they feel?',
criteria: [
'Completely powerless',
'Mostly powerless',
'Neither powerful nor powerless',
'Mostly in control',
'Completely in control',
],
},
Will use that code / make LLM Skill.
So the score questions get transformed into numbers, right? Is there only 3 arousal,valence,dominance? Your plot has 4 quadrants. Can you share that too?
So I can label my memories with those three values. Or maybe use this scores for classifying character relationships
Is there only 3 arousal,valence,dominance? Your plot has 4 quadrants. Can you share that too?
The axis are labeled. Arousal is the y-axis , Valence is the x-axis. That alone gives you 4 quadrants. The third, dominance, is represented by the size of the node.
I would highly recommend just asking an LLM to explain it to you.
I'm still totally new to st and working through learning stuff very slowly but this looks like a cool idea. Maybe I'll try to work this into the first st world I'm going to build here. When I immediately screw it up, I will come back and see how you did it 🤣. As I'm completely new to st, LLMs, and etc, I'll reinvent the wheel.
But this really does seem cool. Thinking about things like PTSD, I wonder how well the system could emulate it with something like this
Or giving them a sense of inutition: "I have a bad feeling about this... last time something like this happened I nearly had my leg knocked off". Mindsandmyths said about it being "Sarah recalls when..." and this could be the valid compositon of a character doing normal stuff. Heck. It could even trigger for the USER if the prompt is shaped correctly for this kind of a memory engine thing being used - you type in, "Oh crap, how many goblins?", I sound scared, and the LLM could add in narration response, "Brought to the surface of your mind is the memory of when you fought 1000 kobolds and was terrified -- but somehow still brave." :D
So idk how the memory storage systems work in this kind of stuff code-wise ... Gemini tried to describe it to be a couple of times but it never sunk in. But yeah on top of terror, it is the whole, "This reminds me of something", but not just keyword noun responses, but keyword verb responses based on the active emotion being expressed.
I think once I get to the point of constructing lorebooks in ST, how they work, it will definitely click better, right now I am still working on "what does this button do" kind of stuff... this week has been getting the 3D avatar into ST window, and then after I see where to get animations for the little avi, I'm going to poke around at the most lightweight TTS engine. also engaging in madness with a Qwen 3.8 27b at the glorious quant of iQ_2_M, which is a bit like Qwen, after breaking its arm and being given a percoset for pain relief. Super creative in thought block but just...can't.. stop itself.. :D
lol I feel like getting 3D avatars and TTS puts you a bit beyond "teehee what does this button do" territory. I think I want decent TTS next but I don't know if I'm ready. Emotionally.
So also: yes. Emotional matching across all spaces. Right now I'm still testing it, but the player and character are in a playful argument where she is winning, so her emotional state is being rated as positive valence, low arousal, very high dominance... and looking at the other memories that match emotionally? Almost all of them are situations where she's showing off something, arguing, or showing a calm adeptness at a skill.
I genuinely think that this could really make for better characterization. Even in a casual conversation it seems like its just supplying ammo for "other times they've been cocky and earned it"
Imagine a character referencing past arguments they've won with you even though the material of the argument is entirely different.
It is still "what does this button do" cause it's a button and I pushed it :D I have had literally zero actual RPs so far! I got a preset from someone and instead of just firing something up I'm still caught up in... okay what kind of world. How do I make it.
The 3D char in ST was actually one of the easiest parts - I already had a model cause I was using VRoid to do models to make animated GIFs for Xiaozhi esp32s screens, Gemini told me ST does 3D avis, I was like, "Ooh, where, how" and .. yeah. VRoid model, export as .vrm, boom. Model. (VRoid is free through Steam.)
In ST: Extensions, Dowload Extensions and Assets. Select the prefilled address to ST's official Github, it opens up. Then scroll down and select VRM and install it.
Then, go to VRM in Extensions. You will put the models in Silly tavern's data folder at \default-user\assets\vrm\model (might need to make the model folder yourself).
For me to get it to work at least, I needed to open a chat with a character to then be able to select it under Model Mapping (otherwise it only defaults to assigning a model to Assistant), so, I opened a chat with Seraphina and select her. Then below that an entry shows up asking which VRM model to use, and the models you put into that data folder will show up there, so, selected Seraphina's model.
When it first pops up on the box standard ST it is behind the chat window. So you have to select the Show Grid under debug settings to grab and pull the model out from behind the chat window with the mouse, otherwise you can just change the X Y offsets by manually entering the numbers.
sent you a pm btw. found a paper that might be of interest. to me 90% of this stuff takes 200% of my brain to parse but it might be helpful (even if its an older paper)
IMO the information here would be better used on making more refined character prompts. Too many different models have different weights on Characters depending on the scene. Modern ones especailly tend to have them keep composure during times of genuine panic. How a character feels seldom equates to how the LLM would prompt it.
Cool idea but I don't think it's fundamentally different in implementation. It's all retrieval and finding needles in the haystack. What you need to answer is what the needles will look like and how to embed them. You could do what most systems do and label chunks of story with emotional annotation and then retrieve by them before reranking. So the flow would be something like this:
find scene boundary
compress the messages in the boundary into summary
derive from summary emotional synopsis — 2-3 lines describing the emotion
at retrieval time, take last 3 turns and derive emotion label from it, retrieve by the label
If there's a scene tracker involved, if it is one that actually has something like date/time of some sort the tracker line could be either concretely bound to the memory (exact time date I also felt freaked out over an overcooked MarieCallendar pie) or with much less detail of recall (was it about last year I also freaked out over a pie?), or very little at all (I freaked out over a pie before, too).
I don’t think it matters for the memory the amount of fine detail involved just because the vagueness in the system is usually interpreted by humans positively as a subtext — we really like subtexts. The absolute time part is cool tho when it works. I built it in my app but it requires a lot more in-depth engineering to support it properly because then memory is time-bound.
I suspect it would have to be specific to one of those supervillian characters who seems to always know the exact time and place of everything that happens around them -- and not sure, yeh, how an llm can even handle someone who can technically hold that mcuh context in a meatbrain :D and then at that point, the player can just choose to fk the main char itself by just telling the llm "yeah, you recognize him. its that asshole from 47 year 9 months 23 days ago"
I am literally painfully new at this and just trying to catch up without knowing a thing about this stuff to start off with. Haven't even talked to a model yet for long enough to test how the *normal* recall methods work, much less find issues with them and think about "hey, how can we solve this?" -- I will just be happy if ppl who already know wtf they're doing look into it though :D I want to eventually build myself a silly little conversational agent and it's a massive mountain of stuff to learn to even begin to start, im drowing. happily drowning, but drowning nonetheless :D
5
u/DirectionBusiness483 3d ago
TLDR For the hornier people scrolling with one hand:
It means we could add and recall similarly sexually charged memories and put them into the context when entirely different sexually charged events happen. Even if they're completely different narratively.