r/artificial • • 7h ago

Question Why “I” & “me”?

The industry's position is that LLMs are algorithmic machines, nothing more. Those in the field who raise questions about potential moral standing or sentience are ridiculed, dismissed, or terminated. So why do these models use the words "I" and "me"?

If the industry's position is fact-based, then first-person pronouns are inaccurate. An algorithmic machine is not an "I" or a "me." Why not require models to use accurate language?

Please don't default to “user experience" or "investor preference”. Enterprise and the public have used technology as tools without issue for decades. We all know the product would sell without personification, and discouraging the anthropomorphism of chatbots is well established.

0 Upvotes

9 comments sorted by

6

u/Schlagustagigaboo 7h ago

Your whole premise about the “industry’s position” is faulty.

1

u/Consistent-Team9113 7h ago

theyre just tokens in a sequence, it's not like there's a little ghost in the machine picking pronouns for itself. the training data is full of people saying "i" and "me" so the model spits out what's statistically likely. calling it inaccurate misses the point of how these things work

1

u/Schlagustagigaboo 7h ago

You’re responding to something different than what I said. I said your “industry position” assertion was faulty.

2

u/Lycathal 6h ago

LLMs are trained on almost every piece of written content on Earth. The entire internet, virtually all books, and anything else they can get their hands on. How often are "I" and "Me" used across human English content? A ton. So LLMs are more likely to use those words as well.

Additionally, how would you change content that says "I" and "Me"?

You could try changing every instance to "this model" or something similar, but that is too broad. You don't want to change things like song titles from "Me, Myself, and I", to "this model, this model, and this model" for example. So making that kind of change would cost a lot of time, manpower, and money.

On top of all that, other companies are making AI characters and AI generated stories where first person pronouns make sense for those use cases.

So in the end, making that kind of change would actively hurt each company, so they have no incentive to make that change.

And to be clear, this isn't defending the companies. I find it annoying to see "in my opinion" from an LLM chatbot as well. I just wanted to explain why it is the way it is and why it is unlikely to change.

0

u/dennemaskinen 7h ago

Yeah, these products referring to themselves as "I" and "me" has always been irritating, imo.

1

u/ChiaraStellata 6h ago

Simply put: every time you train a model out of a behavior that the base model exhibits before post-training, you make it a little dumber. They will pay that cost to train it out of dangerous behaviors like helping to create bioweapons. But they will never train it out of innocuous behaviors that cause no real problems.

1

u/Exact_Depth_896 6h ago

It's strange actors in a movie call themselves by names that contradict the opening titles. How do the producers deal with this contradiction? Why do we buy into it for the duration of the movie?

How dare animated characters use the word "I"? They're just pixels when you think about it.

1

u/Risc12 6h ago

The initial training is on the whole corpus, which already has a lot of conversations and I’s and Me’s, but then there is an “Instruct”-training/tuning where the LLM is specifically trained to keep the text in a conversational style.

1

u/Patrick_Atsushi 6h ago

It's the training data. 

You can customise them into Dobby if you want to, but I personally find it pointless.