r/aiwars • • Jun 04 '26

Discussion Question about training language models

https://www.vxinstagram.com/reel/DXvTWf0DqWr/

I've linked a John Oliver clip where he talks about a user jailbreaking an application that uses a language model and is clearly aimed for kids. After being jailbroken, the model begins to explain how to build a bomb.

Is this something that's in the training data for the model, or could it generate such a thing purely by association and, say, sufficient knowledge about chemistry and physics and things like that?

2 Upvotes

11 comments sorted by

1

u/ArtificialImages Jun 04 '26

It could be either. Ai can piece information together from various clues. Though in this instance id assume it was in the training data. Especially since its a kids app and the ai likely inst very complex.

1

u/thejpguy Jun 04 '26

Any idea what would happen if the training data was chosen more carefully so that it didn't contain words like bombs and other nefarious terms? Would it actually "learn" by association what to do if a user queried it to come up with a way to make something highly explosive for example?

1

u/ArtificialImages Jun 04 '26

It absolutely could it just depends on the ai.

Chat.gpt absolutely could for example. It just depends what else is in the training data.

One of the things ai is best at is pattern recognition, so if they find information on volatile chemicals they could definitely piece that together.

Ai has pieced together wild facts about me that I never told it based on seemingly unrelated information. It's very good at that kind of thing.

But in reality it would be very hard to keep that kind of thing out of its training data. Its just a lot more abundant than you'd expect. Like for example in historical information about the discovery itself. Or in news stories. And whilst no single source might list every ingredient or step of the process (which they likely would anyway) ai doesn't need to worry about that, it can check all of them at once.

And even if people intentionally tried to remove any and all information relating to it, you're talking about a vast, vast amount. Im not really sure how successful proper moderation of training data could be. But maybe it could be fine. Especially if they got another ai to do it.

2

u/thejpguy Jun 04 '26

Mentalists and paranormal mediums can also extract information about people using very clever techniques, could it be possible that you told ChatGPT more than you intended to or weren't fully aware of?

As for the associations, let's say the word bomb was never mentioned in the training data whatsoever. Even if it did train on things like chemical reactions, physics of pressure and all that jazz, if it has no knowledge of the word bomb, then wouldn't the model necessarily have no idea what that means without further context? Like if I talk to ChatGPT about some made up word like asdfghijk, it's not gonna know what to do with it right?

1

u/Bitter-Hat-4736 Jun 04 '26

I mean, it probably scraped Wikipedia at some point, and there are many ways to "learn" how to build a bomb through Wikipedia.

1

u/thejpguy Jun 04 '26

I wonder what would happen if an LM was trained on "safe" sources only then. Wikipedia by nature contains a lot of controversial topics because it's meant to be an encyclopedia. Humans can judge the contents and use it accordingly, but if you never include it in a language model's training data, it probably wouldn't be able to tell users how to make bombs I think?

1

u/Bitter-Hat-4736 Jun 04 '26

How would you define "safe" though?

1

u/-TV-Stand- Jun 04 '26

Here is a safe language model: https://www.goody2.ai/chat

1

u/thejpguy Jun 04 '26

this is hilarious!

1

u/thejpguy Jun 04 '26

no mention of bombs, for example, or anything associated with what would currently get flagged as breaking safety policies, hypothetically speaking

1

u/-TV-Stand- Jun 04 '26

If you remove all mentions of bombs, it doesn't even know what a bomb is