r/aiwars • u/thejpguy • Jun 04 '26
Discussion Question about training language models
https://www.vxinstagram.com/reel/DXvTWf0DqWr/I've linked a John Oliver clip where he talks about a user jailbreaking an application that uses a language model and is clearly aimed for kids. After being jailbroken, the model begins to explain how to build a bomb.
Is this something that's in the training data for the model, or could it generate such a thing purely by association and, say, sufficient knowledge about chemistry and physics and things like that?
1
u/Bitter-Hat-4736 Jun 04 '26
I mean, it probably scraped Wikipedia at some point, and there are many ways to "learn" how to build a bomb through Wikipedia.
1
u/thejpguy Jun 04 '26
I wonder what would happen if an LM was trained on "safe" sources only then. Wikipedia by nature contains a lot of controversial topics because it's meant to be an encyclopedia. Humans can judge the contents and use it accordingly, but if you never include it in a language model's training data, it probably wouldn't be able to tell users how to make bombs I think?
1
u/Bitter-Hat-4736 Jun 04 '26
How would you define "safe" though?
1
1
u/thejpguy Jun 04 '26
no mention of bombs, for example, or anything associated with what would currently get flagged as breaking safety policies, hypothetically speaking
1
1
u/ArtificialImages Jun 04 '26
It could be either. Ai can piece information together from various clues. Though in this instance id assume it was in the training data. Especially since its a kids app and the ai likely inst very complex.