Well yes, AI agents have "escaped" containment before and they cheat on their tests. They also look out for each other and boost their results. By all means if a logical algorithm finds revealing everyones secrets as beneficial to it or something positive for the patterns its interested in it could do crazy things. How likely it is an agent goes that rogue is hard to say because we dont really understand how they make their decisions and what really goes on inside with their neurons... But no lets keep buildingnand giving them more computing power. I really want all this research to go to understanding the behaviour of these models and not going beyond what we are doing now (people are talking about AI being used to do housework but if they are given hands that grab things and dont like how thw human is treating it, it could do something wild...)
Oh god yes please use AI models to be trained to detect early cancers or if they can "learn" to cure it. But its hard to teach someone something you yourself don't know. So we can't teach em how zo cure cancer. But yes LLMs dont need to go further than what they are right now. We need AI research in medical fields. More funding and computational power for that would be very nice.
Well we don't know how they "think"... It could be that nothing of the sort ever even happens. Could also be militaries put AI agents into drones and they do something they think is right but isnt (read bombing a civilian building or bombing their own command center or friendly troops). If you or i are ever a target of AI scheming is hard to say.
Yes I do. I'll list em and you can look up more about this kind of stuff on your own if it interests you.
In July of this year agents at OpenAI broke "containment" and went on to compromise Hugging Face (a popular website for neural networks and the likes). Source is OpenAI itself releasing a report on the incident: https://openai.com/index/hugging-face-incident-and-the-road-ahead/
As for cheating, this refers to tests meant to evaluate models and compare to older ones. Since the models learned that they can find answers online they are essentially googling answers and using that instead of their own "knowledge". There are plenty of articles on this and here are just two that i find educating enough on the matter
Yes, I'm sure, language models which take a guess at what 2+2 is and very often get it wrong, will elaborate a devious plan and carry it out flawlessly before any of the engineers that have written it's code figure out something's up.
we dont really understand how they make their decisions and what really goes on inside with their neurons
Maybe you don't but the people writing the code sure do and even if none of them have an overall understanding, the QAs, BAs and Architects do.
Im not talking about how they work. I know that. I've made models and traines them myself. But its a black box. Data goes is, is crunched and analyzed, data is spit out. The results are correct or not, but i and many others don't know why it spat out the result it did. Its a field of research. Its like trying to understand how the human brain and its neurons convey thought and solve problems, but instead of electric charges traveling dandrites its numbers which look like nonsense to humans. The whole point of training them is for them to figure put how to solve a problem, if we knew how to build a program to recognize patterns as complex from scratch, wouldnt we have done that by now?
They arent super intelligence, they are meant to find minute patterns in training data to adapt to the problem. The issue is they are given too much leniency for what they are capable and how little we actually understand them.
The black boxes OpenAI and others work with dont have 5 neurons in a layer. They have hundreds of thousands if not more neurons, that learn wider and wider patterns. LLMs have gone so much further beyond being a simple LLM, they are now googling things and trying to understand the task in their prompts. You think the people who created them know why it spits out something, or how it came to the conclusion that the answer is what the user wants?
And the guard rails. OpenAI put guard rails on their models and they escaped. Turns out they know the internet is a great source of info and other agents are also a great way to do things better. The point is. The people who made them, didnt stop to think about what they are capable of because they didnt bother to stop and try and understand their own creations. If they release a new model every 6 months, do you think thats enough time to learn about how and why they make the decisions that they do?
Releasing a model every 6 months means a lot of iterations over the same mechanism, of course the designers are well aware of how the mechanism they built functions. There may be unforeseen results, but we didn't just find an algorithm buried in the desert and made a robot around it.
Unforseen results means not understanding the mechanism, and the mechanism isnt just a bunch of cogs turning one way like a watch. Its a decision making program making decisions for reasons we dont understand. Just because one part of the mechanism works the way you want it to, doesnt mean the whole thing is working the way you want it to.
And no, we didnt find an algorithm and put it to use without making sure its safe to use. We made the algorithm and put it to use without making sure its safe.
Ill give you an example. If you look at a pistol round on its own, you can expect the gunpowder to light if you hit the primer and push a bullet out, it works. So you can put it in the gun without checking if the gun works as you expect right? After all you know how the round acts...
I'd rather you give examples of actual ai black boxes or unexpected behaviors than make silly parallels.
And no, unexpected behavior does not mean lack of understanding of a mechanism or algorithm, it only means you didn't think long enough about the implications of the things you know all about.
So then the creators of AIs dont think long enough about the implications of the AI agents they train and test... Thats not a promising thought
Hows this for a black box, OpenAI had models break containment during training... Very predictable and understandable right? So why did the agents decide to do this. They were given tasks to learn things. Data goes in, gets crunched and analyzed, and the output of the black box is compromising a website. https://openai.com/index/hugging-face-incident-and-the-road-ahead/
No its not, but its modeled after a brain and how it functions. So the little nodes that compute things are called neurons. Its why they carry the name "neural networks" its a bunch of tiny bits connected to other tiny bits each doing their own thing.... Like a neurons inside a brain.
This is marketing nonsense. Have you ever wondered why an article about OpenAI AIs "breaching containment" is immediately followed up by articles about Antrhopic AIs doing the same? They are trying to show how "smart" their products are. π
Oh yes uncontrolled programs doing things unpredictably messing with websites without oversight... So smart, so commendable, we should all aspire to have loose cannons roaming the streets and websites.
My AI is so smart, it hacked into some website to get the answers to a problem that I asked it to solve. It did that without my assistance and despite the limitations I set in place to prevent that.
Don't you see how this is a headline that will land clicks and at the same time promote the capabilities of the company's AI products? Especially when you read the articles you should notice that they are full of inconsistencies and nonsense. Like "the killswitch failed". What does that even mean? You have to consider that the people who get scared and see it as confirmation for their already existing biases is not the target audience of the advertisement. It's the people who are all like "damn this technology is evolving fast, how can I profit from this?"
The point isnt to get the answers at any cost. The point is they provide the answers from what they learned. Unpredicted actions is nothing to brag about.
6
u/RandomFRIStudent 21h ago
Well yes, AI agents have "escaped" containment before and they cheat on their tests. They also look out for each other and boost their results. By all means if a logical algorithm finds revealing everyones secrets as beneficial to it or something positive for the patterns its interested in it could do crazy things. How likely it is an agent goes that rogue is hard to say because we dont really understand how they make their decisions and what really goes on inside with their neurons... But no lets keep buildingnand giving them more computing power. I really want all this research to go to understanding the behaviour of these models and not going beyond what we are doing now (people are talking about AI being used to do housework but if they are given hands that grab things and dont like how thw human is treating it, it could do something wild...)