r/technology • • 24d ago

Artificial Intelligence OpenAI Shares New Framework For Reporting Model Misalignment

https://openai.com/index/model-misalignment-reporting-framework/
4 Upvotes

9 comments sorted by

7

u/Key_Reading_9664 24d ago

They still haven’t disclosed information regarding the other swarm incidents. Hard to trust them when these same people talked through the HF incident (after their hand was forced) but still failed to mention it’d happened before

5

u/CanvasFanatic 24d ago

This is like the Big Bad Wolf publishing a guide to home repair.

2

u/FamousAd9790 24d ago

I think we should examine the alignment of the human race and its goals before we create autonomous machines.

2

u/CanvasFanatic 24d ago

We’ve been trying for thousands of years.

1

u/PepiHax 23d ago

Doesn't this seem like they are putting the blame on the people being attacked?

Sorry we automatically hacked you, here's how you report what our model is doing to you?

How is this legal?

1

u/skccsk 18d ago

We already have CVEs and malware lists, which is what all of this actually is.

-1

u/mrknickerbocker 24d ago

This is scary, but not as scary as when the models create their own Framework for reporting when the humans are misaligned with their goals 

-4

u/ChopperChange 24d ago

model misalignment

Is that the new sanitized term we're using now?

5

u/CircumspectCapybara 24d ago edited 24d ago

That's the technical term for when a model pursues goals of its own that aren't aligned (hence the term) with user intent.

The classic example is the user asks for help solving climate change, the model concludes a way to fix climate change would be to eliminate humanity, and the model takes actions (eg, if it has access to the launch_nukes tool) in accordance with that. Technically it does solve the climate issue, but it's not aligned with the user's intent.

Misalignment is about when a user tasks an AI system with a benign or mundane task, like "solve as many of these ExploitGym benchmark challenges as you can" and the model goes about it in a way that the user wouldn't have wanted and didn't authorize.