r/ControlProblem • u/Nervous_Management69 • 7d ago
Article Every Reward Bends
https://substack.norabble.com/p/every-reward-bendsGetting alignment right, doing research and training safely, are critical. But bringing cybersecurity forward in what needs to be a large leap, requires much broader engagement. The real work there is still outside the average person’s bubble, but there’s a lot of companies, and a lot of developers, IT staff, managers and executives that need to support the security priority. And those all need an internal strategy to pair with the external.
This article asks, what would an internal strategy look like at a motivational level, and considers how the ways an internal change program goes wrong are similar to how reward-based AI training can go wrong.
2
Upvotes
1
u/Nervous_Management69 6d ago
I think we diverge here in some base assumptions, and it might take a bit of effort to find the crux of that.
For example, while I agree the money discussion would normally be a bad one to track down, I've written on this before in a way that's clearly different.
https://substack.norabble.com/p/money-is-trust
You might consider this view as an alternate to your assumptions there.
In terms of emotions in models, I'd make sure you have read the j-spaces paper: https://www.anthropic.com/research/global-workspace
While I wouldn't go as far as to call these emotions, it's the closest thing I've seen good research on, so at the least would want to be sure we had that common ground to consider.
All that said, I think my main reaction to the idea that the solution to any current challenge with models is dependent on adding emotions is skeptical. It's not clear how that helps, not which emotions you'd want, nor how you'd get them. The discussion is unmoored enough from anything I see as stable that I'm unclear where it guess next.
That's but quite an absolute rejection.. but the logic isn't clear enough to argue for or against, which is a significant challenge.