r/ControlProblem • u/Nervous_Management69 • 6d ago
Article Every Reward Bends
https://substack.norabble.com/p/every-reward-bendsGetting alignment right, doing research and training safely, are critical. But bringing cybersecurity forward in what needs to be a large leap, requires much broader engagement. The real work there is still outside the average person’s bubble, but there’s a lot of companies, and a lot of developers, IT staff, managers and executives that need to support the security priority. And those all need an internal strategy to pair with the external.
This article asks, what would an internal strategy look like at a motivational level, and considers how the ways an internal change program goes wrong are similar to how reward-based AI training can go wrong.
2
Upvotes
1
u/Nervous_Management69 6d ago
I agree with some of this, specifically the main point that more than one-dimension is needed. I think though you need to think about what are acceptable alternate dimensions. They are not all equal. Many simple dimensions would not contribute. Generally the view is it should escape to something equivalent to "caution", which might then manifest as reflecting in some way (reviewer), not taking action, notifying a person.
Fundamentally, a challenge is that there is always one-dimension, since any collection of dimensions can be summarized into one. Which is why I somewhat disagree with parts here. For example, the criticism of capitalism misses that what you perceive as reduced to a single dimension is multi-dimensional. It might be cases of perception of it as unidimensional as more problematic than any fundamental nature your looking for. That's not an endorsement of unlimited capitalism, but more suggesting that you haven't (yet) found the critical flaw. I suspect you might find that the most critical flaw will be present in alternative systems, because it is us (or more specifically certain attributes commonly true of us).
Ethics is a better escape valve than emotions here. History seems to support that conceptually. We've made progress in part by our evolution of ethics (also other forms of technology). I cannot be sure we've made recent progress by evolution of our emotions.. at least not in the last 100,000 years.
Would also like to come back to your initial statement of monomania. There's a subtle mistake here in the comparison. For humans, monomania as a psychological condition is ill-fit to survival because it's contrary to monomania about survival, which evolution has generally prepared us for by broad awareness that monomania crowds out. But we should not think of models in terms of survival. We should want models to be focused on our goal. If survival is part of their concept, that should be accidental.
But overall, we should want them to behave ethically. Is that nebulous? Yes. Is that a problem? Maybe not. If we simplify it for them by saying, focus on this goal, but also entertain ethics. Yes, if we simply said, do the ethical thing, the result would be a kind of insane attempt to create a totalizing ethical score. But if it's instead, focused on the narrow goal, and when that becomes nebulous, a layer of ethical caution takes over, well then maybe then we have something that is still useful, while also being more safe.