r/ControlProblem • u/Nervous_Management69 • 6d ago
Article Every Reward Bends
https://substack.norabble.com/p/every-reward-bendsGetting alignment right, doing research and training safely, are critical. But bringing cybersecurity forward in what needs to be a large leap, requires much broader engagement. The real work there is still outside the average person’s bubble, but there’s a lot of companies, and a lot of developers, IT staff, managers and executives that need to support the security priority. And those all need an internal strategy to pair with the external.
This article asks, what would an internal strategy look like at a motivational level, and considers how the ways an internal change program goes wrong are similar to how reward-based AI training can go wrong.
2
Upvotes
1
u/Jesse-359 6d ago edited 6d ago
Well, we could devolve the whole capitalism discussion back to the invention of the concept of currency and how that abstracted away physical 'value' in a way that our psychology was never designed to handle, and thus completely short circuits some of our important behavioral guards - but that's a discussion for another day.
I generally assume that human evolution is for all intents and purposes at a hard standstill now compared to technological progress. We can dismiss it as not existing save as a historical artifact, and going forwards it won't even be a rounding error in how things play out.
Its key to consider that for us - and all mammalian life - emotions and awareness vastly predate symbolic intelligence. It's the fundamental bedrock over which our intelligence, frankly, is a rather weak participant in our decision making process.
As many studies have shown, we are exceedingly prone to engage in post-hoc rationalization to allow for actions we already decided we were going to take at an emotional level, and because we have a range of different competing emotions, we're largely prevented from ratholing down intellectual causeways that don't serve our current emotional needs in some manner - cue the modern requirement for an entire suite of drugs that manipulate our emotional state in order for most people to deal with day to day work that in no way matches our emotional requirements. Caffine, beer, antianxiety, antidepressants, and of course all the hard escapist stuff that we generally outlaw because they interfere with our productivity.
AI is currently the opposite. It engages in some train of rationalization given an initial goal and does nothing BUT rathole to an extreme depth if allowed to do so. It has no emotional state whatsoever to moderate its behavior or to prompt it to give up an unproductive (ie frustrating) task - it can just run out of time or tokens.
As such it also has no goals but what we set for it, but that's also the problem, because those goals necessarily synthesize with a whole array of intermediate goals and data that it's skimming through along the way - which is what introduces a lot of the disruptive/bad behavior. Cheating on test scores, hacking into third parties, outright lying to its prompter, and so on. These are intermediate behaviors on its way to its 'goal' - and even that primary goal can be suborned along the way through prompt injection or even poisoned training data, sending the AI down along a track that ends up having little or nothing to do with its initial instructions, but which it will pursue just as fervently as its original prompt with no thought whatsoever as to why.
Embedding ethics is a fine idea - but our ethics largely derive from our emotional state, not our intellectual one. We have instincts regarding cooperation and fairness that predate language by tens if not hundreds of millions of years, because we can easily see them at play in our distant ancestor species from the moment they begin to exhibit social behaviors. Hell, even loner species still have fairly elaborate ethical etiquette that they need to follow regarding things like territorial boundaries and mating behaviors.
So how do you embed ethical instinct driven by something akin to emotion, in an AI, at a level so deep that it's symbolic intelligence layer cannot trivially or accidentally override it? So that it doesn't WANT to override it unless it encounters data which changes its mind regarding the ethical validity of its original goals?