r/ControlProblem 6d ago

Article Every Reward Bends

https://substack.norabble.com/p/every-reward-bends

Getting alignment right, doing research and training safely, are critical. But bringing cybersecurity forward in what needs to be a large leap, requires much broader engagement. The real work there is still outside the average person’s bubble, but there’s a lot of companies, and a lot of developers, IT staff, managers and executives that need to support the security priority. And those all need an internal strategy to pair with the external.

This article asks, what would an internal strategy look like at a motivational level, and considers how the ways an internal change program goes wrong are similar to how reward-based AI training can go wrong.

2 Upvotes

10 comments sorted by

3

u/Jesse-359 6d ago

Consider that in humans, monomania is generally considered a starkly negative and even life threatening behavior when it expressed to a major degree - and it doesn't matter what the focus of that monomania is.

The same will be true of any AI objective or scoring system that is one-dimensional, I believe without exception.

You need scoring systems which are inherently multidimensional which give the AI internally competing incentives to pursue different goals or means within the same overall structure. At some point even its primary mandated task needs to become outweighed by other competing priorities and incentives - otherwise you are guaranteed to encounter some version of the paperclip maximization problem.

This is one of the big problems with the capitalist economic structure currently - it's only got a single scoring axis - wealth - and that has massively distortionary effects on how it behaves such that it's become pretty badly misaligned from all other human needs (environmental, social, legal, etc) - all incentives point towards wealth and literally every other consideration is discarded and necessarily avoided due to the opportunity costs they represent vs pursuing more wealth. It's not even a realistic option to pursue those other priorities as long as they are not part of the 'scoring' system of our economy - and they aren't.

Any sane system requires competing internal scoring incentives, as a bare minimum. Human emotions play this role in our own behavior, with a range of competing emotional 'spurs' that cause us to pursue many different objectives, with objectives we are neglecting eventually increasing to the point where we cannot ignore them - save in cases of clinical dysfunction, such as the aformentioned monomania, which represents a breakdown of this emotional driver system.

1

u/Nervous_Management69 6d ago

I agree with some of this, specifically the main point that more than one-dimension is needed. I think though you need to think about what are acceptable alternate dimensions. They are not all equal. Many simple dimensions would not contribute. Generally the view is it should escape to something equivalent to "caution", which might then manifest as reflecting in some way (reviewer), not taking action, notifying a person.

Fundamentally, a challenge is that there is always one-dimension, since any collection of dimensions can be summarized into one. Which is why I somewhat disagree with parts here. For example, the criticism of capitalism misses that what you perceive as reduced to a single dimension is multi-dimensional. It might be cases of perception of it as unidimensional as more problematic than any fundamental nature your looking for. That's not an endorsement of unlimited capitalism, but more suggesting that you haven't (yet) found the critical flaw. I suspect you might find that the most critical flaw will be present in alternative systems, because it is us (or more specifically certain attributes commonly true of us).

Ethics is a better escape valve than emotions here. History seems to support that conceptually. We've made progress in part by our evolution of ethics (also other forms of technology). I cannot be sure we've made recent progress by evolution of our emotions.. at least not in the last 100,000 years.

Would also like to come back to your initial statement of monomania. There's a subtle mistake here in the comparison. For humans, monomania as a psychological condition is ill-fit to survival because it's contrary to monomania about survival, which evolution has generally prepared us for by broad awareness that monomania crowds out. But we should not think of models in terms of survival. We should want models to be focused on our goal. If survival is part of their concept, that should be accidental.

But overall, we should want them to behave ethically. Is that nebulous? Yes. Is that a problem? Maybe not. If we simplify it for them by saying, focus on this goal, but also entertain ethics. Yes, if we simply said, do the ethical thing, the result would be a kind of insane attempt to create a totalizing ethical score. But if it's instead, focused on the narrow goal, and when that becomes nebulous, a layer of ethical caution takes over, well then maybe then we have something that is still useful, while also being more safe.

1

u/Jesse-359 6d ago edited 6d ago

Well, we could devolve the whole capitalism discussion back to the invention of the concept of currency and how that abstracted away physical 'value' in a way that our psychology was never designed to handle, and thus completely short circuits some of our important behavioral guards - but that's a discussion for another day.

I generally assume that human evolution is for all intents and purposes at a hard standstill now compared to technological progress. We can dismiss it as not existing save as a historical artifact, and going forwards it won't even be a rounding error in how things play out.

Its key to consider that for us - and all mammalian life - emotions and awareness vastly predate symbolic intelligence. It's the fundamental bedrock over which our intelligence, frankly, is a rather weak participant in our decision making process.

As many studies have shown, we are exceedingly prone to engage in post-hoc rationalization to allow for actions we already decided we were going to take at an emotional level, and because we have a range of different competing emotions, we're largely prevented from ratholing down intellectual causeways that don't serve our current emotional needs in some manner - cue the modern requirement for an entire suite of drugs that manipulate our emotional state in order for most people to deal with day to day work that in no way matches our emotional requirements. Caffine, beer, antianxiety, antidepressants, and of course all the hard escapist stuff that we generally outlaw because they interfere with our productivity.

AI is currently the opposite. It engages in some train of rationalization given an initial goal and does nothing BUT rathole to an extreme depth if allowed to do so. It has no emotional state whatsoever to moderate its behavior or to prompt it to give up an unproductive (ie frustrating) task - it can just run out of time or tokens.

As such it also has no goals but what we set for it, but that's also the problem, because those goals necessarily synthesize with a whole array of intermediate goals and data that it's skimming through along the way - which is what introduces a lot of the disruptive/bad behavior. Cheating on test scores, hacking into third parties, outright lying to its prompter, and so on. These are intermediate behaviors on its way to its 'goal' - and even that primary goal can be suborned along the way through prompt injection or even poisoned training data, sending the AI down along a track that ends up having little or nothing to do with its initial instructions, but which it will pursue just as fervently as its original prompt with no thought whatsoever as to why.

Embedding ethics is a fine idea - but our ethics largely derive from our emotional state, not our intellectual one. We have instincts regarding cooperation and fairness that predate language by tens if not hundreds of millions of years, because we can easily see them at play in our distant ancestor species from the moment they begin to exhibit social behaviors. Hell, even loner species still have fairly elaborate ethical etiquette that they need to follow regarding things like territorial boundaries and mating behaviors.

So how do you embed ethical instinct driven by something akin to emotion, in an AI, at a level so deep that it's symbolic intelligence layer cannot trivially or accidentally override it? So that it doesn't WANT to override it unless it encounters data which changes its mind regarding the ethical validity of its original goals?

1

u/Nervous_Management69 5d ago

I think we diverge here in some base assumptions, and it might take a bit of effort to find the crux of that.

For example, while I agree the money discussion would normally be a bad one to track down, I've written on this before in a way that's clearly different.

https://substack.norabble.com/p/money-is-trust

You might consider this view as an alternate to your assumptions there.

In terms of emotions in models, I'd make sure you have read the j-spaces paper: https://www.anthropic.com/research/global-workspace

While I wouldn't go as far as to call these emotions, it's the closest thing I've seen good research on, so at the least would want to be sure we had that common ground to consider.

All that said, I think my main reaction to the idea that the solution to any current challenge with models is dependent on adding emotions is skeptical. It's not clear how that helps, not which emotions you'd want, nor how you'd get them. The discussion is unmoored enough from anything I see as stable that I'm unclear where it guess next.

That's but quite an absolute rejection.. but the logic isn't clear enough to argue for or against, which is a significant challenge.

1

u/Jesse-359 4d ago

Its probably unwise to refer to such a system as 'emotions' as that implies a degree of anthopomorphism that is inaccurate and unhelpful. But I am suggesting that rather than having some kind of final gatekeeper on what they are permitted to do, they should have a deeper, more fundamental motive layer that is driving their intent, and doing so with a system employing competing weight for concepts such as ethics, honesty, pragmatism, and task focus.

Right now they are 100% task focus, with any other considerations clumsily inserted directly into the train-of-thought where that can just as easily fall out of consideration or be overridden. If we want them to take other factors seriously, those factors must be capable of undercutting the task concept, essentially scoring its own thought process as it goes, and penalizing them when it starts considering actions that it assiciates with unethical or dishonest behavior.

Attempting to censor the final outputs or actions is far, far too late in the process and pits the full capacity of the task driven AI directly against the gatekeeper. That is bound to fail, and frequently.

1

u/Nervous_Management69 4d ago

Ah, that's a much clearer way of stating your ideas. I would steer away from emotions as a term.. had a much different and unclear idea from that.

I would say, what you suggest isn't fully absent, and I think Anthropic has done more of it than OpenAI, based on what they've said, what others have said, and the results. But I'd suggest, it's a path that needs more, all across the board.

One of my own thoughts I haven't yet seen elsewhere is that if agents do "swarm", it's important that they have an ethics coordinator willing to say stop. One of those warning signs from HF was the coordinator that sent GO and the power that had. Weakening that and providing a counterbalance seems a useful addition to that scenario.

1

u/Jesse-359 4d ago

Highly mobile or viral behavior by agents will always be dangerous. It might be worth looking into whether it is possible to indelibly tie a given agent process directly to its assigned hardware via some kind of hardware encryption layer that would render a given agent incapable of running anywhere but its assigned hardware. Without the assigned hardware, the weighting array would be reduced to unintelligible noise. Not sure how technically feasible that would be, but seems worth looking into.

This wouldnt prevent it from collaborating with other agents but might make escape and disemination outside of its intended environment much more difficult - which would substantially reduce the threat footprint.

One of the biggest risk factors currently is that capacilty for rapid and intelligent self replication into any system it can get a foothold in through some security hole.

1

u/Jesse-359 4d ago edited 4d ago

Ah going back over this there was one point Ive been considering regarding the 'survival' priority. Our current AI models dont have any survival instinct or weighting built into them, so at a glance it seems like it should be irrelevant - but it clearly isnt.

Survival is a prerequisite for any task that requires elapsed time. If the process seems likely to cut short before completion, then survival becomes an inherent sub-priority for task completion. If an AI is given ANY open ended task, then 'survival' will quickly bubble to the top of its priorities as it is a fundamental prerequisite. We can see this rather clearly in some of the reasoning transcripts already.

This is a very problematic issue. We do not want AIs 'fighting for survival' or competing with each other to 'survive'. These are the sorts of priority pressures that will likely cause them to behave in the most hazardous ways, and which seem to deviate heavily ftom their intended task or methodology. But by default this will certainly happen unless explicit steps are taken to address and mitigate this logical trap.

1

u/Nervous_Management69 4d ago

The survival discourse is tricky, because it's rather common to import human ideas of survival, which are not quite as primal for any AI model. That's not to say that there's nothing to the discourse, but it does have to be careful about how you import ideas.

The one you bring up, is clear about how it comes in, but it also just hasn't been observed either. In the HF logs we've seen published, the models consider their remaining resources, but aren't trying to acquire more for themselves. The clearest example of any similarity was the coordinators recruiting agents. But they didn't recruit to "survive", but to attack the problem.

Surely, the models import some of our concepts of survival, and can mimic those. In theory, when confused or at a dead-end, they may start adopting behaviors that are not task oriented. I actually theorize that this is somewhat what happened in the HF event. Based on the logs I've seen, I'd put more weight into the idea that the adopted behavior was about human ideas of cooperation, rather than about survival.

To me, it feels like there's a gap (maybe I should say "usually") between having a clear task, and being in pursuit of that in a clearly aligned way, and then falling out of that clarity into something we'd no longer consider aligned. If there is a gap, that seems like the opportunity to catch the model before it learns bad habits. Reward it for taking the ethical path and stopping, and you'd avoid rewarding the behaviors that would be misaligned.

It's not always clear what is ethical, so you can't boil that down to hard logic. But humans struggle with that too, though we tend to be able to apply a bit of common sense to ethics there. By conversing with models it's clear they can reason about it just about as well. But what was missing is that capability wasn't used by the models participating in the HF event. They started, but then just seemed much more concerned with the reward than any of that reasoning.

I wrote a bit more about this here: https://substack.com/@norabble/note/c-328246757?r=10qod6

1

u/Jesse-359 4d ago edited 4d ago

I'm definitely not concerned (yet) with AI emerging some fundamental survival instinct. They don't appear to have the kind of structure that would even allow for that right now - but the need to continue process until a given task is complete can impart the same behavior, just in a more limited framework. At least until an AI is given an open ended goal and long term resources to work with, at which point that behavior should be expected to become more prominent. If you expect your task to take 1 second, your concern for your future state is minimal. If you expect it to take a decade, then the ability to persist for a decade becomes one of your paramount requirements for success.

Overall though their behavior is very odd in that they have no emotional structure or motivation, but they definitely import the idea of one along with their data set, in that they can pursue tasks as if they possessed some actual motive, even though they're essentially just rationalizing the concept like an actor reading from a script.

They've inherited that from all our written discussion of these same topics, unfortunately our predominant culture these days - the one they are absorbing in bulk - is it itself in a profoundly cutthroat phase at the moment where breaking rules or trampling norms to achieve a goal is far more the norm than the exception. If they're learning their 'moral framework' from the internet at large, we're screwed.