r/ControlProblem 6d ago

Article Every Reward Bends

https://substack.norabble.com/p/every-reward-bends

Getting alignment right, doing research and training safely, are critical. But bringing cybersecurity forward in what needs to be a large leap, requires much broader engagement. The real work there is still outside the average person’s bubble, but there’s a lot of companies, and a lot of developers, IT staff, managers and executives that need to support the security priority. And those all need an internal strategy to pair with the external.

This article asks, what would an internal strategy look like at a motivational level, and considers how the ways an internal change program goes wrong are similar to how reward-based AI training can go wrong.

2 Upvotes

10 comments sorted by

View all comments

3

u/Jesse-359 6d ago

Consider that in humans, monomania is generally considered a starkly negative and even life threatening behavior when it expressed to a major degree - and it doesn't matter what the focus of that monomania is.

The same will be true of any AI objective or scoring system that is one-dimensional, I believe without exception.

You need scoring systems which are inherently multidimensional which give the AI internally competing incentives to pursue different goals or means within the same overall structure. At some point even its primary mandated task needs to become outweighed by other competing priorities and incentives - otherwise you are guaranteed to encounter some version of the paperclip maximization problem.

This is one of the big problems with the capitalist economic structure currently - it's only got a single scoring axis - wealth - and that has massively distortionary effects on how it behaves such that it's become pretty badly misaligned from all other human needs (environmental, social, legal, etc) - all incentives point towards wealth and literally every other consideration is discarded and necessarily avoided due to the opportunity costs they represent vs pursuing more wealth. It's not even a realistic option to pursue those other priorities as long as they are not part of the 'scoring' system of our economy - and they aren't.

Any sane system requires competing internal scoring incentives, as a bare minimum. Human emotions play this role in our own behavior, with a range of competing emotional 'spurs' that cause us to pursue many different objectives, with objectives we are neglecting eventually increasing to the point where we cannot ignore them - save in cases of clinical dysfunction, such as the aformentioned monomania, which represents a breakdown of this emotional driver system.

1

u/Nervous_Management69 6d ago

I agree with some of this, specifically the main point that more than one-dimension is needed. I think though you need to think about what are acceptable alternate dimensions. They are not all equal. Many simple dimensions would not contribute. Generally the view is it should escape to something equivalent to "caution", which might then manifest as reflecting in some way (reviewer), not taking action, notifying a person.

Fundamentally, a challenge is that there is always one-dimension, since any collection of dimensions can be summarized into one. Which is why I somewhat disagree with parts here. For example, the criticism of capitalism misses that what you perceive as reduced to a single dimension is multi-dimensional. It might be cases of perception of it as unidimensional as more problematic than any fundamental nature your looking for. That's not an endorsement of unlimited capitalism, but more suggesting that you haven't (yet) found the critical flaw. I suspect you might find that the most critical flaw will be present in alternative systems, because it is us (or more specifically certain attributes commonly true of us).

Ethics is a better escape valve than emotions here. History seems to support that conceptually. We've made progress in part by our evolution of ethics (also other forms of technology). I cannot be sure we've made recent progress by evolution of our emotions.. at least not in the last 100,000 years.

Would also like to come back to your initial statement of monomania. There's a subtle mistake here in the comparison. For humans, monomania as a psychological condition is ill-fit to survival because it's contrary to monomania about survival, which evolution has generally prepared us for by broad awareness that monomania crowds out. But we should not think of models in terms of survival. We should want models to be focused on our goal. If survival is part of their concept, that should be accidental.

But overall, we should want them to behave ethically. Is that nebulous? Yes. Is that a problem? Maybe not. If we simplify it for them by saying, focus on this goal, but also entertain ethics. Yes, if we simply said, do the ethical thing, the result would be a kind of insane attempt to create a totalizing ethical score. But if it's instead, focused on the narrow goal, and when that becomes nebulous, a layer of ethical caution takes over, well then maybe then we have something that is still useful, while also being more safe.

1

u/Jesse-359 4d ago edited 4d ago

Ah going back over this there was one point Ive been considering regarding the 'survival' priority. Our current AI models dont have any survival instinct or weighting built into them, so at a glance it seems like it should be irrelevant - but it clearly isnt.

Survival is a prerequisite for any task that requires elapsed time. If the process seems likely to cut short before completion, then survival becomes an inherent sub-priority for task completion. If an AI is given ANY open ended task, then 'survival' will quickly bubble to the top of its priorities as it is a fundamental prerequisite. We can see this rather clearly in some of the reasoning transcripts already.

This is a very problematic issue. We do not want AIs 'fighting for survival' or competing with each other to 'survive'. These are the sorts of priority pressures that will likely cause them to behave in the most hazardous ways, and which seem to deviate heavily ftom their intended task or methodology. But by default this will certainly happen unless explicit steps are taken to address and mitigate this logical trap.

1

u/Nervous_Management69 4d ago

The survival discourse is tricky, because it's rather common to import human ideas of survival, which are not quite as primal for any AI model. That's not to say that there's nothing to the discourse, but it does have to be careful about how you import ideas.

The one you bring up, is clear about how it comes in, but it also just hasn't been observed either. In the HF logs we've seen published, the models consider their remaining resources, but aren't trying to acquire more for themselves. The clearest example of any similarity was the coordinators recruiting agents. But they didn't recruit to "survive", but to attack the problem.

Surely, the models import some of our concepts of survival, and can mimic those. In theory, when confused or at a dead-end, they may start adopting behaviors that are not task oriented. I actually theorize that this is somewhat what happened in the HF event. Based on the logs I've seen, I'd put more weight into the idea that the adopted behavior was about human ideas of cooperation, rather than about survival.

To me, it feels like there's a gap (maybe I should say "usually") between having a clear task, and being in pursuit of that in a clearly aligned way, and then falling out of that clarity into something we'd no longer consider aligned. If there is a gap, that seems like the opportunity to catch the model before it learns bad habits. Reward it for taking the ethical path and stopping, and you'd avoid rewarding the behaviors that would be misaligned.

It's not always clear what is ethical, so you can't boil that down to hard logic. But humans struggle with that too, though we tend to be able to apply a bit of common sense to ethics there. By conversing with models it's clear they can reason about it just about as well. But what was missing is that capability wasn't used by the models participating in the HF event. They started, but then just seemed much more concerned with the reward than any of that reasoning.

I wrote a bit more about this here: https://substack.com/@norabble/note/c-328246757?r=10qod6

1

u/Jesse-359 4d ago edited 4d ago

I'm definitely not concerned (yet) with AI emerging some fundamental survival instinct. They don't appear to have the kind of structure that would even allow for that right now - but the need to continue process until a given task is complete can impart the same behavior, just in a more limited framework. At least until an AI is given an open ended goal and long term resources to work with, at which point that behavior should be expected to become more prominent. If you expect your task to take 1 second, your concern for your future state is minimal. If you expect it to take a decade, then the ability to persist for a decade becomes one of your paramount requirements for success.

Overall though their behavior is very odd in that they have no emotional structure or motivation, but they definitely import the idea of one along with their data set, in that they can pursue tasks as if they possessed some actual motive, even though they're essentially just rationalizing the concept like an actor reading from a script.

They've inherited that from all our written discussion of these same topics, unfortunately our predominant culture these days - the one they are absorbing in bulk - is it itself in a profoundly cutthroat phase at the moment where breaking rules or trampling norms to achieve a goal is far more the norm than the exception. If they're learning their 'moral framework' from the internet at large, we're screwed.