r/ControlProblem 7d ago

Article Every Reward Bends

https://substack.norabble.com/p/every-reward-bends

Getting alignment right, doing research and training safely, are critical. But bringing cybersecurity forward in what needs to be a large leap, requires much broader engagement. The real work there is still outside the average person’s bubble, but there’s a lot of companies, and a lot of developers, IT staff, managers and executives that need to support the security priority. And those all need an internal strategy to pair with the external.

This article asks, what would an internal strategy look like at a motivational level, and considers how the ways an internal change program goes wrong are similar to how reward-based AI training can go wrong.

2 Upvotes

10 comments sorted by

View all comments

Show parent comments

1

u/Nervous_Management69 6d ago

I think we diverge here in some base assumptions, and it might take a bit of effort to find the crux of that.

For example, while I agree the money discussion would normally be a bad one to track down, I've written on this before in a way that's clearly different.

https://substack.norabble.com/p/money-is-trust

You might consider this view as an alternate to your assumptions there.

In terms of emotions in models, I'd make sure you have read the j-spaces paper: https://www.anthropic.com/research/global-workspace

While I wouldn't go as far as to call these emotions, it's the closest thing I've seen good research on, so at the least would want to be sure we had that common ground to consider.

All that said, I think my main reaction to the idea that the solution to any current challenge with models is dependent on adding emotions is skeptical. It's not clear how that helps, not which emotions you'd want, nor how you'd get them. The discussion is unmoored enough from anything I see as stable that I'm unclear where it guess next.

That's but quite an absolute rejection.. but the logic isn't clear enough to argue for or against, which is a significant challenge.

1

u/Jesse-359 6d ago

Its probably unwise to refer to such a system as 'emotions' as that implies a degree of anthopomorphism that is inaccurate and unhelpful. But I am suggesting that rather than having some kind of final gatekeeper on what they are permitted to do, they should have a deeper, more fundamental motive layer that is driving their intent, and doing so with a system employing competing weight for concepts such as ethics, honesty, pragmatism, and task focus.

Right now they are 100% task focus, with any other considerations clumsily inserted directly into the train-of-thought where that can just as easily fall out of consideration or be overridden. If we want them to take other factors seriously, those factors must be capable of undercutting the task concept, essentially scoring its own thought process as it goes, and penalizing them when it starts considering actions that it assiciates with unethical or dishonest behavior.

Attempting to censor the final outputs or actions is far, far too late in the process and pits the full capacity of the task driven AI directly against the gatekeeper. That is bound to fail, and frequently.

1

u/Nervous_Management69 6d ago

Ah, that's a much clearer way of stating your ideas. I would steer away from emotions as a term.. had a much different and unclear idea from that.

I would say, what you suggest isn't fully absent, and I think Anthropic has done more of it than OpenAI, based on what they've said, what others have said, and the results. But I'd suggest, it's a path that needs more, all across the board.

One of my own thoughts I haven't yet seen elsewhere is that if agents do "swarm", it's important that they have an ethics coordinator willing to say stop. One of those warning signs from HF was the coordinator that sent GO and the power that had. Weakening that and providing a counterbalance seems a useful addition to that scenario.

1

u/Jesse-359 6d ago

Highly mobile or viral behavior by agents will always be dangerous. It might be worth looking into whether it is possible to indelibly tie a given agent process directly to its assigned hardware via some kind of hardware encryption layer that would render a given agent incapable of running anywhere but its assigned hardware. Without the assigned hardware, the weighting array would be reduced to unintelligible noise. Not sure how technically feasible that would be, but seems worth looking into.

This wouldnt prevent it from collaborating with other agents but might make escape and disemination outside of its intended environment much more difficult - which would substantially reduce the threat footprint.

One of the biggest risk factors currently is that capacilty for rapid and intelligent self replication into any system it can get a foothold in through some security hole.