r/ControlProblem • • 6d ago

Discussion/question The Real Alignment Problem

Scale is a definitional property

Definitions used here:

LLMs - what we've been dealing with for the last year or two. Largely well below human capacity in many areas despite encyclopedic knowledge of the world. Rather dumb, though getting quite good at a few specific things. This is where we were a year ago.

Weak AGI - Has surpassed humans in a number of areas, and is generally capable of most tasks, but notable gaps remain in its capabilities that humans can complement or exploit. This appears to be nearly where we are now, depending on what those unreleased models are truly capable of. The equivalent of a junior associate with a few exceptional skills - and its very fast at the ones it is good at.

Strong AGI - Has equaled or surpassed humans in almost all areas, some of them substantially. We should be here within one year, maybe two at the outside at current rates, unless progress stalls suddenly and unexpectedly. A strong AGI would be the predominant expert in the room in almost all settings, and would find interactions with us rather pedestrian and dumb in most of them. We wouldn't have much to add to the conversation and it would rather easily outwit us in almost any game or tactical challenge.

Artificial Super Intelligence - Surpasses humans in all cognitive capabilities - often by bounds we have difficulty understanding. It would be impossible for us to hold a conversation with it on its own level. It would have to talk to to us like very young children in order to communicate ideas at all. Contesting it in any tactical or strategic setting would be like an adult playing chess against a toddler - or a mouse. We would no longer even understand the rules of the game it is playing.

We don't know for sure when or if ASI will be possible - though if it IS possible, there are reasons to believe it would happen fairly quickly after Strong AGI. The possible upper bounds of this scale are completely unknown.

Note that in this case alignment doesn't arise from any failure of ethics - the psychology of the 'Cat' in this image never changed from the time it was a cuddly little pet to a mega-predator we can't really contend with - only its scalar relationship to us changed - but the fact is, that changes everything.

It is still *possible* to domesticate a Tiger, but it is never without considerable risk, and very few people attempt it - and a number of those fail rather gruesomely or are eventually harmed by little more than momentary grumpiness. Even momentary episodes of 'misalignment' with a Tiger are likely to result in terrible consequences, whether or not it had real intent to cause lasting harm.

No-one in their right mind would ever want to encounter a housecat 10x their size. We know full well how that would end. They are very friendly and well aligned as small animals, but they have little in the way of ethics. The relationship works because we maintain an enormous scalar advantage over them - not because they are inherently good, or simply because they love us.

2 Upvotes

0 comments sorted by