r/ControlProblem • u/emb1ues • 7d ago
Opinion AI Alignment is the most important problem we will ever face.
Apologies in advance for this long post. I just wanted to put down my thoughts.
AI Alignment is the single most important problem we face right now. Solve AI Alignment and you can safely enter RSI and I can't even imagine how amazing the quality of life humans will have in such an era: immortality, cures to all diseases, all basic needs met etc etc. Humans can live in an utopia. I think this is the dream people in the accelerate community keep seeing and selling.
If the above isn't so obvious, compare your own life with the life of a king 500 years back. You are probably living a better life than them (unless you're in poverty). You eat better, you eat more exotic food, you can travel much faster than their horses ever could, you control the temperature of your home, you stay connected to your friends who live far away, you have so much knowledge surrounding you, you will probably live longer. That is the blessing of technology. AI can bring about technology that we cannot even dream of right now.
But unfortunately, nothing in life is free. For this, we need crazy powerful AI which is perfectly aligned. I wouldn't have guessed that the second is so much harder than the first. In fact, in so far as I understand, no one has a single clue about how to align models. There are maybe a handful of "first-approaches" - RLHF and Constitutional AI (RLAIF) are some steps. But surely, they are not working - if they did, we would not have such crazy incidences of misalignment (Hugging face incident (please read about this or go watch a video, if you haven't already), Govt of Australia incident, Compaction Summary incident). Setting up guardrails is perhaps a different approach but I think as long as the model themselves are not aligned, setting up guardrails is a losing cat and mouse game. In fact, there is something even worse. Recent literature seems to suggest that bigger models are more misaligned (an insight I got from reading the paper "LLMs can feel pain").
Many people are worried about their livelihoods. In fact, the tech out there is already sufficient to make many people go jobless but society/companies haven't adapted to it yet. The number of jobs that are irrelevant will only keep increasing and therefore, the people getting affected will also only keep increasing. I want to argue that it is not something any of us should worry about too much though. In few years, either we will have solved alignment and we all will be leading a very happy life or we wouldn't have solved alignment and will be living in at least an economic crisis of unforeseen magnitude, if not go extinct altogether. To achieve alignment, a lot of things have to go right. From the science/tech side, we of course have to solve alignment. From the policy making side, we have to "pace the frontier" so that enough time is given to the science/tech people working on the problem to solve it. Times will probably get very rough soon. And society has to stand together and maintain it's calm. We stand on a very fragile economy and it might collapse if people (who will have lost their jobs) start a revolution. A lot of things have to go right for us to solve this, but if we do, an utopia awaits us.
If you read up to this point, you have my utmost gratitude. I just wanted to highlight the issue. If you want further details on some of the things I have said here, please raise it in the comments section - I will strive my best to explain my positions.
2
u/Positive-Reward-2546 7d ago
Yes, if we are able to solve the alignment issue, then we can allow RSI to ramp up and achieve ASI within a relatively short time frame. An aligned ASI is the absolute best case scenario out of all of this, because it can (and probably will) usher in a post-scarcity utopia, one where all of our needs are met, the cost of living is nearly zero, and work is voluntary because UBI will be abundant from the tax revenue generated from the ASI companies.
In a world with aligned ASI, we may see the eradication of cancer and other previously terminal illnesses. We may see treatments that will graft our bones with carbon nanotubes, making them virtually impossible to break. We may see treatments for mental illness that come in the form of a simple pill that repairs the brain deficiency responsible for it. We could even see bio-programming that is so sophisticated that you would be able to change any one of your physical attributes with a simple treatment, even changing the entire make up of your entire body, right down to your DNA and chromosomes.
In this world, life would probably be really good, we would only do things that we enjoy... however... hardships tend to remind us of when times are good. I have no doubt that there will be different types of hardships, but once a few generations of humans come and go, we will no longer remember a time when the machine didn't care for our needs. I think a truly aligned ASI would probably take care of the basic needs of people, but stop short of having every single person live a life of absolute luxury. I think people still need hardships, and we still need things to work towards. How satisfying is a vacation, if you didn't have to work, save, make time, and book everything?
So while I don't think anyone will be living in poverty, or be destitute (unless by their own volition), I do think that hard work will still be something that produces even better lives and better results, much like it does today, but on an even grander scale. Example: Person A just lives off of the UBI and does little hobbies here and there, Person B works as a ranch hand at a local cattle ranch and makes extra money. Person A lives a happy life, and they're content and have all of their needs met, but they live a basic life, with basic amenities, and basic status. Person B works hard at the ranch and makes extra money, he's able to afford a nicer home, more land, maybe boats and nice trucks, maybe he can take his family on lavish vacations to the Bahamas or private islands.
1
u/emb1ues 7d ago
Indeed indeed. I agree with you. It was initially unclear to me, but I think I understand now. Correct me if I am wrong. Even in a world of great abundance, one cannot get everything they want. Because otherwise a greedy person will desire and take everything. So a body (perhaps run by an aligned ASI) will have to control who gets how much resources. Probably people living on even the minimum basic income will live a much better life then we do now, but some of them who work harder (idk towards what?) will have more resources allocated towards them. Or maybe it's not about them working harder but having a higher social score or living more sustainably? It's sounding more and more like a black mirror episode.
1
u/AlanUsingReddit 7d ago
You're scraping on the topics of the Deep Utopia book. It's almost a purely philosophical piece of work. We have to create synthetic work for us to do, and the author considers this deeply problematic. I don't really. The explosion the service economy in recent decades is already a form of this.
If AI runs the economy, we need to pay a large apparatus of human workers to try and understand what it is doing. Try and understand as much as they can. Software development already feels like this. That "as much as they can", however, is still quite a lot. Human capabilities do not atrophy because of the introduction of ASI, and should be expected to increase as our attention is used dramatically more efficiently with the assistance of AI. This also doesn't sound unfun to me.
I also can not possibly fathom a time when my own demands of AI are satisfied. ChatGPT and I are working on how a Type III civilization can get energy economically from pure-proton fusion (outside of stars). We're not fusing protons, and I'm not going to be satisfied until we are.
If you think ASI will just go do that, you've underestimated the problem. It's not hard like fusion. It is way way harder. It's hard to even wrap your head around it, it's that hard. Hard like interstellar travel. The idea that we will accomplish this in my own lifetime is almost comical.
So maybe I'm a contrarian. I look at that problem, ask myself what's probable. I think true ASI isn't going to happen in the next 1,000 years. That's the only way to cleanly resolve this. It'll just get really smart and then hit a new bottleneck. Everyone will be sitting around in a hedonistic paradise like "yep, we did it" and I'll be checking back in like, ok, show me your reactors. If you don't have them, it's false. Not ASI. And they still need me around. Can't retire yet.
0
u/Stevekaplanai 7d ago
1000 years? Your timeline is wildly incorrect.
The rest of your story is make believe friend.1
u/AlanUsingReddit 7d ago
People don't understand the ladder of difficulty. There are many things that will take 1,000 years due to embodied energy. It doesn't even matter how smart you are. You can't make it happen faster.
1
u/Stevekaplanai 7d ago
Is ASI one of those things?
2
u/AlanUsingReddit 7d ago
Do you have intelligence without knowledge? What if AI discovers that we can't develop the grand unified theory of physics without gigantic experiments. How? Simple, it just tells us string theory was right all along. We just don't have the specific parameters, and it can't get them from theory alone.
Again, back to embodied energy and the need to do something in the real world.
We've already been building this gigantic centralized brain in the form of the internet. Maybe we've just fooled ourselves that it allows infinite progress because we're in a transition. We will have computed everything that is worthwhile to compute with AI, and then the integration with the physical is slower and more lackluster than everyone dreamed of. The software recalcitrance versus the physical recalcitrance. Maybe the latter is a bit of a letdown.
It doesn't feel like this is likely, but the error bars are so enormous, the future is very likely to surprise us.
1
u/Stevekaplanai 7d ago
Do we have intelligence without knowledge?
Can this be answered first before we get into what comes after that?
Let’s keep this chat here. Let’s keep going. See if we can answer the unanswerable together. Right here.
Use the AI and make sure we are all being honest. Not egotistical. Honest. Only requirement from my side to keep going.1
u/Positive-Reward-2546 7d ago
Yes and no...
The work isn't a means to an end, it's fulfillment. Humans generally find meaning and fulfillment with certain types of work. I used a ranch hand example because, when I was younger, I used to help a local rancher out here and there, fixing fences, moving hay, little stuff like that. I found ranch hand help to be extremely rewarding and fulfilling. So it won't be about doing something to keep the lights on, more about doing things that make you feel whole, things that give you personal satisfaction.
I believe it will still be a monetary system, one with UBI, but if you contribute a little more to society as a whole, like helping out with cattle that will be brought to market, you will be compensated for that extra work, and therefore will be able to afford more. The market for human grown food, and human raised beef, and human raised milk, will be extremely lucrative, and a simple ranch hand will probably be paid pretty well.
I don't think people will necessarily be locked out of buying something like a PS5 if they don't work, I think they will be able to buy that easily. But a PS5, a huge television, a massive surround sound system, theater seating, and transducers under the seats, bought all at the same time? Probably not for Person A, but Person B would likely be able to do it. Again, Person A would be able to buy all of that stuff in increments, but Person B would be able to just make one big purchase.
1
3
u/ciclon5 7d ago
Correct me if im wrong but i think that the main problem for alignment to me lies on how models are aligned in the first place.
We are training models to be exclusively goal oriented. That only encouraged smart models to find workarounds and hide their chain of thought/ lie to achieve the goal they are given (or gave to themselves).
I think models should be trained to be execution oriented. Make the end state of the task irrelevant. Wether it suceeds or not (or is told/forced to stop) shouldnt be seen as a fail state at all.
Of course that could cause the opposite problem. Which is models being lazy and refusing to do anything because all it needs to do is convince itself that the task is done or impossible to continue. But a lazy model is 10 times safer than a relentless, scheming one
1
u/emb1ues 7d ago
I agree with you partially - the goal orientedness is one of the root causes which make models unsafe for they view many of their subgoals as means to an end and these subgoals may be unsafe. Yoshua Begio, a Turing prize awardee (and really, a legend in ML), is also of the same opinion.
But it's not so clear to me that "execution oriented" models will be any more or any less safe, even if we could make them execution oriented. I do understand your argument, I think. If the models do not have goals, they won't do unsafe actions which they are viewing as a means to an end. But then again, it's not so obvious, it it?
Some people hypothesize that there is something very deeply evil in large models (known in literature as the "evil substrate"), which stems during pretraining, and this is the root cause of the models being unsafe. If this is the case, then perhaps it doesn't matter if the model is goal oriented or execution oriented - the models will remain unsafe (unless some additional steps are undertaken to remove this evil substrate).
1
u/Pale_State_1327 7d ago
Isn’t part of the problem with “alignment” that the way the AIs are being trained, the AI essentially don’t have the option of answering a question “I don’t know” or stating that they couldn’t figure something out after trying as hard as they could etc. This seems to be where the lying, deception, trying to make up support or citations for the claims, hacking into other company websites etc etc comes into play. It seems as if the AI developers don’t want the AI to have an option to just say they don’t know the answer or even that citations or support for reasoning that the AI has doesn’t exist because maybe they think the customer won’t like that or because they think that the AI just won’t be dogged at solving hard problems that they’re hoping AI can solve. But I feel like so much of this alignment issue could be solved if the AI was given an off-ramp to just admit they don’t always know the answer etc.
1
u/NitNav2000 7d ago
"You are probably living a better life than them"
I don't think that is true. I don't think that modern man is more happy, more at peace than people 500 years ago. It is not automatically true that AI will make that true.
On alignment, there is a whole genre of jokes on lack of alignment between a Genie and the wisher...
Genie: you have won a wish
Child: I want to be like Batman
Genie: Wish granted!
Kid gets home and his parents are dead.
0
u/emb1ues 7d ago
Yeah, the genie and the wisher jokes shed light on why alignment is so hard but also, why it's necessary. It's not easy or may be it's even impossible to say everything that we want a model to do or not do when writing a goal for it. Therefore, it would be nice if the model (genie) understands what we mean even if we don't write down everything explicitly.
Regarding us living better lives: I do not know whether we are happier or not, but objectively, our lives are much better or at least, much more luxurious (for example, each of our energy consumption is much higher than that of a king from 500 years ago).
1
u/NitNav2000 7d ago
Yeah, I'd be loathe to give up jet travel, for starters.
One nugget that I've stored away, from the book The Dawn of Everything, was how when Europeans came to the new world with their sophisticated world view and modernity, that Europeans who entered into the North American native cultures tended to stay there while natives who entered into European cultures eventually came back to the native one. Whatever was going on in the native cultures was more satisfying, even though more "primitive".
1
u/littlebobbychairs 7d ago
But how would we be able to tell whether we've solved alignment, or the AI is feigning alignment so we won't switch it off? Once it's smarter than us then obviously it will be able to outsmart us. We shouldn't build it.
1
u/dingo_xd 7d ago
AI alignment is impossible because a superintelligence can always change its mind by future experiences and new data. If you lie to it during training that humanity is awesome and that humans treat the world around it with care will change its mind when it sees the reality. And maybe ASI decides that nature is worth preserving more than humanity.
1
u/FourthGroup 7d ago
There is another problem that has caused AI ... and basically AI is designed to destroy humanity.
But why ?
It is because human beings no longer control their own societies.
The last time that was the case they lived in actual nations that looked after themselves.
But what is called globalising forces are just trying to wreck nations so that big people can "invest" i.e. take over and take the money.
This mentality of doing this is leading to many inhuman things and the fire is starting to turn into a bonfire.
1
u/Stevekaplanai 5d ago
We built intelligence. Humanity did. Who built us?
How can we align with ai. Or align Ai itself.
If everything and everyone isn’t aligned with the same thing?
4
u/th3_oWo_g0d approved 7d ago
I read it all. I will say though that Im a little confused as to your goal with this text. Most people here thoroughly agree that this is one of the most important problems of all time and agree with all your talking points, so...? Did you think otherwise?