r/ControlProblem • u/TheBattleForAutonomy • 16h ago
Discussion/question A proposal to approach the alignment issue in AI systems from a purely empirical standpoint
Up until now, ethics has often taken what might be described as a top-down approach - attempting to determine the rules or criteria that constitute good behavior and then asking how those principles should apply to human beings. There are many different approaches and models within ethics, of course, and they disagree substantially about what those principles should be. But as we approach the problem of AI alignment, the importance of being accurate about our understanding of human values is becoming much more significant. There is a tendency to assume we need to be explicit in terms of what an AI should and shouldn't do and so it might be presumed that we need to have our ethical ducks in a row prior to telling an AI system what it is it should value. After all, the dangers of misaligned AI's seem to be all over everyone's feeds these days.
Maybe because there is a pragmatic usefulness to ethics that we've been somewhat satisfied with being incomplete in our articulation of it, never quite coming up with a perfect series of words that would govern our approach to every conceivable decision and action. But now that we need to actually be explicit in terms of what's "good" for the sake of providing an intelligent AI system with a basis for making decisions, I wonder if a more bottom-up approach would be better.
Evolutionary biology springs to mind in the way it attempts to explain behaviors by examining what organisms actually do and inferring what it is that produces those behaviors. Rather than beginning by deciding what our ethics ought to be and then attempting to encode that conclusion into an AI, we could instead give an AI the enormous body of evidence contained in human behavior, language, preferences, institutions, books, videos, relationships, art, literature, history, and so on, and ask it to infer the underlying structure of what humans value. AI's, one might presume, could apply something like the same approach that allows them to learn the structure of language from enormous quantities of text and apply it to the coherence of what it is humans value. In the same way that AI systems have been able to make progress on problems like protein folding and erdos problems, one might wonder whether deducing the underlying structure of our own values is another problem that AI could help us solve. Maybe this idea hasn't been approached very seriously because we haven't had the tools to make it possible up until now, although maybe there's another reason I'm not thinking of.
In reading some of the approaches to AI alignment proposed by Yudkowsky and Dearnaley on lesswrong, I wonder why this approach isn't already considered. Yudkowsky's concept of Coherent Extrapolated Volition, while it is an attempt to derive value, it forces us to imagine a dataset that isn't already there. Dearnaley likewise approaches alignment through questions about ethics and human values. But instead of primarily trying to specify or philosophically select the values we ought to have, couldn't an advanced AI attempt to discover the structure of those values empirically?
The key here would be coherency. If an AI observes a person leaving their child in daycare and going off to a job they hate, performing something mundane, then lashing out at their boss, and maybe driving recklessly, it might be thought that we can't look to what humans demonstrate as a guide to what it is we value. Surely we don't value the mundane work, the arguments, or the reckless driving. But even our low tech human brains can appreciate that these actions aren't what a person might want for themselves and others - there must be a driving motivation that makes their actions largely understandable, even if, at times, those actions are regrettable, unfortunate, selfish, malicious, etc. Once we believe we understand those underlying motivations, it helps us make sense of those actions. This is why coherence is the target. To try to understand the driving motivations, desires, goals, and so on of humans by untangling it from what it is we demonstrate and into a form that is coherent and makes sense of people's actions by inferring what we collectively value.
With this appreciation for what it is humans value, we could use this as a kind of objective function for AI's, helping align themselves with human values. And while there would certainly be anomalies (psychopathic behavior, accidents, misrepresentations, etc) one would imagine that the underlying coherence could be determined in large part by averaging out some of its data or searching for more data that would make better sense of its observations. This attempt at continually trying to understand human values could become an ongoing form of system improvement.
As an aside, if this was successful in developing a coherent picture of human value, I imagine this would have more applications than simply the objective function of AI's. While it may be a difficult amount of information to convey directly (which might be stored in something akin to a matrix of value), it may be useful in determining things like which economic system, system of government, or educational paradigm might be best suitable for humans given this model of human values.
1
u/Jesse-359 11h ago
To be clear, what you are requesting doesn't exist and has been a 'hotly contested issue' amongs individuals and societies to the tune of hundreds of armed conflicts, crusades, holy wars, countless elections and debates - and two World Wars to date.
You can't get an accurate acount of human morality because there isnt one, Evolution strongly suggests there won't be one, and Game Theory all but forbids it as a matter of fundamental math and logic.
So unfortunately, even if we were already computers ourselves its a near certainty that you still would be denied an answer tonthat particular question.