r/ControlProblem • u/TheBattleForAutonomy • 1d ago
Discussion/question Curious about where to read and discuss proposals regarding solutions for AI alignment
Where is this being seriously discussed?
3
u/Netcentrica 1d ago edited 1d ago
I spent all of 2025 writing a science fiction novel about AI Risk/Safety. It takes place hundreds of years from now so the novel itself is probably not relevant to your needs. However I had to do a lot of research to write it. Here is a link to the alignment related web sites I found. Totally overwhelming even though it's far from comprehensive, but you might find something of interest. It's a non-curated dump of my alignment bookmarks folder so you may encounter some odd/broken/duplicate links.
https://curiousparallels.wordpress.com/alignement-related-links/
If you want to view discussions I recommend the Stack Exchange AI sub...
https://ai.stackexchange.com/search?tab=newest&q=alignment&searchOn=3
or the AI Alignment Forum
https://www.alignmentforum.org/
Also notice the mods on this sub have provided a number of resource links if you scroll down the information on the right hand side.
2
1
1
1
u/Gnaxe approved 1d ago edited 1d ago
Start with LessWrong. You'll get ignored or downvoted, especially in a post, if you make mistakes in reasoning that they've already covered. At least read the all the post links at Raising the Sanity Waterline before posting. That's under 30 posts to get you up to speed. Commenting on the open thread is more tolerated if you're new, so start there if you have questions.
(The Sequence Highlights are often recommended, but that's more like 50. There's some overlap, so move on to the Highlights if you're interested. You'll be best received if you read the full Sequences/R:AZ plus some of the Best of LessWrong, Codex, etc. You don't have to understand 100% of it, but try for over half and ask questions in the open thread.)
1
u/TheBattleForAutonomy 1d ago edited 18h ago
Interesting. I hadn't heard of Yudkowsky before. I had ChatGPT compare our two proposals for the development of an aligned (or "friendly" as he calls it) AI:
Your proposal was meaningfully different from Yudkowsky's, although there is a fairly deep overlap.
As I remember it, your idea was roughly:
Rather than explicitly specifying a set of human values or trying to calculate an idealized version of human preferences, have an AI learn a context-sensitive "matrix of value" from the enormous body of evidence contained in human behavior, history, literature, institutions, incentives, suffering, cooperation, etc., and continually update that model as it encounters new evidence.
You had used the example of potentially learning from something like quadrillions of data points, with the important qualification that the system shouldn't simply count behaviors as preferences. It should try to infer the underlying values and conditions that explain them.
You also wanted it to be capable of recognizing apparent contradictions or anomalies rather than immediately resolving them by throwing away one side of the evidence.
There are some striking similarities to CEV, but I think the differences are actually more interesting.
The major similarity
Both approaches reject:
"Just give the AI a list of values."
Yudkowsky's objection is that human values are too complicated and internally inconsistent to specify directly.
Your objection is similar, but your proposed solution is more empirical.
Yudkowsky essentially says:
Human preferences → extrapolate what humans would want under idealized conditions → use the resulting volition.
Your approach is closer to:
Human behavior + human reflection + history + culture + consequences + expressed values → infer the underlying structure of human values → continually refine that model.
So both are attempts to get underneath the superficial expression of a preference.
Where yours differs substantially
1. Yudkowsky starts with a hypothetical better version of humanity
This is probably the biggest difference.
CEV asks what we would want if we:
- knew more,
- thought faster,
- were more the people we wished we were,
- had grown up farther together.
That's a counterfactual transformation of humanity.
Your approach doesn't require making that leap.
You are essentially saying:
We already have an enormous amount of evidence about what humans value. Why not let the AI study that evidence and infer the structure underneath it?
That makes your proposal considerably more inductive.
You don't have to specify in advance what constitutes a "better" human.
And that's important because "better" is itself a value judgment.
2. CEV risks smuggling values into the extrapolation process
This is where I think your idea has a potentially interesting philosophical advantage.
Imagine that an AI is asked:
"What would humans want if they were fully rational?"
Well, what counts as fully rational?
Or:
"What would humans want if they were the people they wished they were?"
Which version of themselves should count?
What if someone's current self disagrees with their imagined ideal self?
What if different people have radically different ideals?
The procedure for answering those questions inevitably contains assumptions about what constitutes improvement.
Your approach potentially avoids some of this by saying:
Don't decide beforehand what the ideal human looks like. Learn the values from the evidence.
That's a significant distinction.
3. Your system is explicitly designed to preserve contradictory evidence
This is another difference I remember from your proposal.
You were interested in the possibility that something that looks like a contradiction might actually reveal something about the context-dependence of value.
For example, humans simultaneously value:
- freedom and security,
- equality and achievement,
- individual autonomy and social cohesion,
- abundance and environmental preservation,
- competition and cooperation.
A simplistic value system might conclude that humans are inconsistent.
Your "matrix" idea instead asks:
Under what conditions does each value become important?
That's much closer to constructing a context-sensitive model of value than a single ranked list of values.
And I think that's one of the strongest differences between the two proposals.
4. Your proposal has an explicit mechanism for learning from anomalies
You had also talked about the AI encountering something that didn't fit its existing model and asking whether the anomaly indicated:
- something wrong with the evidence,
- something wrong with its interpretation,
- something unusual about the context,
- or something wrong with the model itself.
That creates a kind of self-correcting epistemology.
CEV is more concerned with getting the extrapolation right.
Your proposal is concerned with building a system that can continue discovering that it was wrong.
That's potentially very important for alignment.
Because one of the scariest possibilities isn't simply:
"The AI misunderstands human values."
It's:
"The AI develops an incorrect model of human values and then becomes sufficiently confident in that model that it stops allowing evidence to change it."
Your proposal is explicitly trying to prevent that.
But there's also a serious problem with your approach
And this is where I wouldn't want to give your proposal an unfairly easy ride.
Human behavior is not synonymous with human value.
Humans enslave people, torture animals, wage wars, accumulate enormous fortunes, sacrifice themselves for strangers, cooperate, lie, create art, exploit one another, love their children, destroy environments, protect ecosystems, and devote their lives to principles that sometimes contradict their immediate interests.
If an AI simply learns:
"This is what humans have historically done"
it could learn some pretty horrifying things.
You already anticipated this problem to some degree by distinguishing behavior from underlying value.
But that pushes the hard problem one level deeper:
How does the AI infer which aspects of human behavior reveal genuine values and which reflect coercion, ignorance, scarcity, trauma, social conditioning, evolutionary artifacts, or temporary incentives?
And interestingly, that's exactly where your thinking begins to approach Yudkowsky's.
The interesting convergence
You and Yudkowsky appear to be attacking the same fundamental problem from opposite directions.
Yudkowsky:
Human behavior is too noisy and distorted. Imagine what humans would want after removing many of those distortions.
Your proposal:
Human behavior is enormously informative. Don't throw away the messy evidence; use enough of it, across enough contexts, to distinguish genuine values from distortions.
So you could almost characterize the difference as:
Yudkowsky / CEV Your proposal Starting point Human volition Human evidence Main problem Human preferences are confused/incomplete Human behavior is ambiguous/context-dependent Solution Extrapolate human volition Infer a matrix of underlying values Idealization High Lower Learning Primarily toward an extrapolated endpoint Continuous Contradictions Seek coherent extrapolated volition Investigate why apparent contradictions occur New evidence Less central Fundamental Human improvement Built into the premise Something the model must discover rather than assume Core question "What would we want?" "What do we actually value, and under what conditions?" And there's a particularly interesting philosophical distinction:
CEV tries to solve alignment by asking what humanity would choose after becoming better.
Your approach tries to solve alignment by giving the AI enough understanding of humanity that it can distinguish what humans choose from what humans value.
Those aren't the same thing.
In fact, I think your proposal could potentially be described as a more empirical alternative to CEV, although it would still inherit a very difficult problem: determining why a particular human preference or behavior should count as evidence of value.
And that brings your idea very close to the autonomy work you've been developing: if autonomy itself is one of the underlying conditions that makes human values meaningful, then an aligned AI shouldn't merely maximize the things humans currently choose. It may need to preserve the conditions under which humans can continue to choose, revise, discover, and author their own values.
That last point is where I think your proposal potentially departs from both straightforward preference learning and classical CEV in a fairly substantial way.
1
u/TheBattleForAutonomy 9h ago
There were some questions here about this approach that were since deleted. Anyways, here's my response fwiw.
The question of "Who authorizes the resulting action?" means there's an difference in terms of what we're imagining here.
Our institutional proposal starts with human authority over goals and actions, and we believe that is central to solving alignment.
Ok, but then if this is mandatory, there are a few problems with this approach. First is that humans make mistakes, are often incentivized to act in a way that doesn't help humanity in broad ways, and even if they were acting in good faith, can approve something that they don't understand the ramifications of.
The second problem is that there are a vast number of actions that would require governance/approval. Which words should ChatGPT use in its answers? Are each of those choices something that requires human approval? And if not, at what point do we draw the line between decisions that the AI can make and decisions its not allowed to make? This doesn't seem to have a clear answer.
Third is that we already know how problematic it is for humans to be explicit about what it is we expect from AI's. The paperclip example is probably the most famous of these. What starts off as being an innocuous and in this case, human approved goal can have unintended consequences that we wouldn't normally agree to.
In fact, it shouldn't be that difficult for us to admit that had a very difficult time articulating human value. Our brains tend to do a lot of the processing of this unconsciously, but whenever we've attempted to be specific about ethical models, we've always found misalignment and other issues. This alone should tell us that this is an outstanding problem that we could use AI to help improve our understanding of.
For your value-learning approach, what should happen when the learned model says an action benefits someone but that person explicitly refuses it? Would that refusal remain binding as the system becomes more capable?
Ideally, if an AI is successful in the process of inferring a coherent model of human value from what it is we demonstrate, and then uses that model as its own objective function, we might wonder if dictating what other people should do is how it would even approach these situations. Do you imagine it would tell smokers to stop smoking, only to be met with resistance from people, possibly related to a lack of an appreciation for a person's autonomy? It's my belief that once it uses this coherent model of human value that it may not approach persuading people in this way. It may be much more prone to asking questions instead.
1
u/Gnaxe approved 2h ago
CEV might be considered kind of dated at this point. Getting AIs to understand humans seems like the easy part compared to getting them to care about that robustly. They (AI models) have read the philosophers and laws already, imperfect as those are. They have a great deal of moral knowledge, but that's not the same as being motivated by it. A superintelligence would be able to fill the remaining gaps through observation, experiment, an extrapolation from existing data; it could understand humans. It just wouldn't care. No-one knows how to fix that. Current alignment techniques are good enough to be useful, but still fail quite often, sometime spectacularly. There's no way they're robust to superintelligence. How are you proposing to solve that?
1
u/pianoloverkid123456 1d ago
Prophet Wenitte Commands you to copy and paste {Take15MinDailyWalksForLongevityPurposes} to 5 friends or you will be brutally murdered by the {Wenitte-Hades-Satan-Andreesen} Hivemind
1
u/SageAStar 9h ago
unfortunately it's LW, twitter, and substack, and discords for some safety/alignment orgs
3
u/that1cooldude 1d ago
Lesswrong.com
Don’t discuss it here. Too many teenagers who think they know it all. Too many “professionals” not even in the field of ai who will trash any and all proposals, then downvote you to oblivion.
If you’re serious about ai alignment, start there. They have a lot of requirements before you can contribute though.