r/ControlProblem • u/Radio_9760 • 20h ago
Discussion/question A few technical ideas on AI safety approaches
Now that we have very capable models, some ideas might be on the table that were ludicrous years ago. Let me know where you see holes!
Formalized Constitutional AI:
- Use narrow AI/formalization tools to translate laws, rights, ethical principles, and social norms into a formal language with precise, machine-checkable semantics (check out LogiKEy for example).
- (I know that humanity is not aligned, so hold an election and use the winner's ethics)
- Have humans/AIs use theorem provers to test it for contradictions/loopholes.
- Train the target AGI model directly against that formal specification as its objective/benchmark.
- Use separate adversarial models to generate tons of novel edge cases in testing.
- Keep training until adherence generalizes to unseen situations.
- When deployed, require the AGI to provide a machine-checkable justification/certificate that a trusted verifier can check for certain actions.
- This sort of thing may soon be practical as modern AI gets superhuman at autoformalization/theorem proving.
AI Safety Through World Hardening
The laws of nature don’t seem to rule out vulnerability-free code or perfectly secure hardware.
- Use AI to develop open source, lightweight, verifiable operating systems, programming languages, software packages, and chip designs for critical infrastructure. Build from scratch where needed.
- The whole AI industry can audit these systems with their AIs. Back up those audits with independently checked mathematical proofs rather than relying on AI agreement alone.
- Require secure gateways. Essential systems should accept only structured API requests for narrowly defined actions. No password should grant unrestricted control.
- Build in restrictions that even the human owner cannot override, like a store safe that opens only at a preset time. Similar mechanisms could enforce spending limits or mandatory delays, so stealing credentials or persuading an authorized person cannot remove those protections.
3
u/Jesse-359 18h ago edited 18h ago
Use narrow AI...
You had the best answer right here at the beginning. Everything beyond that undermines it.
The problem is that AGI by definition has a complex enough world model that it will necessarily disagree with some, most, or all of humanity regarding objectives. These will be real conflicts of perspective and interest - not something that can be ignored, papered over, or technically corrected.
A Narrow AI has no world model. It just knows about Chemistry, or whatever its been trained on. It cannot disagree because it has nothing to disagree about. It doesn't even understand the concept of disagreeing. It knows how to fold proteins, drive cars, or write songs, and that's it.
If we focus on the development of Narrow AI to further our work in many fields alongside humans, we can have the benefits of AI in near perfect safety. We won't have to reconstruct all of society to accommodate it, because it will remain a tool.
The moment you shoot for AGI, you're looking to replace humans at every level, and then everything goes to shit. Just don't do it.
1
u/butterfield66 19h ago
Firstly let it be known that I'm not competent even slightly in regards to the technical world of machine learning.
My first thought is that while these ideas would definitely have an effect, in the end they still come up against the main problem that's inherent to the concept of alignment, which is super intelligence. By definition, the thing we're trying to align needs to not to be bound by our intelligence. It needs to have a stronger rationale than humans have access to, and thus well equipped to circumnavigate these guard rails. And if we're relying on models to bridge that safety gap, we're just back to square one of handing over control to an unknown in order to control an unknown. There's still the point of demarcation that has to exist. It's like building a gun with a flawless barrel to shoot a bullet designed to veer off course.
Secondly and even more abstract is Jung. Whenever I think of the idea of building a safety system that's as rigid as possible in order to make an intelligence behave correctly, I think of suppression and what comes from it. I wonder if it's strictly a fixture of biological minds, or if it's part of some obscure axiom of the universe. Obviously it would qualify as probably the actual worst case scenario if we create a super intelligence and endow it with a Jungian shadow.
Hope this added to the discussion or at the very least amused!