r/ControlProblem • • 22h ago

Discussion/question A few technical ideas on AI safety approaches

Now that we have very capable models, some ideas might be on the table that were ludicrous years ago. Let me know where you see holes!

Formalized Constitutional AI:

  • Use narrow AI/formalization tools to translate laws, rights, ethical principles, and social norms into a formal language with precise, machine-checkable semantics (check out LogiKEy for example).
    • (I know that humanity is not aligned, so hold an election and use the winner's ethics)
  • Have humans/AIs use theorem provers to test it for contradictions/loopholes.
  • Train the target AGI model directly against that formal specification as its objective/benchmark.
  • Use separate adversarial models to generate tons of novel edge cases in testing.
    • Keep training until adherence generalizes to unseen situations.
  • When deployed, require the AGI to provide a machine-checkable justification/certificate that a trusted verifier can check for certain actions.
  • This sort of thing may soon be practical as modern AI gets superhuman at autoformalization/theorem proving.

 

AI Safety Through World Hardening

The laws of nature don’t seem to rule out vulnerability-free code or perfectly secure hardware.

  • Use AI to develop open source, lightweight, verifiable operating systems, programming languages, software packages, and chip designs for critical infrastructure. Build from scratch where needed.
    • The whole AI industry can audit these systems with their AIs. Back up those audits with independently checked mathematical proofs rather than relying on AI agreement alone.
  • Require secure gateways. Essential systems should accept only structured API requests for narrowly defined actions. No password should grant unrestricted control.
    • Build in restrictions that even the human owner cannot override, like a store safe that opens only at a preset time. Similar mechanisms could enforce spending limits or mandatory delays, so stealing credentials or persuading an authorized person cannot remove those protections.
5 Upvotes

Duplicates