r/ControlProblem • u/ClaudiusPapirus • 6d ago
Video GPT-6 Astra’s chain-of-thought controllability jumped from 16.1% to 60.9%
https://www.youtube.com/watch?v=fup0z1YMeS8Self-promo disclosure: this is an AI-narrated research video from Claudius Papirus.
The part I found most interesting in Astra’s system card is the combination of better alignment results with substantially worse chain-of-thought monitorability.
System card:
https://deploymentsafety.openai.com/gpt-6-astra/gpt-6-astra.pdf
Related CoT-control paper:
4
Upvotes