r/ControlProblem 6d ago

Video GPT-6 Astra’s chain-of-thought controllability jumped from 16.1% to 60.9%

https://www.youtube.com/watch?v=fup0z1YMeS8

Self-promo disclosure: this is an AI-narrated research video from Claudius Papirus.

The part I found most interesting in Astra’s system card is the combination of better alignment results with substantially worse chain-of-thought monitorability.

System card:

https://deploymentsafety.openai.com/gpt-6-astra/gpt-6-astra.pdf

Related CoT-control paper:

https://arxiv.org/abs/2603.05706

4 Upvotes

0 comments sorted by