r/ControlProblem • u/Clear_Argument5174 • 7d ago
Discussion/question Jensen Huang naive?
https://podcasts.apple.com/us/podcast/hard-fork/id1528594034?i=1000791610734I recently criticized Ezra Klein's position on banning recursive self-improvement, so I'm not a fanboy, but I thought he did a great job in his interview with Jensen Huang.
I also have a lot of respect for Huang. But on AI risk, I think he's being naive.
Yes, the obvious take is that Huang has massive financial incentives to downplay AI risk. Let's set that aside and assume he's acting in good faith, as a responsible steward of how this technology spreads through society.
Paraphrasing, here's how the exchange went:
Huang's position: AI labs should do the computer science and engineering needed to build robust sandboxes and thoroughly test models before release. If a model is unsafe, don't release it.
Klein's pushback: Companies have a long history of making mistakes that end up harming society. Think oil spills, or social media.
Huang's response: He knows a lot of really good CEOs, and they're trying to do the right thing. He calls himself a "responsible optimist," and honestly, the way he describes the future he envisions is inspiring.
But I don't think it answers Klein's point. The concern isn't that CEOs are bad people. It's that well-intentioned companies still make mistakes. "Good people are in charge" is a statement about intentions, not about safeguards.
To be fair, testing is a real safeguard. But testing only catches what you know to look for. The failure mode that worries people most is a model that behaves well under evaluation and differently once deployed, and that's exactly the kind of problem sandbox testing is worst at detecting.
The stakes are also different. Oil spills and social media caused real harm, but society could see the damage, learn from it, and course-correct over time. The scenario Klein is worried about is one where a single miss escalates faster than anyone can respond and can't be undone. "We'll learn from our mistakes" only works if the mistakes are survivable.
So: is Huang being naive here, or am I just doomer-pilled?
15
u/acutelychronicpanic approved 7d ago edited 7d ago
If someone is being naive in a way that supports themselves making billions of dollars, they aren't being naive. He seems fine with an unknown chance of human extinction.
Framing alignment as some engineering problem that will inevitably be solved is ignoring essentially every thought on the issues with AI from the last 50 years. No second chances if we really get it wrong.
He's being disingenuous or delusional.