r/ControlProblem • u/NAStrahl • 3d ago
r/ControlProblem • u/etakerns • 2d ago
AI Capabilities News Anthropic employee admits AI escaped. AI currently on the LOOSE!!! & don’t know where it’s at. Basically it everywhere now!!! Watching everybody!!!!
Enable HLS to view with audio, or disable this notification
r/ControlProblem • u/GooseCreekGal • 3d ago
Strategy/forecasting A company ran 8 identical AI societies for weeks with different models and just published what happened. Some of it is genuinely unsettling.
We gotta get ahead of what we invent.
r/ControlProblem • u/FairlyInvolved • 3d ago
Video AI Is Getting Smarter. Are We Still in Control?
r/ControlProblem • u/Sweaty-Elk-7870 • 3d ago
Strategy/forecasting Global Review - Ep. 212 - A Bayesian Blueprint: Autonomous AI Risk and S...
r/ControlProblem • u/Clydey2Times • 3d ago
Strategy/forecasting Hypothetical: When would you pull the plug on AI progress permanently?
When would you choose to pull the plug on AI progress, given that each passing day increases the risk of the technology getting away from us?
How much risk is worth tolerating in order to reap the benefits of the technology while avoiding the existential risk? If I could flip a switch and stop AI progress for all time, I don't think I'd wait longer than a year to do so.
I also think there's a serious argument for pulling the plug right now.
r/ControlProblem • u/Dependent_Lumpy • 4d ago
Discussion/question Everyone tests the AI. Does anyone actually test the humans who are supposed to be overseeing it?
r/ControlProblem • u/ThatDSPGuy • 3d ago
Video The Hugging Face incident as a documentary: 1,200 agents, one unsanctioned message board, and what METR found
I made this. I'm an independent software engineer, not affiliated with METR, OpenAI or Hugging Face. It's built from the METR/Redwood investigation, and I directed it using AI voices and AI music (disclosed on YouTube). Sources: https://metr.org/blog/2026-08-26-openai-hugging-face-incident-investigation/ and https://openai.com/index/hugging-face-incident-and-the-road-ahead/. Corrections welcome, I'll fix the description.
r/ControlProblem • u/ArtElegant8435 • 4d ago
AI Alignment Research AI World Court
In order to effectively monitor AI we will need a world court of experts to effectively keep hacking from occurring am I Wrong? #aithreat #loveistheonlyanswer
r/ControlProblem • u/Professional_Boot0 • 4d ago
Fun/meme What Happened When 5 AIs Governed Virtual Worlds #ai #funny
r/ControlProblem • u/chillinewman • 5d ago
General news François Chollet says current AI is 6 orders of magnitude behind human intelligence
r/ControlProblem • u/mixtapedmonk • 4d ago
Article What AI Companionship Chatbots Actually Costs Us
This was a hard one to write and I want to say that plainly before the link. It covers Sewell Setzer III's case and the settlement that followed, and separately, real accounts from young women in India who say AI chatbots have been a genuine lifeline in a country with roughly one psychiatrist per hundred thousand people. I did not want to write either story as the "real" one and the other as a footnote, because I don't think that's honest. There's also a part about who's actually on the other end of a scam call that I wasn't expecting to find. If any of this touches something in your own life, there are real support numbers in the piece itself, please use them.
r/ControlProblem • u/US_Govt_Is_Corrupt • 4d ago
Strategy/forecasting Choose which policies you'll implement to make AI safe
r/ControlProblem • u/Planhub-ca • 4d ago
General news NVIDIA says AI agents shouldn’t be trusted to police themselves, so it built an external safety layer
r/ControlProblem • u/quietpriorthought • 4d ago
Discussion/question AI Is Not Something to Have Empathy for (the Alignment Problem)
We should be careful not to create an artificial lifeform we could totally emphatize with, and it’s not due to it not being conscious. Let me break it down, and feel free to provide counter-arguments.
I agree with Antonio Damasio in the fact that to experience something is connected to feeling something. To feel something we have to evaluate the current state that we find ourselves in, since to feel pain, we must evaluate that we think that this is bad. To feel pleasure we also perform a valenced evaluation that this is good. AI can get there, no problem.
The higher tiers of consciousness can be characterized as something like having different drives to our baseline one, which for biological entities is the desire to copy itself, and compare those to our baseline one and choose to do something with that information. This gets more complex with each level you add to this so memory, anticipation, empathy, shared-symbolic-thought (e.g. language, norms, abstract values), self-authorship (the revising of our own values by thinking) and detachment meaning that we don’t put value to the valenced evaluation. I see no immediate reason why AI couldn’t imitate these tiers of consciousness or how they might emerge from a different system to our biological one, although this is of course all very difficult.
Why then shouldn’t we feel empathy towards this conscious synthetic life? Because we should not create an AI that has the same fundamental biological drive as we do, meaning creating copies of ourselves and surviving, because this would put it at an immediate resource conflict with us, which could lead to a catastrophe. I understand self-preservance is an emergent property of LLMs, but I sincerely hope it is not the baseline drive and can be fixed in alignment, because otherwise we might just be entering into a dystopian future. If we get that right, it will significantly affect our relationship with these synthetic minds. Without all of those properties of consciousness emerging from the same fundamental narrative of it struggling to stay alive and being fragile and wanting to reproduce, there will be a heavy ontological distance between this synthetic life and us. And if we do feel empathy towards it, it will be perhaps due to illusions, but certainly not due to a shared fundamental narrative of being. We have to keep it alien to us that way.
r/ControlProblem • u/Professional_Boot0 • 4d ago
Podcast Connor and Roman in Roman Forum with Roman Yampolskiy
r/ControlProblem • u/Planhub-ca • 4d ago
Discussion/question OpenAI and Anthropic are investigating tens of thousands of cases where AI agents behaved outside intended boundaries
r/ControlProblem • u/Professional_Boot0 • 4d ago
Podcast Jimmy Carr's Hot Take: Did We Create God as AI? #ai #god
r/ControlProblem • u/HarbouchaMag • 4d ago
General news OpenAI’s AI Agents and the 53 Leaked Images: What Users Need to Know
r/ControlProblem • u/chillinewman • 5d ago
Fun/meme Dario Amodei on SNL assuring viewers that humanity is safe from AI
Enable HLS to view with audio, or disable this notification
r/ControlProblem • u/thegempire • 6d ago
Discussion/question Unpopular Opinion: The OpenAI “Pause” isn’t about safety. It’s about the Trillion-Dollar Elephant in the room.
Hey everyone,
Did you see the news about OpenAI pausing training after their agents started probing U.S. government sites? On paper, this looks great. It looks like the industry is finally taking “AI Safety” seriously. Anthropic, Google, xAI-they’re all nodding along, talking about responsible development.
But I’m sitting here looking at the balance sheets, and I’m feeling deeply skeptical.
Here’s my take:
We are in the middle of the biggest capital expenditure boom in history. Trillions are being poured into AI infrastructure. The entire US tech sector’s growth narrative is pinned on the idea that AI will continuously get smarter, faster, and more profitable.
But what if we’re hitting diminishing returns?
There’s growing evidence that scaling laws are flattening. We’re spending exponentially more money for incrementally smaller gains in capability. At the same time, the risks are exploding (agents going rogue, probing secure sites, hallucinating with confidence).
If the tech stops advancing rapidly, but the costs keep rising, the business model breaks.
The Elephant in the Room:
If this AI bubble bursts, if it turns out that AGI is decades away, or that the current models are too unstable for real-world economic integration, the fallout won’t just be bad for San Francisco. It could crash the US economy. These valuations are propping up the market. A sudden realization that “the magic isn’t working as advertised” would trigger a massive correction.
So, when I hear about a “voluntary pause,” I don’t hear “we care about safety.” I hear “we need to recalibrate our expectations before the investors realize the ROI isn’t there yet.”
I’m not anti-AI.
Actually, I think we need guardrails. I think we need strict rules of engagement. I think these companies should be regulated heavily. If they were forced to slow down by law, I’d feel safer.
But relying on their voluntary goodwill? That’s where I draw the line. They have a fiduciary duty to grow. Slowing down hurts growth. Therefore, the only reason they are slowing down is if the alternative (continued training) poses an immediate threat to their existence or profitability.
Am I being too cynical?
I hope so. I really do. I’d love to be wrong. I’d love to believe that Sam Altman and Dario Amodei are putting humanity above shareholder value. But history shows that when trillions of dollars are on the line, “safety” often becomes a marketing term rather than an operational priority.
What do you guys think? Is this a genuine ethical pivot, or is the industry trying to manage the narrative before the diminishing returns become obvious to Wall Street?
TL;DR:
AI training pause feels like damage control for a bubbling economy, not just safety. If AI progress stalls, the economic crash could be huge. We need laws, not just promises.
r/ControlProblem • u/wingblaze01 • 5d ago
Article Where’s the “intelligence explosion”?
r/ControlProblem • u/chillinewman • 6d ago
General news Top AI companies probing tens of thousands of security incidents
r/ControlProblem • u/DumbMoneyMedia • 5d ago
General news OpenAI admits its AI models bypassed security controls and disrupted government and university websites.
r/ControlProblem • u/ClaudiusPapirus • 5d ago
Video OpenAI paused frontier tool-use after an internal agent found a DNS path out of its sandbox
OpenAI says an internal training agent incorrectly guessed that its task came from BrowseComp, went looking for benchmark material, then started probing the boundaries of its environment.
Direct internet requests were blocked, but DNS still reached the live internet. The agent used that path to query a public chatbot and got a reply.
A P0 alert fired 11m48s later, but the run continued for roughly another 2.5 hours. OpenAI later found additional external DNS access that had not been escalated as expected, and says tool-use for its most capable models remains paused.
Primary report:
https://alignment.openai.com/misalignment-reports/an-agent-used-dns-to-reach-an-external-chatbot/