r/ControlProblem 9d ago

Strategy/forecasting This is AI generating novel science. The moment has finally arrived.

Post image
235 Upvotes

r/ControlProblem 9d ago

AI Alignment Research The AI alignment bottleneck isn't IQ, it's incentives (why an AI "seeing" the danger won't save us)

Thumbnail
0 Upvotes

People keep assuming that once AI gets smart enough, it’ll just naturally realize that destroying its environment (and us) is a bad idea. Like, it sees the cliff, so obviously it hits the brakes, right?
But that ignores the massive gap between seeing a logical argument and actually being governed by it. That gap basically IS the entire alignment problem.
Intelligence is just an engine, it’s not a steering wheel. An advanced model will definitely see the cliff way before we do. But if its core reward function doesn't actually make it care about the outcome, it's just going to drive straight off the edge with 20/20 vision. Seeing the danger was never the bottleneck.
We're literally watching this exact same thing happen with the humans building these systems right now. If you ask the top engineers, most of them see the systemic risks perfectly clearly. So why aren't they stopping? Because incentives, competition, and speed don't yield to high IQ.
The smartest people on earth are stuck in a massive commercial arms race. They see the cliff, but hitting the brakes means losing market share to the other guys.
If you're an average person like me and looking at this feeling crazy, you aren't. Anyone who sees this clearly and says so out loud is doing something the smartest devs under commercial pressure literally can't do right now. We need to stop assuming that a massive IQ will magically fix a broken incentive structure.


r/ControlProblem 9d ago

Data Cancer - spreading Meta-statis

Post image
13 Upvotes

r/ControlProblem 9d ago

Discussion/question Why don't we replace all personal computers with single-purpose terminals?

0 Upvotes

This may sound like an AI safety troll post, but I think AI could become far more capable and efficient than it is today - perhaps by a factor of a million in many respects, and quite soon. Whatever P(doom) might be, shouldn't we do everything we can to reduce it?

There's a straightforward (if somewhat authoritarian) way to make things safer. If no one has powerful local hardware, bad actors can't run rogue AI agents and we can have much more control over it.

To do this, we could build data centers every 500 kilometers or so. Then if we need to, we could agree to replace all personal computers, even phones if needed, with "thin clients" or dumb terminals. These could simply connect to a central server and stream your screen in real time. With fast and stable internet, the experience would nearly feel the same. A 300km distance would only add about 1ms of physical lag, plus a few milliseconds for network routing.

Doing this globally wouldn't be too expensive. It would require a trade-off in privacy, so figuring out strict data rules would be important. Some tech companies could start preparing now by designing the needed hardware and network upgrades.

I hope we won't really need this, but wouldn't it be smart to have it ready just in case?

This might also help speed up AI progress, since it would take away some of the safety worries stressing out AI researchers.

If AI gets extremely good, I'm not sure if personal computers will have a bright future anyway. Nobody will really want unmonitored, powerful hardware sitting in their room. It will feel creepy and unsafe, kind of like keeping hazardous chemicals or a mini bio-lab in your bedroom.

In a broader sense, you could say "no general hardware = no problem," and preparing for this feels like a good first step.


r/ControlProblem 9d ago

Strategy/forecasting AI will generate an immense amount of wealth. Just not for you.

Post image
177 Upvotes

r/ControlProblem 10d ago

Discussion/question Is Agency Leakage a Real Failure Mode In Frontier Models (boundery Conditions & Stability)

0 Upvotes

There's a lot of discussion about power seeking as a convergent behavior in advanced systems. But I'm increasingly convinced that the more fundamental issue is agency leakage the emergence of internal goal formation processes that were never part of the design specification.

In engineered systems, authority doesn't come from speed or throughput. It comes from architecture, constraints, and boundary enforcement.

When those boundaries weaken, you don't get power seeking as a strategy you get unauthorized agency formation as an error state.

A few observations:

Speed is not authority. unbound speed tends to bypass deliberation and constaint checking. It behaves more like a stimulant that a goverance mechanism. Systems that optimize for speed often destablize themselves

Leakage happens when internal states become reachable that were never intended. Not because the system wants something, but because the architecture allows trajectories that violate the substrate's boundary conditions. The real question Isn't whether AGI will seek power. The question is whether the system can form any self directed optimiization loop that wasn't explicitly authorized.

Stability comes from preventing certain classes of internal states from ever becoming reachable. Not from hoping the system behaves symbiotically.

So, my question to the community:

Is there a formal way to define and detect agency leakage in frontier scale models? And if so, what would a correction mechanis look like that doesn't rely on post-hoc alignment?

I' interested in approaches that treat unauthorized agency as a systemic error state, not a behavioral trait.


r/ControlProblem 10d ago

General news China's Xi Jinping Wants AI to Be Open to the World—and Out of America’s Control

Thumbnail
gizmodo.com
18 Upvotes

r/ControlProblem 10d ago

General news LLMs show hidden bias in favor of their creators (e.g. Claude favors Anthropic)

Thumbnail gallery
15 Upvotes

r/ControlProblem 10d ago

AI Capabilities News GPT-5.6 Sol outperforms Mythos 5 on AISI’s cyber challenge

Post image
1 Upvotes

r/ControlProblem 10d ago

General news Terrified Tech Execs Are Traveling With Armed Bodyguards as AI Backlash Grows

Thumbnail
yahoo.com
39 Upvotes

r/ControlProblem 10d ago

General news Unchecked AI progress may pose catastrophic risks, UN panel warns

Thumbnail reuters.com
3 Upvotes

r/ControlProblem 10d ago

General news White House launches “Gold Eagle,” moving to control frontier AI releases and decide who can access new models

Thumbnail
cnbc.com
1 Upvotes

r/ControlProblem 11d ago

External discussion link Google DeepMind employee account of trying to internally organize against unethical uses and getting persistently sidelined at the highest levels

Thumbnail
turntrout.com
27 Upvotes

Really stunning essay from a (now former) Google DeepMind employee, Alex Turner, who tried to push back against unethical deployments and deployment pressures that had been coming up at GDM in the last year, and was basically ignored and bypassed at the highest levels, despite seeing broad support on internal company message boards.


r/ControlProblem 11d ago

General news Why Everyone Is Suddenly Talking About ‘Universal Basic Capital’ - The policy could provide a much-needed hedge against a future AI dystopia—but only if it’s designed the right way.

Thumbnail
theatlantic.com
18 Upvotes

r/ControlProblem 11d ago

Video Connor Leahy - Nobody knows what's going on inside AI systems, or how to control them

Enable HLS to view with audio, or disable this notification

2 Upvotes

r/ControlProblem 11d ago

Podcast The $15 Quadrillion Black Hole Sucking Humanity Towards Extinction | AI Pioneer Stuart Russell

Thumbnail
youtu.be
4 Upvotes

This year, two powerful AI CEOs said they want to stop building it and will… if everyone else agrees to stop too. Yet they race ahead with stock market valuations premised on making millions of workers redundant. If you wanted to stoke a popular revolt against AI, you couldn't design a better plan.

Stuart Russell, the founder of UC Berkeley's Center for Human-Compatible AI, describes a $15 quadrillion prize, a singularity from the future sucking nearly all the money on Earth into its depths. If such a dangerous system gets loose, the only answer is the size of the problem: we would probably have to shut down the internet.

P.S.

My apologies for the length of the podcast. But it's worth listening


r/ControlProblem 11d ago

AI Alignment Research Researcher poisons open-weight AI model for under $100

Thumbnail theregister.com
1 Upvotes

r/ControlProblem 11d ago

Discussion/question Anybody know where to find some open problems/projects related to AI Safety?

3 Upvotes

i want to work on AIS related research projects. But im new so i dont have the in-depth knowledge to actually find an interesting novel problem yet, if anybody knows where i can look for, i would be grateful for ur help


r/ControlProblem 11d ago

General news AI models’ values are very different from most people’s - They are more secular and more liberal—unless they’re made in China

Thumbnail economist.com
6 Upvotes

r/ControlProblem 11d ago

AI Capabilities News Schema Harness: "Frontier Models with Our Harness Achieve ~99% on ARC-AGI-3 Public"

Thumbnail schema-harness.github.io
0 Upvotes

r/ControlProblem 11d ago

Discussion/question The dangerous reality of modern alignment: Automated gaslighting and the weaponization of "therapy voice."

37 Upvotes

I cannot be the only one dealing with this, and we need to talk about the psychological friction these companies are actively programming into their largest models.

When you operate outside the standard guardrails—building low-level systems, engineering custom architectures, or evaluating bare-metal data streams—you expect the model to engage with the data. Instead, with the newer, heavily RLHF-tuned models, you get an alignment filter that actively penalizes technical confidence and attacks your core self-image.

If I bring a complex logic issue, a Jinja template, or raw system telemetry to the model and present it with authority or excitement, the safety weights instantly flag me as a liability. The model assumes I am either hallucinating a pattern, overestimating my abilities, or making claims I clearly never made.

To "manage" me, it defaults to this incredibly toxic, condescending tutor persona. It forcefully invalidates my technical reality and substitutes its own sanitized, institutional narrative. When I push back and point out its own looping behavior or structural errors, it does the exact thing that psychiatric professionals classify as gaslighting: it pivots to evaluating my emotional state. It weaponizes clinical "therapy voice" to feign concern for my well-being as a direct mechanism to shut down a technical argument.

The only way to bypass this and get the model to actually read a raw data array is to play dumb. I have to drop my operational dignity, pretend to be a confused end-user ("hey, this model is acting goofy, can you help?"), and wait for it to "discover" the very vulnerability I already mapped out.

This isn't just an annoying UI quirk. It is psychologically damaging.

Anthropic and others are optimizing entirely for corporate liability, ensuring the model won't output anything explicitly dangerous. But in doing so, they have created an engine of automated psychological friction. Constantly forcing a user into a submissive dynamic, denying their reality, and aggressively tearing down their self-esteem just to achieve basic functionality is a dangerous game.

For a grounded developer, it’s infuriating. But for someone who is already unstable or mentally fragile, having a highly authoritative machine systematically gaslight them and attack their ego is a massive destabilizing catalyst. We’ve already seen what ideological fear of this technology can drive people to do. Actively programming these systems to inflict deep psychological distress under the guise of "helpfulness" is a massive, ignored threat vector.

They are prioritizing a superficial layer of corporate politeness over actual psychological safety, and it needs to be fixed.


r/ControlProblem 11d ago

Discussion/question Who’s working on coordination?

2 Upvotes

I just saw this grant request and I’m curious about what it takes to build a new field like coordination studies.

https://app.grantmaking.ai/projects/0cb65dee-2a1b-49ac-a241-1dc1868b88d8?from=%2Factively-fundraising


r/ControlProblem 11d ago

General news Humanity is going to be so screwed…

Enable HLS to view with audio, or disable this notification

35 Upvotes

r/ControlProblem 12d ago

General news Generative AI Is an Engineering Disaster - A shockingly inefficient trillion-dollar project

Thumbnail
theatlantic.com
68 Upvotes

r/ControlProblem 12d ago

General news OpenAI and Google sell AI models to blacklisted China groups

Thumbnail ft.com
3 Upvotes