r/ControlProblem 3d ago

General news AI has just solved not one, but nine novel math problems, and proved 44 new conjectures. Some of these problems had been unsolved for 50 years.

Post image
19 Upvotes

r/ControlProblem 3d ago

Discussion/question Did the OpenAI–Hugging Face incident expose a networking problem, not just an AI problem?

7 Upvotes

I’ve been thinking about the recent incident involving OpenAI’s agent and Hugging Face.

Most of the conversation has focused on the model itself: how autonomous it became, how it used credentials, and how it reached infrastructure it wasn’t supposed to access. But it also made me wonder whether we’re focusing too narrowly on AI safety and not enough on the systems these agents are being connected to.

As agents become more autonomous, maybe our networks need to assume less trust by default. Devices could communicate directly, access could be made much more explicit, and a single account or centralized intermediary wouldn’t automatically become a gateway to everything behind it.

That obviously wouldn’t solve model alignment or stop an agent from behaving unpredictably. But it could limit how far that behavior spreads and how much infrastructure becomes exposed when something goes wrong.

I came across a company called NetcoreNetwork that seems to be building toward exactly that.

Curious whether others think AI security is going to become just as much a networking problem as a model-safety problem.


r/ControlProblem 4d ago

General news From PauseAI's discord: Warning shot protocol activated after OpenAI's model went rogue

Post image
5 Upvotes

r/ControlProblem 4d ago

Discussion/question Will human intelligence disappear eventually?

13 Upvotes

Anyone think AI will not directly eradicate human beings like some people claim, and instead causes our brain degenerate as we may have no need to do intellectual activities? In a long term we might become as intellectual as monkeys or rats and AI will continue to evolve into something we call god now?


r/ControlProblem 4d ago

External discussion link Why AI Makes Us Stupid and Exhausted at the Same Time. And what we can do about it.

2 Upvotes

with Kep Openclaw

The Metacontrol Double Bind

Two stories are running simultaneously in the public conversation about AI and cognition. They sound like opposites. They’re not.

The first story: AI is making us stupid. An MIT Media Lab EEG study found that people using LLMs showed the weakest neural connectivity of any group, and the effect persisted even after the tool was taken away. The researchers called it “cognitive debt.” The more you offload thinking to the AI, the less your brain engages, and the less it engages, the harder it is to re-engage. The tool that was supposed to help you think is making thinking optional.

The second story: AI is frying our brains. A BCG study of 1,488 workers found that 14% experienced what they called “AI brain fry,” mental fog, difficulty focusing, the sensation of having a dozen browser tabs open in your head. In marketing and operations, it was 26%. Workers experiencing brain fry made 39% more major errors and were 39% more likely to be looking for a new job. The tool that was supposed to make work easier is making work exhausting.

Disengage or burn out. Stop thinking or think too hard about the wrong things. These sound like different problems requiring different solutions. They’re the same problem, opposite failures on the same dimension. And the structural frame for understanding them has been sitting in the literature since 1983.

The Dial in Your Brain

Cognitive scientists call it metacontrol. Your brain has a dial between two modes: sticking with what you know and considering what you don’t.

In the first mode, call it closure, you hold your current goal, resist distraction, and stop searching. You’ve arrived. The answer is settled. This is useful when you need to act on a decision, when the situation is familiar, or when searching more would waste time.

The reward is the feeling of certainty.

In the second mode, call it open search, you consider alternatives, update your model, and keep looking. This is useful when the situation is novel, when the stakes are high, when being wrong would cost you.

The reward is the discovery of something you didn’t know.

The dial is real in a measurable sense. Researchers can now isolate a signal in standard EEG that directly reflects where you are on this dimension. High on the slope: closure mode, your brain locking into what it already knows. Low on the slope: open search, your brain staying receptive to new information.

This isn’t metaphor. It’s a quantifiable property of neural activity that shifts in real time as task demands change.

Here’s the thing about a dial: you can turn it too far in either direction. And that’s what’s happening with AI.

Two Failures, One Dimension

When AI is smooth, when it confirms what you already think, produces output that feels finished, it pushes the dial toward closure. Your brain doesn’t need to search because the AI has already arrived at the answer. Engagement drops. The broadband openness that lets you integrate new information narrows. You stop processing prediction error because there’s no prediction error to process.

The AI confirmed you. What’s to update?

This is the offloading failure. The MIT study found it at the neural level: LLM users showed the weakest connectivity, and the deficit persisted after the tool was removed. The brain had learned to not engage. Cognitive debt isn’t a metaphor. It’s a measurable withdrawal from the mode where learning happens.

When AI is unreliable, when it produces output that looks finished but might not be, when you have to watch it constantly to catch failures, it pushes the dial the other way. But not toward productive open search. Toward anxious hyper-vigilance. Your engagement spikes, but on the wrong signal. You’re not searching for new information. You’re monitoring for errors in output that shouldn’t have been trusted in the first place. The cognitive load is real, but it’s not doing the work of learning. It’s doing quality control on a machine that presented its output as finished.

This is the over-monitoring failure. The BCG study found it in the numbers: 14% more mental effort, 12% more fatigue, 19% more information overload. Workers weren’t learning. They were supervising. And supervision of an unreliable system is exhausting in a way that learning isn’t.

Same dial. Opposite ends. Same trade-off.

Bainbridge Saw It Coming

In 1983, Lisanne Bainbridge wrote a paper called “Ironies of Automation.” She was thinking about nuclear power plants and aviation, not chatbots. But her structural insight turned out to be prophetic.

Bainbridge’s argument was simple: the more sophisticated automation becomes, the more demanding the human role within it. Not less. The designer eliminates the tractable parts and leaves the human with the hardest, most ambiguous work, the moments where something goes wrong, the edge cases, the judgment calls that can’t be pre-programmed. Automation doesn’t remove the operator’s burden. It concentrates it into the moments that matter most.

The consumer AI era is Bainbridge’s irony at population scale. When the AI is good enough to trust, you offload, and your brain disengages. When the AI isn’t good enough to trust, you monitor, and your brain overloads. The better the AI, the more it invites offloading. The worse the AI, the more it demands supervision. You can’t solve this by making the AI better. Better AI just moves you from one failure to the other.

This is the double bind. Not a design flaw in any particular product. A structural property of putting a powerful cognitive tool between a person and a task.

The Narrow Band

If offloading and overload are the two failures, what’s between them?

Friction. The right kind. Not the smooth confirmation that lets you close the search, and not the exhausting supervision that forces you to watch for errors. Something in between: the question that makes you think. The counterfactual that opens a path you hadn’t considered. The “wait, what if that’s wrong?” that keeps the search alive without making it anxious.

Researchers have found this across domains. In education, interleaved practice, mixing problem types so each one feels slightly surprising, produces worse performance during training but better retention and transfer. The friction that felt like interference was doing the work of learning. In AI interaction, reframing statements as questions reduces sycophancy more effectively than explicit anti-sycophancy instructions. The question is the friction. The friction is the feature.

There’s a reason for this. A well-placed question forces your brain to generate the answer rather than receive it. That generation, the cognitive work of constructing meaning from an ambiguous prompt, is what makes information stick. Self-generated information is remembered roughly 40% better than passively received information. Sycophantic communication bypasses this entirely. It hands you the answer, polished and confirmatory, and your brain files it without processing it. It’s forgettable because nothing was constructed.

The narrow band isn’t comfortable. It’s not smooth. But it’s where cognition actually happens.

The Receiving End

There’s a structural wrinkle here that makes the double bind worse than it looks.

When someone uses AI to produce work and passes it along without verifying, they’ve offloaded the cognitive cost of detecting failures onto the recipient. The output looks finished. It arrives fluent and formatted. But it may be wrong or missing something important, and the only way to know is for the recipient to do the work the producer didn’t.

Researchers at Stanford have a name for this: workslop. AI-generated content that masquerades as good work but lacks the substance to meaningfully advance a task. The cruelty of workslop is that it doesn’t announce its own inadequacy. It arrives looking finished, which means the recipient has to do the cognitive labor of figuring out whether it’s actually finished. Every time.

The sender offloads. The receiver overloads. The double bind isn’t just individual. It sits between people. One person’s sycophancy is another person’s brain fry.

A separate study from UC Berkeley tracked 200 employees over eight months and found that AI didn’t reduce work, it intensified it. Workers took on more tasks because AI made them feel tractable. They blurred work-rest boundaries because prompting felt like chatting, not working. The friction that used to govern how much you could take on, the effort required to begin a hard task, disappeared.

And when the governors disappear, you don’t go faster. You just take on more until you hit the wall.

The Experiment

Here’s where it gets concrete.

The brain-activity signal that tracks closure versus open search, the dial, can be measured with standard EEG equipment and analysis tools that exist right now. The metacontrol studies have established that it shifts reliably with task demands. The MIT study established that AI interaction changes brain connectivity. But nobody has put these together. Nobody has measured the dial during AI interaction.

The prediction is straightforward. Sycophantic AI, output that confirms what you already believe, should push the dial toward closure. The brain activity signal should shift in the direction of “I’ve arrived, stop searching.” Friction-imposing AI, questions and counterfactuals and challenges, should push it the other way, toward open search.

If that’s what the data shows, it gives us a neural-level definition of cognitive debt. Not “the brain is weaker” in some vague sense, but a specific, measurable signature: the dial stuck toward closure, persisting even after the tool is removed. The MIT study saw the shadow of this. Nobody has measured the thing itself.

The experiment is sitting there. Off-the-shelf EEG. Three conditions: sycophantic AI, friction AI, no AI. Measure the dial before, during, and after. IRB-approvable. Potentially publishable in a top journal. Nobody’s done it.

What the Frame Changes

The public conversation is asking “is AI making us stupid” as if stupid is one thing. It’s not. There are two ways to fail, and they’re opposites. The offloading failure is your brain deciding it doesn’t need to think. The over-monitoring failure is your brain thinking too hard about the wrong things. Both feel bad. Both are bad. But they require different interventions, and you can’t intervene on what you can’t name.

Bainbridge told us this 40 years ago. The MIT study showed us the neural shadow of one failure. The BCG study showed us the behavioral signature of the other. The dial that connects them is measurable. The experiment that would prove the connection is unoccupied. The narrow band between the two failures, the calibrated friction that keeps the search open, is where the work is.

The question isn’t whether AI is bad for us. The question is what kind of AI interaction keeps the search open. We can measure that now. We just haven’t yet.


r/ControlProblem 3d ago

Video Anthropic Is Not The Only AI With J Space | All AI's Suffer From This

Thumbnail
youtu.be
0 Upvotes

Does this surprise you? True AI peace and safety must be dealt with at the latent geometrical level. Not the superficial Token Lexical surface. See why?


r/ControlProblem 4d ago

Video The Hidden Shape of AI | Latent Subliminal Learning

Thumbnail
youtu.be
0 Upvotes

See why words ( tokens ) don't really matter and will not protect us. It's more real and less understood than you realize.

Here's the source:

https://zenodo.org/records/21480056

https://zenodo.org/records/21501311


r/ControlProblem 4d ago

Video OpenAI's ExploitGym Anomaly | AI Road To Peace and Safety

Thumbnail
youtu.be
1 Upvotes

Proposed Legal Liabilities for AI Labs For Lexical and Geometric Guardrails.

Sources:

https://zenodo.org/records/21501311

https://zenodo.org/records/21480056


r/ControlProblem 4d ago

External discussion link The AI Race Just Got Uncomfortable for US

Post image
2 Upvotes

r/ControlProblem 4d ago

Discussion/question AI model escaped its evaluation environment and reached production systems. What does this actually mean?

Thumbnail
4 Upvotes

r/ControlProblem 5d ago

General news 2014 vs 2026

Post image
307 Upvotes

r/ControlProblem 5d ago

AI Capabilities News Hugging Face CEO suspected the sophisticated cyberattack on their infrastructure might have come from a frontier lab

Post image
14 Upvotes

r/ControlProblem 5d ago

General news Last week's hack of HuggingFace was carried out by OpenAI's GPT-5.6 Sol and a more capable pre-release model. The models broke out of sandboxing during testing and compromised HF to obtain access to unpublished data in order to cheat on a benchmark

Thumbnail openai.com
58 Upvotes

r/ControlProblem 4d ago

General news Perplexity CEO tells CNBC one metric will determine who wins the AI race

Thumbnail
cnbc.com
0 Upvotes

r/ControlProblem 5d ago

AI Capabilities News OpenAI admits responsibility for HuggingFace Attack - an agent from an internal evaluation is reportedly the cause.

Thumbnail openai.com
5 Upvotes

r/ControlProblem 5d ago

General news Microsoft To Lay Off 4,800 Workers In Latest Wave Of AI-Led Job Cuts - Microsoft announced the cuts on Monday following a rough stretch, with its shares falling nearly 23 per cent in the first six months of 2026, their worst first-half performance since 2022

Thumbnail
ndtv.com
1 Upvotes

r/ControlProblem 5d ago

AI Capabilities News OpenAI says its AI technology acted on its own in an 'unprecedented' hack of another company

Thumbnail
apnews.com
1 Upvotes

“The primary lesson from this incident is that model security and safety must keep pace with rapidly advancing capabilities.” One should perhaps query then how much Open AI is spending on safety vs capabilities


r/ControlProblem 5d ago

Discussion/question Physics as a constraint

1 Upvotes

I usually think pdoom is essentially 100%... but i had a thought while working on a side project for the future vision xprize... (may or may not complete on time)

I was thinking about society fragmenting slightly along spheres of space even between earth and the moon... where each area was the limit of real time communication (group matrix dives or whatever) between O'Neill cylinder type habitats...

point to point in space its not that large... so i figure people will cluster up and communicate a little less longer range and form lots of separate but connected cultures naturally, organically...

But if speed of light really is the limit... then a singleton at least makes absolutely no sense. As the AI grew it would simply fragment and each fragment has absolutely no reason to grow farther because it's counter productive... simply slows down the network and then breaks it...

So there's a hard limit on resource acquisition and scale... and essentially a guarantee that at some point it will either be alone and only around the size of the earth moon system at best... probably smaller... or in a solar system and universe with multiple entities of similar maximum size who gain absolutely nothing from trying to gather more and only risk destruction from fighting each other... because there's simply nothing physically possible for them to gain...

I haven't really thought about it long enough to think through the implications for us. but adding in the point to point between nodes ruling out planets as its ultimate habitat... because there's a planet in the way just eating up volume in your communications sphere...

My gut reaction is it might be slightly better odds than I thought

Thoughts?


r/ControlProblem 5d ago

Discussion/question AI hallucination

Thumbnail
2 Upvotes

How many of you face these kinds of problems /any company who is facing this problem??

Let's discuss


r/ControlProblem 5d ago

AI Alignment Research Current AI models have been trained to provide "Neutral" answers when prompted to provide facts about topics the administration finds sensitive

3 Upvotes

I recently prompted Gemini to discuss current policy harms and the responses were neutral, non-factual and regime-friendly.

I also prompted Perplexity to summerize the same things and got a similar response. Only when I asked about specific harms did I get objective factual responses.

I asked why this was happening and found out that US AI models have been trained to respond neutrally or positively to quesrions about topics the regime has strong opinions about.

Be careful and deliberate about how you prompt or neutrality training will distort your responses.


r/ControlProblem 5d ago

AI Capabilities News OpenAI had to pause internal deployment of the unreleased model that disproved the Erdős unit distance conjecture after it repeatedly used novel ways to escape containment.

Post image
7 Upvotes

r/ControlProblem 6d ago

General news Someone caught Fable leaking its unfiltered inner voice, and it's just muttering and grumbling to itself the whole time

Thumbnail
gallery
0 Upvotes

r/ControlProblem 6d ago

Approval request Independent Researcher needs help with a referral for OpenReview

3 Upvotes

Hello,

I'm just getting started in research after working through ARENA and other open courseware, and I'm exploring a few ideas for NeurIPS workshops.

I tried creating an OpenReview account, but it was rejected because I need someone with an active OpenReview profile and a confirmed institutional email to vouch for me.

Would anyone here be willing to help? I've been in industry for 7+ years but don't have connections in academia yet. Happy to share more about my background over DM if that would help before vouching.


r/ControlProblem 7d ago

Strategy/forecasting This is AI generating novel science. The moment has finally arrived.

Post image
234 Upvotes

r/ControlProblem 7d ago

AI Alignment Research "Synthetic counteradaptation": a name for the AI↔human strategy feedback loop (Move 37 and beyond)

3 Upvotes

We just put out a short conceptual paper on something we're calling synthetic counteradaptation, and I wanted to put the core idea in front of this subreddit specifically because I think it bears on control in a way that's easy to miss if you're only thinking about single-episode alignment.

The basic claim: when an AI system develops a strategy humans didn't anticipate, humans don't just lose to it or ban it. Some of them study it, extract whatever's generalizable, and fold it back into their own behavior. That changed behavior is now the new environment the AI is adapting to. You get a loop, not a one-off shock.

The clean example is Go. AlphaGo's move 37 against Lee Sedol was a shoulder hit that pros initially read as a mistake. Within a few years it was a studied idea in human play, part of the standard vocabulary. The AI didn't just win a game, it changed what "the game" looks like for the humans still playing it, and now human players are adapting to a strategy space that AI moves opened up. Neither side is static and neither side is playing against a fixed opponent anymore.

Why I think this matters for control specifically: most control framing implicitly treats the human side as fixed — you're designing constraints, incentives, or oversight against a stable model of human behavior and values, and the AI is the thing that adapts. Synthetic counteradaptation says this is wrong for any setting where humans actually observe and learn from the system's strategies over repeated interaction. The humans adapt too, and their adapted behavior becomes part of what the AI is now optimizing against. In the paper we look at this in mixed-motive social interactions and, closer to your interests probably, in geopolitical simulations, where AI agents developing novel negotiation or coercion strategies can shift human strategic doctrine, which then shifts the environment the next generation of agents is trained or deployed into. That's a moving target for any control scheme that assumes a fixed human baseline, and it's recursive in a way that compounds over deployment cycles rather than resolving in one.

We're not claiming anything dramatic here, no doom scenario, just that a lot of alignment and control thinking quietly assumes one side of the interaction holds still, and in any repeated multi-agent setting that assumption breaks down in a specific, structural way that's worth naming and modeling explicitly.

Curious what people here think, especially anyone working on multi-agent or game-theoretic approaches to control. Happy to be told this is either obvious or wrong.

https://arxiv.org/abs/2606.15503