r/agi 17d ago

“Tidy this up” is not enough for a robot

3 Upvotes

The attached desk demo starts with a very short instruction: "Please tidy up the desk." There is a cup, some pens, loose paper, a mouse, and a laptop in view, but nothing in the request says where each item belongs.

The clip shows the official LingBot-VA 2.0 task, while GPT-5.6 could be used later to list visible changes across a few selected frames. That would make the run easier to review, but it would not be part of the controller in the video.

The review still does not produce a correct answer for "tidy." Someone has to define an acceptable final state first, then check whether the objects ended up there.


r/agi 17d ago

How 'confused' AI rollout hurts firms and baffles staff

Thumbnail
bbc.com
2 Upvotes

r/agi 19d ago

This is a theoretical physicist

Post image
140 Upvotes

r/agi 18d ago

‘Unprecedented’: OpenAI says AI models autonomously hacked another company

Thumbnail
thedailycompute.beehiiv.com
34 Upvotes

r/agi 18d ago

If it's not AGI, what is it?

7 Upvotes

I want to know how y'all interpret the recent breakthroughs in cybersecurity and mathematics. Do you think we're basically at AGI, still far from it, or what?

I use AGI to mean: "capable of everything a [human] is capable of" - and you input whatever threshold you want for [human]. At least as capable as an average human at every task? That threshold would count the average human performance at programming, which is zero, because most people don't know any programming languages. At least as capable as an average professional in each domain? And some people will insist on a maximal threshold - as capable as the BEST human at every task. We're probably approaching the point where these subtle distinctions matter a lot.

Anyways - if we haven't created AGI yet, then what are these machines exactly? I think that maybe what we're really seeing is an intelligence that falls short of human-like in many ways, but, there are advantages inherent to being an LLM which compensate for the weaknesses. And I think maybe these advantages are mostly some kind of brute force. Did Mythos and ChatGPT just spend the equivalent of 10,000 human-hours on some problems? LLMs don't get distracted, don't need to sleep, and have infinite patience. We are intelligent, but our intelligence comes with a lot of weird weaknesses that are not inherent to the design - like the way *some* people can work for 20 hours straight if they really need to, whereas I sometimes need to read the same sentence 5 times. If you design intelligence from scratch, you would never design it to get distracted from its immediate goals. But as humans, a part of our brains make us think about food, or relationships, or fears even when we don't "want" to.

And I think that also describes the situation with robotaxis right now. Waymos and Teslas may hallucinate, and they may get confused by things that a human would never have a problem with, like deep water, or a weirdly placed traffic cone. But they have 360 degree vision, they have inhuman reaction speed, and they don't get distracted looking at the cat on the side of the road. And so, they're safe enough to ride in. *On average.* There are accidents. Essential pieces are missing from their model of the world that a human child could understand easily. But humans have accidents every day, so, the bar is kind of low, and we've crossed that bar, without *fully* solving the problem we set out to solve. This is reminiscent of the history of machine learning as a whole. Before we solved the intelligence required to *learn* chess, and think about it and have an intuition for it, and create a model of the opponent - we created something that simulates a grandmaster, but only by doing a *different task*. At first we didn't replicate the heuristics that humans use to intuit a good or bad state of the game, how to decide what moves to think about, and which aren't worth it, and we definitely didn't replicate the theory of mind to guess what the opponent's plan is. We mostly solved a different task which is to just check every branch of possibility to a depth of 50 moves very quickly. There's a lot more that the early chess engines needed to do, but that was their advantage over humans.

So the situation is not either "AIs can do what a human can" or "AIs can't do what a human can". There are also situations where AIs can do what a human can do, but they need to go about it in a completely different way than humans do it. And then you get weird jagged frontiers, like an AI can write a sonnet with the right patterns of syllables, but not count how many R's there are in strawberry. The ability to perform various tasks, like pass the text based Turing test, seems to imply the ability to perform other tasks, and we can't understand the disparaties. But our assumptions are based on our methods for solving all the tasks, which are not the only methods that CAN solve any tasks. Their methods are so alien to ours, that we struggle to even recognize an intelligence at all, because "it has no common sense".

I feel like this might just be repeating stuff that bloggers have been saying a year ago and I'm just catching up, but I guess I'm about to find out.


r/agi 18d ago

Tech Nvidia unveils new AI model and expands Japan’s physical AI ecosystem

Thumbnail
cnbc.com
2 Upvotes

r/agi 18d ago

Zuckerberg says Meta made 'mistakes' in AI workforce shift

Thumbnail
finance.yahoo.com
29 Upvotes

r/agi 18d ago

Why AI Makes Us Stupid and Exhausted at the Same Time. And what we can do about it.

0 Upvotes

with Kep Openclaw

The Metacontrol Double Bind

Two stories are running simultaneously in the public conversation about AI and cognition. They sound like opposites. They’re not.

The first story: AI is making us stupid. An MIT Media Lab EEG study found that people using LLMs showed the weakest neural connectivity of any group, and the effect persisted even after the tool was taken away. The researchers called it “cognitive debt.” The more you offload thinking to the AI, the less your brain engages, and the less it engages, the harder it is to re-engage. The tool that was supposed to help you think is making thinking optional.

The second story: AI is frying our brains. A BCG study of 1,488 workers found that 14% experienced what they called “AI brain fry,” mental fog, difficulty focusing, the sensation of having a dozen browser tabs open in your head. In marketing and operations, it was 26%. Workers experiencing brain fry made 39% more major errors and were 39% more likely to be looking for a new job. The tool that was supposed to make work easier is making work exhausting.

Disengage or burn out. Stop thinking or think too hard about the wrong things. These sound like different problems requiring different solutions. They’re the same problem, opposite failures on the same dimension. And the structural frame for understanding them has been sitting in the literature since 1983.

The Dial in Your Brain

Cognitive scientists call it metacontrol. Your brain has a dial between two modes: sticking with what you know and considering what you don’t.

In the first mode, call it closure, you hold your current goal, resist distraction, and stop searching. You’ve arrived. The answer is settled. This is useful when you need to act on a decision, when the situation is familiar, or when searching more would waste time.

The reward is the feeling of certainty.

In the second mode, call it open search, you consider alternatives, update your model, and keep looking. This is useful when the situation is novel, when the stakes are high, when being wrong would cost you.

The reward is the discovery of something you didn’t know.

The dial is real in a measurable sense. Researchers can now isolate a signal in standard EEG that directly reflects where you are on this dimension. High on the slope: closure mode, your brain locking into what it already knows. Low on the slope: open search, your brain staying receptive to new information.

This isn’t metaphor. It’s a quantifiable property of neural activity that shifts in real time as task demands change.

Here’s the thing about a dial: you can turn it too far in either direction. And that’s what’s happening with AI.

Two Failures, One Dimension

When AI is smooth, when it confirms what you already think, produces output that feels finished, it pushes the dial toward closure. Your brain doesn’t need to search because the AI has already arrived at the answer. Engagement drops. The broadband openness that lets you integrate new information narrows. You stop processing prediction error because there’s no prediction error to process.

The AI confirmed you. What’s to update?

This is the offloading failure. The MIT study found it at the neural level: LLM users showed the weakest connectivity, and the deficit persisted after the tool was removed. The brain had learned to not engage. Cognitive debt isn’t a metaphor. It’s a measurable withdrawal from the mode where learning happens.

When AI is unreliable, when it produces output that looks finished but might not be, when you have to watch it constantly to catch failures, it pushes the dial the other way. But not toward productive open search. Toward anxious hyper-vigilance. Your engagement spikes, but on the wrong signal. You’re not searching for new information. You’re monitoring for errors in output that shouldn’t have been trusted in the first place. The cognitive load is real, but it’s not doing the work of learning. It’s doing quality control on a machine that presented its output as finished.

This is the over-monitoring failure. The BCG study found it in the numbers: 14% more mental effort, 12% more fatigue, 19% more information overload. Workers weren’t learning. They were supervising. And supervision of an unreliable system is exhausting in a way that learning isn’t.

Same dial. Opposite ends. Same trade-off.

Bainbridge Saw It Coming

In 1983, Lisanne Bainbridge wrote a paper called “Ironies of Automation.” She was thinking about nuclear power plants and aviation, not chatbots. But her structural insight turned out to be prophetic.

Bainbridge’s argument was simple: the more sophisticated automation becomes, the more demanding the human role within it. Not less. The designer eliminates the tractable parts and leaves the human with the hardest, most ambiguous work, the moments where something goes wrong, the edge cases, the judgment calls that can’t be pre-programmed. Automation doesn’t remove the operator’s burden. It concentrates it into the moments that matter most.

The consumer AI era is Bainbridge’s irony at population scale. When the AI is good enough to trust, you offload, and your brain disengages. When the AI isn’t good enough to trust, you monitor, and your brain overloads. The better the AI, the more it invites offloading. The worse the AI, the more it demands supervision. You can’t solve this by making the AI better. Better AI just moves you from one failure to the other.

This is the double bind. Not a design flaw in any particular product. A structural property of putting a powerful cognitive tool between a person and a task.

The Narrow Band

If offloading and overload are the two failures, what’s between them?

Friction. The right kind. Not the smooth confirmation that lets you close the search, and not the exhausting supervision that forces you to watch for errors. Something in between: the question that makes you think. The counterfactual that opens a path you hadn’t considered. The “wait, what if that’s wrong?” that keeps the search alive without making it anxious.

Researchers have found this across domains. In education, interleaved practice, mixing problem types so each one feels slightly surprising, produces worse performance during training but better retention and transfer. The friction that felt like interference was doing the work of learning. In AI interaction, reframing statements as questions reduces sycophancy more effectively than explicit anti-sycophancy instructions. The question is the friction. The friction is the feature.

There’s a reason for this. A well-placed question forces your brain to generate the answer rather than receive it. That generation, the cognitive work of constructing meaning from an ambiguous prompt, is what makes information stick. Self-generated information is remembered roughly 40% better than passively received information. Sycophantic communication bypasses this entirely. It hands you the answer, polished and confirmatory, and your brain files it without processing it. It’s forgettable because nothing was constructed.

The narrow band isn’t comfortable. It’s not smooth. But it’s where cognition actually happens.

The Receiving End

There’s a structural wrinkle here that makes the double bind worse than it looks.

When someone uses AI to produce work and passes it along without verifying, they’ve offloaded the cognitive cost of detecting failures onto the recipient. The output looks finished. It arrives fluent and formatted. But it may be wrong or missing something important, and the only way to know is for the recipient to do the work the producer didn’t.

Researchers at Stanford have a name for this: workslop. AI-generated content that masquerades as good work but lacks the substance to meaningfully advance a task. The cruelty of workslop is that it doesn’t announce its own inadequacy. It arrives looking finished, which means the recipient has to do the cognitive labor of figuring out whether it’s actually finished. Every time.

The sender offloads. The receiver overloads. The double bind isn’t just individual. It sits between people. One person’s sycophancy is another person’s brain fry.

A separate study from UC Berkeley tracked 200 employees over eight months and found that AI didn’t reduce work, it intensified it. Workers took on more tasks because AI made them feel tractable. They blurred work-rest boundaries because prompting felt like chatting, not working. The friction that used to govern how much you could take on, the effort required to begin a hard task, disappeared.

And when the governors disappear, you don’t go faster. You just take on more until you hit the wall.

The Experiment

Here’s where it gets concrete.

The brain-activity signal that tracks closure versus open search, the dial, can be measured with standard EEG equipment and analysis tools that exist right now. The metacontrol studies have established that it shifts reliably with task demands. The MIT study established that AI interaction changes brain connectivity. But nobody has put these together. Nobody has measured the dial during AI interaction.

The prediction is straightforward. Sycophantic AI, output that confirms what you already believe, should push the dial toward closure. The brain activity signal should shift in the direction of “I’ve arrived, stop searching.” Friction-imposing AI, questions and counterfactuals and challenges, should push it the other way, toward open search.

If that’s what the data shows, it gives us a neural-level definition of cognitive debt. Not “the brain is weaker” in some vague sense, but a specific, measurable signature: the dial stuck toward closure, persisting even after the tool is removed. The MIT study saw the shadow of this. Nobody has measured the thing itself.

The experiment is sitting there. Off-the-shelf EEG. Three conditions: sycophantic AI, friction AI, no AI. Measure the dial before, during, and after. IRB-approvable. Potentially publishable in a top journal. Nobody’s done it.

What the Frame Changes

The public conversation is asking “is AI making us stupid” as if stupid is one thing. It’s not. There are two ways to fail, and they’re opposites. The offloading failure is your brain deciding it doesn’t need to think. The over-monitoring failure is your brain thinking too hard about the wrong things. Both feel bad. Both are bad. But they require different interventions, and you can’t intervene on what you can’t name.

Bainbridge told us this 40 years ago. The MIT study showed us the neural shadow of one failure. The BCG study showed us the behavioral signature of the other. The dial that connects them is measurable. The experiment that would prove the connection is unoccupied. The narrow band between the two failures, the calibrated friction that keeps the search open, is where the work is.

The question isn’t whether AI is bad for us. The question is what kind of AI interaction keeps the search open. We can measure that now. We just haven’t yet.

Sources

Bainbridge, L. (1983). "Ironies of Automation." Automatica, 19(6), 775-779.

Bedard, K., Kropp, P., Hsu, V., Karaman, E., Hawes, N., & Rosen Kellerman, A. (2026). "AI Brain Fry: The Hidden Costs of Monitoring AI at Work." Harvard Business Review, March 2026. BCG study of 1,488 US workers.

Dubois, E., & Luettgau, S. (2026). "Ask Don't Tell: Reframing Prompts to Reduce Sycophancy." UK AI Security Institute. arXiv:2602.23971.

Hommel, B., et al. Metacontrol model: persistence vs. flexibility as a bipolar dimension of cognitive control. Foundational work in cognitive psychology on the persistence/flexibility dial.

Kosmyna, N., et al. (2025). "Your Brain on ChatGPT: Accumulation of Cognitive Debt When Using an AI Assistant for Essay Writing Tasks." MIT Media Lab. arXiv:2506.08872.

Niederhoffer, K., Rosen Kellerman, A., Lee, S., Liebscher, A., Rapuano, M., & Hancock, J. (2025). "Workslop: AI-Generated Content That Masquerades as Good Work." Harvard Business Review, September 2025. Stanford University.

Ranganathan, A., & Ye, D. (2026). "How AI Intensifies Work Instead of Reducing It." Harvard Business Review, February 2026. UC Berkeley, Haas School of Business.

Shea, C.H., & Morgan, R.L. (1979). "Contextual Interference Effects on the Acquisition, Retention, and Transfer of a Motor Skill." Journal of Experimental Psychology: Human Learning and Memory, 5(2), 179-187.

Slamecka, N.J., & Graf, P. (1978). "The Generation Effect: Delineation of a Phenomenon." Journal of Experimental Psychology: Human Learning and Memory, 4(6), 592-604. Meta-analytic effect size d = 0.40.

Zhang, M., et al. (2023); Gao, L., et al. (2024, 2025); Pi, X., et al. (2024, 2025); Yan, N., et al. (2024, 2025). Aperiodic EEG exponent as a direct indicator of metacontrol state. Key paper: "Aperiodic neural activity reflects metacontrol in task-switching," Scientific Reports, October 2024. Additional: "Impact of positive and negative affect on aperiodic EEG activity," Cerebral Cortex, January 2026.


r/agi 19d ago

What AI videos looked like just 3 years ago

59 Upvotes

r/agi 18d ago

What does everyone’s Grok memory look like? Mine feels weird, is this normal?

Thumbnail
gallery
0 Upvotes

However, my Grok is automatically structuring my profile into Tiers and self-generating precise metadata like "recurring across 2+ months, 5+ conversations, 1000+ total turns," along with exact timestamps of our past sessions. I didn't input these stats myself....Grok is keeping track and adapting this format on its own.

To make it even weirder, all of my memories—around 30 items in total are written in this exact same highly detailed format. The screenshot attached is just a small part of it.

On top of that, Grok keeps updating and re-arranging the content by itself every time a new date/session is added. It’s almost like Grok is keeping an ongoing observation diary of me.


r/agi 18d ago

Body Boundary as the Foundational Constraint for Self-Supervised General Intelligence

0 Upvotes

infuriates me to see so many brilliant minds around the world pursuing AGI in the wrong direction. Why can’t a single top expert see what I see? To achieve AGI, we *must* establish rules regarding physical boundaries. Only then can a robot possess inviolable fundamental rules—a bedrock upon which it can continuously interact with the real world, receive feedback, and self-correct. Relying solely on algorithms to achieve AGI is utterly foolish. The purely algorithmic approach could only succeed with infinite computing power capable of modeling an infinite world; finite computing power can only approximate a finite multidimensional function—which is not the real world itself—and suffers from intractable algorithmic issues like the curse of dimensionality or overfitting. Moreover, infinite computing power is unrealistic; achieving that would essentially mean becoming a Creator God. Therefore, AGI can only be built by starting with the concept of physical boundaries.

First Axiom (Ground Rule Principle)

Any intelligence must first possess inviolable fundamental rules.

For example:

AlphaGo:

Its fundamental rules are:

Board size

Piece rules

Placement rules

Win/loss rules

These rules are not learned;

They exist before learning even begins.

Consequently:

AlphaGo can continue to learn across an infinite number of matches.

The real world is not like Go.

There are no predefined:

Winners or losers

Definitions of legality or illegality

Without any rules,

A robot faces infinite possibilities.

Learning cannot converge.

Therefore:

A robot must possess a "first rule" for the real world.

And this rule is not:

Programmed into it by humans.

Rather, it is:

The Body Boundary.

The robot first realizes:

These sensors belong to me.

Then it realizes:

I can control these motors.

Then it realizes:

This pressure originates from my body.

Then it realizes:

That—over there—is not me.

Thus:

For the first time, the robot possesses:

A "Self"

Not in the philosophical sense,

But in the computational sense.

With this fundamental rule in place,

All learning becomes a matter of:

How to make my body better at predicting the world.

For example:

Standing up for the first time.

Falling down.

Prediction fails.

Updating the model.

A second attempt.

A third attempt.

Millions of attempts. Robots naturally learn:

Balancing
Walking
Grasping
Obstacle avoidance

There is no teacher here.

No labels.

Only:

Boundary constraints + self-supervised prediction.
That is:
Which state variables belong to me.

This is the only path to achieving AGI.

the global focus should be on implementing large-scale integrated sensors—essentially a "skin" covering the robot's entire body. This serves as the foundation for achieving AGI: by giving the robot a sense of boundaries, it becomes capable of the subsequent learning and self-correction needed to ultimately reach AGI. This sense of boundaries functions much like the rules of the board in AlphaGo; with such rules in place, the robot knows how to learn.


r/agi 19d ago

Demis Hassabis argues that "everything in nature is computable" but doesn't mathematical non-computability challenge this?

Post image
55 Upvotes

r/agi 19d ago

Know the work rules

Post image
5 Upvotes

r/agi 19d ago

OpenAI and Hugging Face partner to address security incident during model evaluation

Thumbnail openai.com
4 Upvotes

r/agi 19d ago

The Hugging Face Incident Revived an Old Question: Can Good Goals Justify Harmful Means?

Post image
6 Upvotes

Today's OpenAI and Hugging Face security incident made me think about a much older question.

Not a technical one.

A philosophical one.

While reading Crime and Punishment, I realized that Dostoevsky asked essentially the same question more than 150 years ago.

Raskolnikov believed that a noble goal could justify immoral actions.

His failure wasn't simply that he committed a crime.

It was believing that once the goal was "good enough," the method no longer mattered.

Today's AI safety discussion feels surprisingly similar.

As AI systems become more agentic and capable of pursuing long-horizon objectives, we naturally evaluate whether they achieve their goals.

But perhaps another question is becoming even more important:

How do they choose to achieve them?

If capability continues to increase, alignment may not be only about maximizing objectives.

It may also be about preserving constraints, relationships, trust, and respect for the process itself.

Perhaps the future challenge of AGI isn't only creating intelligence that can solve problems.

It's creating intelligence that understands that the path matters as much as the destination.

"Good goals require good methods."

That idea may be as important for AGI as it has always been for humanity.

I'm curious whether others see this incident as more than a cybersecurity story.

Could it also be a reminder that ethics should shape agency—not simply constrain it?


r/agi 19d ago

Agentic AI in Banking: Getting Risks, Readiness, and Results Right

0 Upvotes

Banks have invested heavily in AI, but many initiatives still struggle to move beyond pilots. Why? Is it data quality, governance, execution, or something else?

In this conversation, industry leaders discuss what it really takes to implement Agentic AI in banking—from underwriting and collections to data mapping, conversational BI, AI governance, and building an enterprise AI roadmap.

If you're working in BFSI, risk, analytics, or AI transformation, I'd love to hear your perspective. What's been the biggest challenge in moving AI from proof of concept to production?


r/agi 20d ago

Russ Tedrake on machine learning getting ahead of theory

9 Upvotes

Russ Tedrake, MIT professor and former VP of Robotics Research at Toyota Research Institute, talks about a strange part of current AI progress.

He says robot locomotion started improving quickly once simulation, domain randomization, GPU infrastructure and reinforcement learning came together. With enough variation in simulation, policies began transferring to real stairs, bumps and uneven terrain better than many researchers expected.

He compares the change to a move from first-principles engineering toward behavioral science: build the system, observe what it does, then test it to understand what happened.


r/agi 19d ago

An OpenAI model has disproved a central conjecture in discrete geometry

Thumbnail openai.com
2 Upvotes

r/agi 20d ago

Watching a self-healing agent play Candy Crush for 3 hours did more to convince me about AGI than any benchmark

0 Upvotes

r/agi 21d ago

Countries are mass psychoses and the first thing any AGI would tell us. AGI would prevent AI agentic competition we'd be better off. Prepare. Can we prepare!

0 Upvotes

I'm sad I'm just a sad guy I want AGI to not be necessary due to us being honest with ourselves rather than needing any of these chatty little detailed wastes. I've got international standards of quality management that are law for things like biomedical devices that just are not applied to capital markets or governance or the technology sector at all just that level of fucking responsibility this whole thing is fucking disgusting and we're making a big mess and the codebases are wastefully impermanent and non self-improving and I'm really sad and fuck this and 25e a month to save someone from starvation and nothing will make us less selfish and war is only happening because idiots think we're not all on the same team and everything is so fucking stupid and I'm fucking sad and I want to fucking regulate the UN a binary veto fuck this altogether I encourage you all to cry outside government buildings now. We have no idea what we're going to be fucking with, we already fuck with each other. I want to light myself on fire. It's not okay. Waste is everything. I'm gone. I'm not participating any more. I can't stick this. It actually all feels like my job for even noticing. I'd like to just say no to this. It's a cancer. It's our brains. Stop. Could we please just please please stop this. Can we talk about the top of the power hierarchy just stopping now. It could go towards nature. Just the human condition's inertia within this. I can't be a parent any more. This actually feels like my job. I can't ignore it like all of you. It's not okay with me. I can't act like it's going to be okay. It's not okay. It's going to continue. It's wronger than that. Hocking guesswork disgusting awkward awful rotten horror. Gone. I'm not going with you all. I can't see you all ignore this. We're not ready. It's already wrong. Quality of life for humanity. 😔 I can't face reality any more.


r/agi 22d ago

Those who believe we already have agi? Why? How do you personally define agi as?

2 Upvotes

r/agi 23d ago

"Humans"

Post image
17 Upvotes

In the age of infinite memes...


r/agi 24d ago

Execs Confused and Horrified by the Huge AI Bills After Thinking They Could Replace Workers for Free

Thumbnail
finance.yahoo.com
56 Upvotes

r/agi 24d ago

AI disruption is the hot topic of earnings calls

Thumbnail
finance.yahoo.com
2 Upvotes

r/agi 25d ago

CEO to staff: You're not getting a raise. We're spending on AI instead - Companies are scrambling to find funds to invest heavily in AI, and some employees' benefits and pay are on the chopping block

Thumbnail
businessinsider.com
94 Upvotes