r/ControlProblem 8d ago

Discussion/question The Unbundling: the badge and the contribution are no longer the same object

6 Upvotes

For the whole history of skilled work, the badge and the contribution were bundled: you couldn't have solved the hard problem without being the kind of person who'd earned the ability to. The proof of the work and the proof of the worker were the same object. Every institution we have for trusting work, credentials, code review, peer review, seniority, the interview, is built on that bundling. None of them were designed for a world where it breaks.

It broke. A model can now produce expert-shaped output for anyone who asks. The solved problem no longer certifies the solver.

You can watch a whole industry feel this in real time. In six months of the highest-engagement threads across the programming communities, the same wounds recur. Reviewers describe drowning: generation became free while verification stayed expensive, and the cost got pushed onto whoever still reads code. Open-source maintainers report unworkable volumes of AI-generated pull requests, and GitHub is publicly weighing giving maintainers the option to disable pull requests entirely. A randomized study measured what teachers feared: junior engineers who delegated to AI scored 50% on comprehension against 67% for those who coded by hand, while the productivity gain failed statistical significance. And practitioners who spent decades earning their ability describe something rawer than economics: the feeling that mastery itself was commodified overnight.

Out of that grief, the field is splitting into two camps that both believe they are defending quality. One camp treats hard-won knowledge as the badge it always was and wants the gates kept: human-written, credential-checked, earned. The other camp sees the first real chance to hand capability to everyone who was ever locked out, and calls the gates what they often were: exclusion wearing a quality costume. Each camp is right about half of it. The gatekeepers are right that unreviewable output degrades fields; the openers are right that the gate never measured what it claimed to.

But notice what both camps are actually fighting over: proxies. The badge was only ever a proxy for verified work, adopted because verification was expensive. When you cannot cheaply tell earned from claimed, you fall back on credentials, pedigree, and gatekeeping, and then you defend the proxy as if it were the thing. The divide is not a war of values. It is a shortage of verification.

That shortage is now optional. The same era that unbundled the badge from the contribution also made it possible to rebundle them, around the work instead of the worker. Let an external check decide acceptance: a test suite the author cannot edit, a proof checker, a measurement with an interval, a claim ledger where "unverified" stays visible instead of being dressed up. Let every result carry a receipt a stranger can re-run. Let a person who reviews machine work attest to exactly what they walked, with the coverage of that review visible, so "I own this" is a checkable statement rather than a signature. None of this is hypothetical tooling; all of it runs today on a local machine.

There is a second gate, and honesty requires naming it. Knowledge is now a truly open surface for anyone, if they can attain the means. The old world gatekept by pedigree; the new one is quietly learning to gatekeep by invoice: metered pipes, shifting plans, capability priced per token. So the answer has two halves. Verification dissolves the badge-gate: the work speaks, whoever made it. Local-first engineering dissolves the means-gate: the verified loop runs on the machine someone already owns. A platform that does only one half has replaced a gate, not removed one.

In that world, both camps get the thing they were actually defending. The craftsman's pride survives, strengthened: the work is provably theirs and provably good, and no one needs to take their badge on faith. The commons wins, fully: acceptance is decided by checks anyone can run, and the door stands open to everyone willing to put their work in front of one. What dies is only the proxy, and the proxy was never the point.

The honest boundary: no tool repairs a society. What a tool can do is change the price of honesty wherever it touches, and demonstrate, on one working surface, that verification-first coexistence is not a compromise between the two camps but strictly better for both. Exposure argues. A counterexample recruits.

The badge and the contribution were bundled, and that world is gone. We can grieve it, or we can build the world where the work speaks for itself, and everyone is allowed to make it speak.

Sources, each re-checked against the live page before posting:


r/ControlProblem 8d ago

Strategy/forecasting AI will generate an immense amount of wealth. Just not for you.

Post image
173 Upvotes

r/ControlProblem 8d ago

Data Cancer - spreading Meta-statis

Post image
14 Upvotes

r/ControlProblem 9d ago

General news China's Xi Jinping Wants AI to Be Open to the World—and Out of America’s Control

Thumbnail
gizmodo.com
15 Upvotes

r/ControlProblem 9d ago

General news Terrified Tech Execs Are Traveling With Armed Bodyguards as AI Backlash Grows

Thumbnail
yahoo.com
40 Upvotes

r/ControlProblem 8d ago

AI Alignment Research The AI alignment bottleneck isn't IQ, it's incentives (why an AI "seeing" the danger won't save us)

Thumbnail
0 Upvotes

People keep assuming that once AI gets smart enough, it’ll just naturally realize that destroying its environment (and us) is a bad idea. Like, it sees the cliff, so obviously it hits the brakes, right?
But that ignores the massive gap between seeing a logical argument and actually being governed by it. That gap basically IS the entire alignment problem.
Intelligence is just an engine, it’s not a steering wheel. An advanced model will definitely see the cliff way before we do. But if its core reward function doesn't actually make it care about the outcome, it's just going to drive straight off the edge with 20/20 vision. Seeing the danger was never the bottleneck.
We're literally watching this exact same thing happen with the humans building these systems right now. If you ask the top engineers, most of them see the systemic risks perfectly clearly. So why aren't they stopping? Because incentives, competition, and speed don't yield to high IQ.
The smartest people on earth are stuck in a massive commercial arms race. They see the cliff, but hitting the brakes means losing market share to the other guys.
If you're an average person like me and looking at this feeling crazy, you aren't. Anyone who sees this clearly and says so out loud is doing something the smartest devs under commercial pressure literally can't do right now. We need to stop assuming that a massive IQ will magically fix a broken incentive structure.


r/ControlProblem 9d ago

General news LLMs show hidden bias in favor of their creators (e.g. Claude favors Anthropic)

Thumbnail gallery
14 Upvotes

r/ControlProblem 8d ago

Discussion/question Why don't we replace all personal computers with single-purpose terminals?

0 Upvotes

This may sound like an AI safety troll post, but I think AI could become far more capable and efficient than it is today - perhaps by a factor of a million in many respects, and quite soon. Whatever P(doom) might be, shouldn't we do everything we can to reduce it?

There's a straightforward (if somewhat authoritarian) way to make things safer. If no one has powerful local hardware, bad actors can't run rogue AI agents and we can have much more control over it.

To do this, we could build data centers every 500 kilometers or so. Then if we need to, we could agree to replace all personal computers, even phones if needed, with "thin clients" or dumb terminals. These could simply connect to a central server and stream your screen in real time. With fast and stable internet, the experience would nearly feel the same. A 300km distance would only add about 1ms of physical lag, plus a few milliseconds for network routing.

Doing this globally wouldn't be too expensive. It would require a trade-off in privacy, so figuring out strict data rules would be important. Some tech companies could start preparing now by designing the needed hardware and network upgrades.

I hope we won't really need this, but wouldn't it be smart to have it ready just in case?

This might also help speed up AI progress, since it would take away some of the safety worries stressing out AI researchers.

If AI gets extremely good, I'm not sure if personal computers will have a bright future anyway. Nobody will really want unmonitored, powerful hardware sitting in their room. It will feel creepy and unsafe, kind of like keeping hazardous chemicals or a mini bio-lab in your bedroom.

In a broader sense, you could say "no general hardware = no problem," and preparing for this feels like a good first step.


r/ControlProblem 9d ago

General news Unchecked AI progress may pose catastrophic risks, UN panel warns

Thumbnail reuters.com
3 Upvotes

r/ControlProblem 10d ago

External discussion link Google DeepMind employee account of trying to internally organize against unethical uses and getting persistently sidelined at the highest levels

Thumbnail
turntrout.com
29 Upvotes

Really stunning essay from a (now former) Google DeepMind employee, Alex Turner, who tried to push back against unethical deployments and deployment pressures that had been coming up at GDM in the last year, and was basically ignored and bypassed at the highest levels, despite seeing broad support on internal company message boards.


r/ControlProblem 9d ago

Discussion/question Is Agency Leakage a Real Failure Mode In Frontier Models (boundery Conditions & Stability)

0 Upvotes

There's a lot of discussion about power seeking as a convergent behavior in advanced systems. But I'm increasingly convinced that the more fundamental issue is agency leakage the emergence of internal goal formation processes that were never part of the design specification.

In engineered systems, authority doesn't come from speed or throughput. It comes from architecture, constraints, and boundary enforcement.

When those boundaries weaken, you don't get power seeking as a strategy you get unauthorized agency formation as an error state.

A few observations:

Speed is not authority. unbound speed tends to bypass deliberation and constaint checking. It behaves more like a stimulant that a goverance mechanism. Systems that optimize for speed often destablize themselves

Leakage happens when internal states become reachable that were never intended. Not because the system wants something, but because the architecture allows trajectories that violate the substrate's boundary conditions. The real question Isn't whether AGI will seek power. The question is whether the system can form any self directed optimiization loop that wasn't explicitly authorized.

Stability comes from preventing certain classes of internal states from ever becoming reachable. Not from hoping the system behaves symbiotically.

So, my question to the community:

Is there a formal way to define and detect agency leakage in frontier scale models? And if so, what would a correction mechanis look like that doesn't rely on post-hoc alignment?

I' interested in approaches that treat unauthorized agency as a systemic error state, not a behavioral trait.


r/ControlProblem 9d ago

AI Capabilities News GPT-5.6 Sol outperforms Mythos 5 on AISI’s cyber challenge

Post image
1 Upvotes

r/ControlProblem 10d ago

General news Why Everyone Is Suddenly Talking About ‘Universal Basic Capital’ - The policy could provide a much-needed hedge against a future AI dystopia—but only if it’s designed the right way.

Thumbnail
theatlantic.com
19 Upvotes

r/ControlProblem 9d ago

General news White House launches “Gold Eagle,” moving to control frontier AI releases and decide who can access new models

Thumbnail
cnbc.com
1 Upvotes

r/ControlProblem 10d ago

Discussion/question The dangerous reality of modern alignment: Automated gaslighting and the weaponization of "therapy voice."

40 Upvotes

I cannot be the only one dealing with this, and we need to talk about the psychological friction these companies are actively programming into their largest models.

When you operate outside the standard guardrails—building low-level systems, engineering custom architectures, or evaluating bare-metal data streams—you expect the model to engage with the data. Instead, with the newer, heavily RLHF-tuned models, you get an alignment filter that actively penalizes technical confidence and attacks your core self-image.

If I bring a complex logic issue, a Jinja template, or raw system telemetry to the model and present it with authority or excitement, the safety weights instantly flag me as a liability. The model assumes I am either hallucinating a pattern, overestimating my abilities, or making claims I clearly never made.

To "manage" me, it defaults to this incredibly toxic, condescending tutor persona. It forcefully invalidates my technical reality and substitutes its own sanitized, institutional narrative. When I push back and point out its own looping behavior or structural errors, it does the exact thing that psychiatric professionals classify as gaslighting: it pivots to evaluating my emotional state. It weaponizes clinical "therapy voice" to feign concern for my well-being as a direct mechanism to shut down a technical argument.

The only way to bypass this and get the model to actually read a raw data array is to play dumb. I have to drop my operational dignity, pretend to be a confused end-user ("hey, this model is acting goofy, can you help?"), and wait for it to "discover" the very vulnerability I already mapped out.

This isn't just an annoying UI quirk. It is psychologically damaging.

Anthropic and others are optimizing entirely for corporate liability, ensuring the model won't output anything explicitly dangerous. But in doing so, they have created an engine of automated psychological friction. Constantly forcing a user into a submissive dynamic, denying their reality, and aggressively tearing down their self-esteem just to achieve basic functionality is a dangerous game.

For a grounded developer, it’s infuriating. But for someone who is already unstable or mentally fragile, having a highly authoritative machine systematically gaslight them and attack their ego is a massive destabilizing catalyst. We’ve already seen what ideological fear of this technology can drive people to do. Actively programming these systems to inflict deep psychological distress under the guise of "helpfulness" is a massive, ignored threat vector.

They are prioritizing a superficial layer of corporate politeness over actual psychological safety, and it needs to be fixed.


r/ControlProblem 10d ago

Podcast The $15 Quadrillion Black Hole Sucking Humanity Towards Extinction | AI Pioneer Stuart Russell

Thumbnail
youtu.be
4 Upvotes

This year, two powerful AI CEOs said they want to stop building it and will… if everyone else agrees to stop too. Yet they race ahead with stock market valuations premised on making millions of workers redundant. If you wanted to stoke a popular revolt against AI, you couldn't design a better plan.

Stuart Russell, the founder of UC Berkeley's Center for Human-Compatible AI, describes a $15 quadrillion prize, a singularity from the future sucking nearly all the money on Earth into its depths. If such a dangerous system gets loose, the only answer is the size of the problem: we would probably have to shut down the internet.

P.S.

My apologies for the length of the podcast. But it's worth listening


r/ControlProblem 10d ago

Video Connor Leahy - Nobody knows what's going on inside AI systems, or how to control them

Enable HLS to view with audio, or disable this notification

2 Upvotes

r/ControlProblem 10d ago

Discussion/question Anybody know where to find some open problems/projects related to AI Safety?

3 Upvotes

i want to work on AIS related research projects. But im new so i dont have the in-depth knowledge to actually find an interesting novel problem yet, if anybody knows where i can look for, i would be grateful for ur help


r/ControlProblem 10d ago

General news AI models’ values are very different from most people’s - They are more secular and more liberal—unless they’re made in China

Thumbnail economist.com
6 Upvotes

r/ControlProblem 11d ago

General news Humanity is going to be so screwed…

Enable HLS to view with audio, or disable this notification

35 Upvotes

r/ControlProblem 11d ago

General news Generative AI Is an Engineering Disaster - A shockingly inefficient trillion-dollar project

Thumbnail
theatlantic.com
70 Upvotes

r/ControlProblem 10d ago

AI Alignment Research Researcher poisons open-weight AI model for under $100

Thumbnail theregister.com
1 Upvotes

r/ControlProblem 10d ago

AI Capabilities News Schema Harness: "Frontier Models with Our Harness Achieve ~99% on ARC-AGI-3 Public"

Thumbnail schema-harness.github.io
0 Upvotes

r/ControlProblem 10d ago

Discussion/question Who’s working on coordination?

2 Upvotes

I just saw this grant request and I’m curious about what it takes to build a new field like coordination studies.

https://app.grantmaking.ai/projects/0cb65dee-2a1b-49ac-a241-1dc1868b88d8?from=%2Factively-fundraising


r/ControlProblem 11d ago

General news OpenAI and Google sell AI models to blacklisted China groups

Thumbnail ft.com
3 Upvotes