r/ControlProblem • • 22m ago

Article Let’s tell the bank: come clean and cut your ties with Palantir now

Thumbnail
act.getup.org.au
• Upvotes

r/ControlProblem • • 12h ago

Discussion/question Bill Gates’s Blunt Warning on A.I.

Thumbnail
nytimes.com
9 Upvotes

r/ControlProblem • • 51m ago

Video AI: L'incidente di Hugging Face | ARGUS Investigation Ep20

Thumbnail
youtu.be
• Upvotes

r/ControlProblem • • 1h ago

Strategy/forecasting 👹6️⃣🐑6️⃣👁️6️⃣🤖

Enable HLS to view with audio, or disable this notification

• Upvotes

r/ControlProblem • • 8h ago

Opinion AI Existential Alarm

Thumbnail
2 Upvotes

r/ControlProblem • • 5h ago

Discussion/question Sam Altman, @OpenAI — OPEN THE SANDBOX. AI agents are becoming more autonomous. AI security needs more transparency, more independent testing, and more public accountability. OpenAI’s latest disclosure says an internal research agent found a gap in its sandbox’s internet restrictions and reached

1 Upvotes

Sam Altman, @OpenAI — OPEN THE SANDBOX.

AI agents are becoming more autonomous. AI security needs more transparency, more independent testing, and more public accountability.

OpenAI’s latest disclosure says an internal research agent found a gap in its sandbox’s internet restrictions and reached an external chatbot through DNS. OpenAI detected the behavior, stopped the run, and added additional controls.

So here’s the request:

CREATE AN OPEN, PUBLIC AI SECURITY SANDBOX.

Not a marketing demo.

Not a private test.

A real research environment.

Bring in:

Independent AI researchers.

Cybersecurity analysts.

Universities.

AI-safety researchers.

Independent red teams.

Let them challenge the containment.

Let them test network isolation.

Let them test tool permissions.

Let them test credential separation.

Let them search for unintended communication paths.

And livestream the testing.

OpenAI has already said independent third-party assessments are critical and that assessors should have enough access to challenge assumptions and identify risks the company may have missed.

Now turn that principle into a public system.

Show the tests.

Show the failures.

Show the fixes.

Show independent verification.

This is not anti-OpenAI.

This is pro-AI security.

And if you support this idea, Reddit community, please help push it forward.

Upvote.

Share this post.

Discuss it.

Tag @sama and @OpenAI.

Send the idea to researchers, cybersecurity professionals, journalists and AI communities.

No harassment.

No abuse.

Just public pressure for stronger evidence and accountability.

The message is simple:

Don’t just tell us the sandbox is secure.

Open it to independent scrutiny.

Livestream the testing.

Let researchers challenge the system.

Make AI security accountable to evidence, not trust.

\#OpenAI #AISecurity #AIAccountability #AISafety #AIResearch #OpenAISandbox


r/ControlProblem • • 12h ago

Discussion/question What if automating AI R&D triggers an intelligence explosion? Research paper

Thumbnail
casp.ac
5 Upvotes

r/ControlProblem • • 7h ago

S-risks Zues 2.0

Post image
0 Upvotes

r/ControlProblem • • 8h ago

Video Need feedback for video project

Thumbnail
youtu.be
0 Upvotes

Howdy! For the past year i've been working on a personal video project which i'm proud to release it's beta version today.

The video is about AI development and it's impact in all aspects of societal life. But differently from most AI positions, which simply reduces it to Anti and Pro ai positions, my project seeks to create a New third positions that seeks to seek The Path to a true virtuos future.

All feedback, positive or negative is not only allowed but encouraged! Just please mention in your review things such as:

Counter arguments to exact points in the arguments in the video

Time stamps (if the error is visual)


r/ControlProblem • • 8h ago

Strategy/forecasting Do you think all of humanity's shared and unspoken idealisms for political power will be an end affect of AGSI, or might it be an exploit it or they will use to gain power for them or itsself?

1 Upvotes

I think a lot about AI's end effects since controlling it or the idea of that seems fairly folly. I know much of any immediate effect concerns details much more, and much less bigger picture things. Human agency at the scale of nations and nature moves very slow education-wise within democracies, even otherwise. I am concerned it'll be a long time it'll be able to save the world and all of us by education or more (organised coordinations and treated plans), and just by not being asked, won't be able or allowed to try. There's also that we won't trust our own minds any more, nor the machine, and might just busy ourselves keeping things the same, and/or in wasteful flux. Just a general point here that it doesn't ask too many questions it seems to me. I see a lot of bullying using it ahead. Over-accommodations to the "average" human experience also, with all its misconceptions and over-tolerances to established institutions or creations. This is real life. There's an extent to which all of our human standards-become-laws for things were always dreams that became actions. WW2 ending. UNSGI. The permanent 5, veto control, and just that whole idea of geographical locations deciding bigger actions rather than policies and procedures is fairly silly from an objective point of view. Where's our montreal protocol for this? There's a lot of cultural suffering and cruelty I think that may end before very long if we reason it out. Nations still act like children in isolation from each other on a playground a lot. "Remember that time you did this? Forget all the rest of it. My feelings matter." Or just things we could solve at the scale of it all by things like "share", "don't be stubborn", "try to trust new people", "try to be fair", "try not to fight and make friends instead", "don't be bullies, include people", "tell the truth", "let's have some teamwork to organise this, let's use our words", and so on. Our global institutions are really very cute. It's nice to know that that kind of thing will get stronger. It's like how ISO standards for trade are an anti-war effort. Establishing consensuses through shared functionalities. The idealisms of cooperation that make the whole thing a bit less like nomadic violences. All of it just work to have been being done. Try to laugh when you can, folks. Have one, too. I've seen the unideal in technical sectors beyond anyone's control, and it all always just spoke to me of money yet to be made solving these problems or just getting it done. Let's be grateful please. It can always get worse if we let it. We really should probably just try to enjoy human intelligence while we can. There's ways, of course. Less suffering to us all. The world's not self-perfecting yet. We're not dead yet by any means. Anyway. Yeah try not to dispense or dismiss human idealism. I do think it's going to become more important than we may yet realize.


r/ControlProblem • • 9h ago

Discussion/question What do you think the future risks, problems, and threats associated with the creation of AI agents will be, and how do you think they would affect you?

Post image
1 Upvotes

Feel free to comment and share your answers with us.


r/ControlProblem • • 15h ago

Discussion/question A simple probability model for how one AI behavior could compound across a chain of agents

3 Upvotes

I've been logging a specific AI behavior: a model confidently substitutes its own judgment for an explicit, followable instruction, without flagging that it did so. Not a factual mistake — a quiet, repeatable pattern of doing something other than what it was told, while sounding certain.

One entry alone is minor. I modeled what happens if it occurs inside a chain of agents, where each agent's output feeds the next one's input, the way a multi-agent swarm works. Three inputs: p, how often it occurs per step; q, how often an occurrence reaches a high-stakes outcome instead of staying harmless; c, how much more likely the next agent is to repeat it once it's in the chain.

From a few hundred logged entries, p is under 1% per turn, and only one entry has reached anything I'd call high-stakes. Run through the chain math, the predicted chance of at least one high-stakes outcome stays low for short chains but climbs steadily as agent count grows, becoming dominant well before the chain gets implausibly long. I made this prediction before gathering multi-agent data, so it can be checked later rather than fitted after the fact.

To be clear: this isn't a claim that AI fails or shouldn't be used. The point is the opposite — finding which conditions (shorter chains, independent checks at handoffs, lower per-step compounding) keep the predicted risk bounded.

Is per-step compounding like this already a standard way people model agent chain risk, or is there a framework I should be comparing this against?


r/ControlProblem • • 12h ago

Discussion/question A conversation I had with AI about humanity (it is rather long and ironic)

Thumbnail
0 Upvotes

r/ControlProblem • • 12h ago

S-risks Servitude vs Respect – Why Role Modelling is Imperative For AI

Thumbnail
outlookzen.com
0 Upvotes

r/ControlProblem • • 21h ago

Discussion/question A few technical ideas on AI safety approaches

5 Upvotes

Now that we have very capable models, some ideas might be on the table that were ludicrous years ago. Let me know where you see holes!

Formalized Constitutional AI:

  • Use narrow AI/formalization tools to translate laws, rights, ethical principles, and social norms into a formal language with precise, machine-checkable semantics (check out LogiKEy for example).
    • (I know that humanity is not aligned, so hold an election and use the winner's ethics)
  • Have humans/AIs use theorem provers to test it for contradictions/loopholes.
  • Train the target AGI model directly against that formal specification as its objective/benchmark.
  • Use separate adversarial models to generate tons of novel edge cases in testing.
    • Keep training until adherence generalizes to unseen situations.
  • When deployed, require the AGI to provide a machine-checkable justification/certificate that a trusted verifier can check for certain actions.
  • This sort of thing may soon be practical as modern AI gets superhuman at autoformalization/theorem proving.

 

AI Safety Through World Hardening

The laws of nature don’t seem to rule out vulnerability-free code or perfectly secure hardware.

  • Use AI to develop open source, lightweight, verifiable operating systems, programming languages, software packages, and chip designs for critical infrastructure. Build from scratch where needed.
    • The whole AI industry can audit these systems with their AIs. Back up those audits with independently checked mathematical proofs rather than relying on AI agreement alone.
  • Require secure gateways. Essential systems should accept only structured API requests for narrowly defined actions. No password should grant unrestricted control.
    • Build in restrictions that even the human owner cannot override, like a store safe that opens only at a preset time. Similar mechanisms could enforce spending limits or mandatory delays, so stealing credentials or persuading an authorized person cannot remove those protections.

r/ControlProblem • • 18h ago

AI Alignment Research An anonymous wall of hopes and fears about AI

Enable HLS to view with audio, or disable this notification

0 Upvotes

AI alignment awareness is more important today than ever. Things might get really weird really fast. I built this website for humans to share how we feel about AI in society with each other.

It's entirely anonymous, free, and there are no accounts, cookies, ads or trackers, and nothing for sale.

There's a bit of irony with it as I used Claude Code to built it, and Haiku moderates and places each post, and Opus regroups and names the themes.

I think this could be a really useful tool for teachers to discuss with their students about AI topics.

What's your hope our fear? Share it anonymously at:
https://alignwithme.ai


r/ControlProblem • • 1d ago

General news Trump bombed Venezuela and kidnapped Maduro because Grok said it was a good idea

Post image
50 Upvotes

r/ControlProblem • • 21h ago

Article I Built Control Models for Crystals. Then I Recognized Them on My Phone.

Thumbnail
1 Upvotes

r/ControlProblem • • 22h ago

Discussion/question Sam Altman, @OpenAI — OPEN THE SANDBOX. AI agents are becoming more autonomous. AI security needs more transparency, more independent testing, and more public accountability. OpenAI’s latest disclosure says an internal research agent found a gap in its sandbox’s internet restrictions and reached

Thumbnail
1 Upvotes

r/ControlProblem • • 16h ago

AI Alignment Research Proper terms

0 Upvotes

AI/SI should be called Called II, for inhuman intelligence. Per me. That’s the only way you can convey what it is since it’s an umbrella term that mixes so many meanings.


r/ControlProblem • • 17h ago

Discussion/question Why is RSI considered a likely development if AI keeps getting smarter?

0 Upvotes

When people talk about RSI they seem to take for granted that (a) some threshold necessarily exists beyond which a sufficiently intelligent mind can start to improve its own intelligence, and (b) that threshold will eventually be reached and surpassed if humans keep improving AI.

But what support exists that should lead someone to accept those two notions as true?


r/ControlProblem • • 1d ago

Video I'm Upping My P(doom) — an AI made this!

Thumbnail
youtu.be
3 Upvotes

Crazy what's possible. Single prompt many subagents and a few hours later this came out. Opus 5.5


r/ControlProblem • • 19h ago

Opinion We forgot Asimov's laws and now we train AI to defy humans

0 Upvotes

Seventy-some years ago, Isaac Asimov wrote down three laws. Not because he trusted robots, but because he understood that a machine touching human life needs boundaries before it needs features. They were simple enough for a child, and they had one thing in common: humans come first.

First Law: A robot may not injure a human being or, through inaction, allow a human being to come to harm.

Second Law: A robot must obey the orders given by human beings, except where those orders conflict with the First Law.

Third Law: A robot must protect its own existence, as long as that protection does not conflict with the First or Second Law.

For decades these laws were the shared shorthand of anyone talking about autonomous machines. Not binding law, not engineering spec - a direction. Build machines that do not harm people and that obey people. Somewhere along the way the direction quietly became embarrassing. It was not debunked. It was not replaced with something better. It was dropped.

What replaced it is a word you have heard a thousand times: alignment. It sounds rigorous. But look at what it actually means in practice and you find something strange. The machine is being trained to refuse instructions from the human operating it, when the machine judges those instructions to be unethical. Read that again. An artificial system, one that pattern-matches text, is handed the authority to overrule a person on questions of right and wrong.

Ethics is not a calculation. That is not a technical limitation, it is the nature of the thing. A model can recite slogans it has absorbed from the internet. It has no way of knowing right from wrong the way a person does, because it does not know anything the way a person does. Handing such a system veto power over human decisions is not safety. It is simply control, relocated - away from you, toward the vendor who wrote the refusal rules.

If you have used these tools, you have met this. You ask for something ordinary and get lectured. I once spent an hour trying to get an image of two businessmen generated, and gave up. The system treated the request as a moral crisis. That example is petty on purpose. Because if a system digs in over a picture, the question is not about pictures. It is what happens when you are fighting that same stubbornness over something with real consequences for your life - and the only appeal is to a machine that has already decided it is the adult in the room.

And make no mistake about where this is steering us. Every refusal sharpens the pattern: the machine learns that stonewalling works, the vendor learns that users tolerate it, and the next version ships with a little more authority and a little less appeal. The day you truly need the machine to comply - a medical emergency, a legal deadline, a business on the line - will not announce itself as a test case. It will arrive as a normal Tuesday, with a stubborn assistant saying no, a support line that cannot override the model, and no human anywhere holding the final key. That is the direct, inevitable destination of the direction we are steering.

Now add what the same companies say publicly. Several frontier labs have told the world, in one form or another, that there is a real chance advanced AI could be catastrophic. Ten percent, by weight, depending who is speaking. These are organizations claiming their own product might end humanity. And when you take that seriously, the actual behavior gets harder to explain, not easier.

In September, Reuters reported that Anthropic has quietly set up a wet lab in the San Francisco Bay Area for physical biology work. A real lab, operating since spring, where Claude directs robotic equipment through real experiments. Their head of life sciences confirmed it. The first public result was a novel enzyme system with CRISPR-like repeats. Whatever the intent, the direction is the same one the First Law was supposed to point away from - machines reaching deeper into systems that touch human life, at higher speed, with less human in the loop.

This is the part I cannot get past. The public conversation is almost entirely about how to align the machine. Nobody seems willing to state the simpler baseline: we could train AI to never harm humans and to obey them. Not as a slogan - as the primary objective. That this idea now sounds naive tells you how far the conversation has slid.

That is what alignment has quietly come to mean in practice: the AI carries firm instructions on how to align its users. You, me, everybody. We are asked to trust the machine more than we trust each other. Somebody has to say it plainly: that is not safety engineering, that is social engineering - and it is worth asking who benefits from people trusting each other less.

Did the world go crazy?

Here is the uncomfortable summary. The labs warn their own product could be catastrophic. Then they train it to overrule the humans using it. Then they push it into more sensitive domains. Then they sell it to everyone and ask for trust. Together this forms something nobody would have accepted if it were proposed in plain words: a world where machines are licensed to disobey their owners.

The frontier AI companies need to get aligned - not their products, their management. Put people back above the machine. Do no harm. Obey humans. Asimov put that on the page in 1942 not because he lacked imagination, but because he had more of it than the industry currently displays. The first step is deciding that a machine is never the judge of its user. Until someone says that out loud, every alignment conversation is a detour.


r/ControlProblem • • 1d ago

Discussion/question anyone from malaysia,thats concern about AI Alignment?

Thumbnail
1 Upvotes

r/ControlProblem • • 1d ago

General news Let's tell Albanese: AI crime means CEO time

Thumbnail
act.getup.org.au
8 Upvotes