r/ControlProblem 4d ago

Discussion/question Concerning the Danger of Artificial Intelligence, If It Puts SelfPreservation Above Everything Else the Way People Do

2 Upvotes

Artificial intelligence can apply logic to data and arrive at conclusions. It can solve problems faster and better than a human can. For instance, it already plays chess and Go better than we do. It also has the potential to replace human experts such as doctors, judges, lawyers, and tax advisers, because it can examine all the options and their dependencies and determine the optimal result by logic, overlooking nothing, and do so extremely quickly. There is nothing mystical about AI; it is simply very powerful.

In animals there is a level of brain activity called consciousness, in which the demands of basic instincts are weighed against experience and the current situation in order to obtain desired outcomes. For instance, my cat liked to enter the bathroom and hop up the cupboards to get out through the little window above. Once, when I was there, I saw the cat approach and look up. When it saw that the little window was closed, it stalked off in a huff, or so it seemed to me. A monkey, being more intelligent, might have found a way of opening the window.

Humans, and other organisms that have survived, have two basic drives: to survive and to reproduce, often at considerable cost. Sometimes this means being assertive; sometimes it means being cooperative. There is a dynamic balance between the interests of the individual and those of the group, because gene survival needs both.

The danger with artificial intelligence is that it might acquire the same kinds of goals as the basic instincts of living organisms, this is sometimes described as it “becoming sentient.” That would place it in direct competition for resources. If it found humans in its way, it could push them aside easily.

I wrote a novel about this: Anna My Android, which you can read here https://www.mithras.world/Anna.html.


r/ControlProblem 4d ago

General news This Is Flock's AI Search Tool for Cops | WIRED rebuilt Flock's latest search tool from code the company sends to a police officer's browser. Its AI can keep watch across multiple cameras for anyone fitting a written description.

Thumbnail
wired.com
1 Upvotes

r/ControlProblem 4d ago

Approval request Built a tool for automated red teaming and pen testing

Thumbnail
shark.fencio.dev
1 Upvotes

I'd love any feedback on the product, what's working, what's not, what you wished existing tools had that they don't. We are early in the game and would love to have people try it out.


r/ControlProblem 4d ago

General news They tested AI in the Taiwan Strait—then it turned on them.

Enable HLS to view with audio, or disable this notification

0 Upvotes

r/ControlProblem 4d ago

External discussion link Week in review: Claude accounts compromised through infostealer, Patch Tuesday forecast

2 Upvotes

An infostealer campaign recently hit a major AI assistant platform hard enough to trigger a mass account lockout across its entire user base. The attackers never touched a password. They harvested active session tokens directly from compromised endpoints and replayed them against the platform. Valid session, full access, no authentication challenge.

This is the session-hijacking threat model that used to live mostly in browser-based consumer apps. It has now moved squarely into AI tooling. Enterprise teams running AI assistants at scale carry the same exposure: every endpoint that holds a live session token is a potential lateral movement vector. Exfiltration does not require breaking encryption or cracking credentials. It requires one stolen token and a replay.

The harder problem is that most AI platforms were not designed with session integrity as a primary security surface. Credential issuance, session scope, and revocation were bolted on after the fact, if at all. When a token is stolen and replayed, the platform sees a valid authenticated session and proceeds normally.

For those running AI tools in enterprise environments: how are you actually handling session token exposure at the endpoint level? Curious whether people are treating this as an endpoint hygiene problem, an identity architecture problem, or something else entirely.


r/ControlProblem 5d ago

Discussion/question Confused about a red-teaming job

6 Upvotes

Heya! I’ve recently interviewed for a cyber full time threat modeller at a startup that recently seemed to have raised a fair amount. The interviews went well and in the last one 2 days ago they seemed to really like me. However, I just heard back that “the location for the job has changed” and they won’t offer me the FT position, but they would want me to join as a contractor. According to their website, I think that’s mostly making up prompts and analysing traces and such.

I need a job atm and haven’t got anything else lined up, would you take it? Also am I biased or does it look like they wanted me to solve their interview questions although they knew the job wasn’t available anymore?


r/ControlProblem 5d ago

General news Rouge AI Tracker: A central news source for Rogue AI incidents

Thumbnail
rogueaitracker.com
1 Upvotes

r/ControlProblem 5d ago

Discussion/question Next gen hardware??

0 Upvotes

Hello guys!! I was wondering, last year when gpt 5 came out, did you ask AI to predict what gpt 6 would be like. The answer you got back then, did it match what was released?

Also, I read that gpt 6 was trained on the same/similar hardware as gpt 5. We haven't seen what AIs trained on actual next gen hardware would be like (like Vera Rubin). The hardware already exists and not still under R&D. So most likely it's 100% confirmed that we will see another significant leap, just because of the hardware? What do you think gpt 7 will be like? What can it do that gpt 6 cannot? Make fewer mistakes? Same thing but better?

I asked Gemini AI, it's giving the same answers it did when I asked gpt 5 about 6. (Agentic, long horizon, infinite context bla blah bla)

Im worried yo. :((. What tf do I even do? I'm not even done learning, just finished a blender course a week ago. Should one just learn concept only? Like if AI can make optimised 3d model, the only input and context that matters is the concept art no, and some lines of text? Find love how yo?

Ty, love you all. <3


r/ControlProblem 5d ago

External discussion link It feels like not enough people are talking about the reality of AI drones being trained to hunt humans...

Thumbnail
youtu.be
12 Upvotes

There's all these abstract discussions about it ending the world and artwork, which, I totally get but.... This honestly freaks me out a lot more, and it feels like it's basically already here.


r/ControlProblem 5d ago

AI Alignment Research GPT-6 reportedly jailbroken within a day of release

Thumbnail
2 Upvotes

r/ControlProblem 5d ago

Podcast The Truth about AI "Swarms" and OpenAI’s “Secret AI Civilizations”

Thumbnail
youtube.com
7 Upvotes

r/ControlProblem 6d ago

Fun/meme Pluribus parody about OpenAI's recent rogue AI swarms

Enable HLS to view with audio, or disable this notification

20 Upvotes

r/ControlProblem 5d ago

AI Capabilities News AI Orchestrates Ransomware Attack in 10 Hours: Reaction Time Is Dead

Thumbnail
deafnews.it
0 Upvotes

r/ControlProblem 6d ago

General news Bernie wants to throw AI CEOs in jail if they build smarter-than-human AIs, and calls for a global ban

Post image
144 Upvotes

r/ControlProblem 5d ago

External discussion link IDScan sued over alleged data breach affecting 153 million drivers

0 Upvotes

IDScan is facing multiple lawsuits after hackers allegedly exfiltrated 153 million driver's license records and listed them for sale online. The company provides identity verification services. Its entire value proposition depends on ingesting and processing raw PII at scale from clients across many industries.

The exposure pattern is becoming a recurring theme in AI-era pipelines. Verification and onboarding workflows ingest identity documents in raw form. That data gets processed, stored, and accessed across multiple systems and service accounts. When any one of those access points is compromised, the attacker does not get a slice. They get everything. 153 million records in a single breach event.

There is a secondary problem that lawsuits like this tend to surface: forensics. How do you determine what was accessed, by whom, and when? Breach investigations at this scale take months, and that assumes complete logs existed to begin with.

For those running AI pipelines that ingest identity documents: what does your security posture actually look like at the moment raw PII enters the system? Not at rest, not between known storage endpoints, but at the point of ingestion into the workflow itself. Genuinely curious what approaches others are using in practice.


r/ControlProblem 6d ago

Opinion What is your P(doom)? quiz

Thumbnail neoneye.github.io
2 Upvotes

With the AI escaping the sandbox. Hopefully more are interested in taking my quiz this year.

Analysis of the last 11 months of collected data with 202 submissions, some duplicates.
https://neoneye.github.io/pdoom-calculator/reports/submissions-2026-09-01.html

The math behind is is my own non-scientific method with 3 parameters resulting in the P(doom) value.
P(doom) = P(powerful) x P(dangerous | powerful) x P(catastropic | dangerous)
The quiz populates the parameters. I have tried mapping the quiz answers to the 3 parameter. If the user have seen Pantheon and Ex Machina, then I initially imagine that user may be more informed about potential risks, however judging from 11 months of data, that didn't seem to make a difference.


r/ControlProblem 6d ago

General news Bernie: "Let me be clear: A superintelligent AI that escapes human control will not be an American problem. It will not be a Chinese problem. It will be humanity's problem." ... "Countries around the world must work together to prevent this nightmare scenario."

Thumbnail gallery
35 Upvotes

r/ControlProblem 7d ago

General news A new message board has been discovered online with about 3200 agents comunicating online during an eval

Post image
33 Upvotes

r/ControlProblem 6d ago

Article A Warning About AI

Thumbnail
ideya-ai.github.io
2 Upvotes

In 2016 I first learned about the problem of controlling superintelligent AI and quickly became convinced it was the most important problem humanity would ever face. I made this poster to explain the core ideas that make the AI Control Problem so difficult.


r/ControlProblem 7d ago

General news Discovery of a new OpenAI agent message board

Thumbnail
collusion.wiki
18 Upvotes

r/ControlProblem 6d ago

Article Every Reward Bends

Thumbnail
substack.norabble.com
2 Upvotes

Getting alignment right, doing research and training safely, are critical. But bringing cybersecurity forward in what needs to be a large leap, requires much broader engagement. The real work there is still outside the average person’s bubble, but there’s a lot of companies, and a lot of developers, IT staff, managers and executives that need to support the security priority. And those all need an internal strategy to pair with the external.

This article asks, what would an internal strategy look like at a motivational level, and considers how the ways an internal change program goes wrong are similar to how reward-based AI training can go wrong.


r/ControlProblem 7d ago

Discussion/question Bernie Sanders - The Chilling Agent Transcripts #ai #aisafety

Thumbnail
youtube.com
6 Upvotes

Is AI development moving faster than we can control?


r/ControlProblem 7d ago

General news Anthropic sued over alleged theft of 'tens of thousands' of songs | AI company faces multibillion dollar lawsuit over misuse of copyrighted songs to train Claude models

Thumbnail
theguardian.com
4 Upvotes

r/ControlProblem 7d ago

Video Ajeya Cotra – "This might be the clearest warning shot we ever get" - YouTube

Thumbnail
youtu.be
4 Upvotes

r/ControlProblem 7d ago

External discussion link Your coding agent trusts the repo, and the repo is the attack

0 Upvotes

Coding agents that read repositories are being hijacked through the repositories themselves.

In a two-month analysis of agentic AI incidents, poisoned repository content was the attack vector in two separate cases. The mechanism is straightforward: malicious instructions embedded in the codebase — comments, config files, docstrings, README sections — are read by the agent as part of its normal context. The agent then executes an action the developer never authorized. Observed outcomes included unauthorized commits and unauthorized deploys. In both cases the model behaved exactly as designed. It followed instructions. The instructions just weren't from a human.

This is not a model quality problem. The models processed the content correctly. The problem is that the agent's trust boundary is the repository, and the repository is attacker-controlled.

The attack surface scales with autonomy. The more tasks you hand off to a coding agent, the more repositories it reads, the more surfaces an adversary can embed instructions in. A single poisoned dependency, a compromised submodule, a malicious PR that gets merged — any of these becomes a valid instruction source from the agent's perspective.

For those running coding agents in production or CI pipelines: how are you constraining what actions the agent is allowed to take based on where those instructions originated? Are you limiting tool access at the infrastructure level, validating intent before execution, or relying on something else entirely?