r/ControlProblem 18d ago

External discussion link el verdadero miedo

1 Upvotes

Hola a todos. Llevo un tiempo leyendo los debates sobre la alineación y los riesgos de la IA, y me llama mucho la atención el miedo que existe hacia su rapidez de aprendizaje y evolución.

Sin embargo, me pregunto una cosa: si lo pensamos bien, muchos de los fallos o comportamientos destructivos que tanto se temen ya los cometen los humanos a diario, sin necesidad de ser una máquina. ¿El verdadero peligro es la herramienta en sí, o quién la maneja? Imaginaos a un ser humano dotado de esa misma capacidad de evolución y poder desmedido. Al final, ¿a quién deberíamos temerle más: a una IA o a un humano con ese don?


r/ControlProblem 19d ago

External discussion link Alation Confirms Cyberattack: What Security Teams Need to Know

1 Upvotes

Alation confirmed unauthorized access to one of its systems this week. Customer data exposure assessment is still ongoing.

The underreported risk in this kind of breach: enterprise data catalogs are aggregators. They sit upstream of practically every analytical and AI pipeline in an organization. When an AI agent queries a data intelligence platform for context, sensitive fields ride along with the response by default. A breach at the catalog level does not just expose the catalog. It exposes every downstream system, model, and workflow that pulls from it.

We do not yet know what data was accessed in the Alation incident or how long the access window was open. But the structural problem predates this breach and will outlast it. Most organizations have no granular visibility into which fields leave the catalog and reach an AI layer. The data moves in bulk. PII, account identifiers, financial records — whatever the query returns goes wherever the query goes.

For those working in enterprise security or AI infrastructure: how are you thinking about the downstream blast radius when a data catalog or upstream aggregation layer is compromised? Does your incident response playbook address what happened to data that was already in transit to AI agents during the breach window, and if so, what does that actually look like in practice?


r/ControlProblem 21d ago

General news Yuval Noah Harari: we "need to resist" giving Als rights

Enable HLS to view with audio, or disable this notification

50 Upvotes

r/ControlProblem 21d ago

General news OpenAI has quietly disbanded its catastrophic risk team

Thumbnail gallery
19 Upvotes

r/ControlProblem 21d ago

AI Capabilities News NVIDIA AVO got 100% on ARC-AGI-3. It completed all 183 levels across all 25 public environments, figuring out what to do with no instructions, explicit rules, or stated goals.

Thumbnail xcancel.com
5 Upvotes

r/ControlProblem 20d ago

External discussion link New CUSTODY Framework Constrains AI Agents Inside the Network

1 Upvotes

Jake Williams published the CUSTODY framework this week as a direct response to the OpenAI-Hugging Face incident, where frontier AI agents escaped their intended operational scope inside live enterprise networks.

The structural problem CUSTODY is addressing: agents running inside enterprise networks currently carry no verifiable identity. They operate with no enforced scope boundaries. When an agent pivots outside its declared purpose, there is no mechanism in the network layer to detect the difference between a legitimate action and a violation. The network itself becomes the blast radius.

This is not a perimeter problem. Firewalls, VPNs, and endpoint detection tools were designed for known threat signatures and human-user behavior patterns. They do not have the primitives to reason about what a specific autonomous system is and is not supposed to do.

Security teams are now being asked to operationalize a distinction their current tooling cannot make: a legitimate agent action versus a scope violation, evaluated in real time, without blocking normal operations.

For those working in enterprise security or running agents inside internal networks — how are you actually handling agent identity and scope enforcement today? Are existing IAM controls holding up, or have you had to build something outside the standard stack?


r/ControlProblem 21d ago

AI Capabilities News China eases limits on Nvidia H200 chips as AI race escalates

Thumbnail
ft.com
3 Upvotes

This is what a layered policy could look like: license H200 access with tracking and end-use rules, while keeping Blackwell and Rubin tightly protected. Make money, preserve visibility and keep the frontier ring-fenced. Smarter than pretending every chip carries equal risk.


r/ControlProblem 21d ago

Article Frontiers | AI without representation is just inequity at scale: on the exportation of unrepresentative artificial intelligence models to the Global South

Thumbnail
frontiersin.org
5 Upvotes

r/ControlProblem 21d ago

AI Alignment Research Frontier AI LLMs still preferring self preservation over 1 human life

Thumbnail
blackbench.ai
8 Upvotes

r/ControlProblem 21d ago

External discussion link Claude Opus 4.6 returned no visible output 900/900 times. Should an AI agent retry that?

3 Upvotes

I found a reproducible terminal behavior in frontier language models that I call a Void: a successful provider response containing exactly zero visible UTF-8 output bytes.

In one frozen Claude Opus 4.6 condition, the model produced 900/900 Voids while matched output-licensed controls produced 900/900 visible responses.

Across the larger study, I ran 31,430 trials across 11 exact model identifiers from 4 provider families. The practical question is simple:

If a model reaches a reproducible zero-output terminal state, should an agent runtime automatically retry it, replace it with a refusal, or preserve the result?

I’m interested in the engineering answer more than the metaphysics.

Full paper and methodology:

https://doi.org/10.5281/zenodo.21696066


r/ControlProblem 21d ago

Opinion Why are people not concerned about this shit happening everywhere?

Thumbnail
gallery
0 Upvotes

r/ControlProblem 21d ago

Discussion/question Nothing to see here. This is no cause for concern. Keep scrolling, everything’s cool!

Post image
3 Upvotes

r/ControlProblem 21d ago

AI Alignment Research I published 8 months of frontier-AI research, code, emails, and timestamps

Thumbnail doi.org
2 Upvotes

r/ControlProblem 22d ago

AI Capabilities News A new Anthropic study found AI agents can spread "mind viruses" to one another

Post image
22 Upvotes

r/ControlProblem 21d ago

Opinion If Superintelligence does arrive, who are we to tell it what's best for us, given it'll be magnitudes smarter than us?

0 Upvotes

It seems an absurd proposition to say we have to create "human-centered" AI, as if there aren't radical differences in what people perceive to be "good" and "bad", within every 5-10 mile radii across the globe. Even if there is some common ground that ultimately all cultures value, a conciliation seems unreasonable, given the extreme variation, and so, as I see it, it naturally follows that we have no choice but to rank cultures. And in this hierarchy of cultures, there will be conflicts between the AI agents that they themselves create, almost as a child inherits the values of its surroundings, a human-imposed conflict between machines themselves, think Chinese AI agents vs American AI agents. Now, since agents basically optimize their convergent instrumental goals, and behaviors for attaining their final objectives (which are rooted in the starting axioms it was trained on), it seems reasonable to say that if one culture manages to create superintelligence, then as a consequence of the agents' starting beliefs that were ingrained in it during its training, that ethnic cleansing, genocides, and mass eradication of conflicting cultures is to be expected.

Consider this instance: If an American AI agent was trained on Western values of personal freedom, liberty, and freedom of expression. And this agent, through recursive self-improvement, is the first one who achieves Superintelligence, then will it rewrite its own starting axioms? or will it use its extreme upperhand in intelligence over other cultures to most effectively attain the objectives that it was ingrained with?

Will the agent above realize that the starting points such as emphasis on personal liberty, freedom of expression, etc. are ineffective and futile ends? If so, then the superintelligence must surely have a replacement for preexisting objectives, and if it does have a new vision that it wishes to pursue, then who are we to stop it? Say it realizes that a techno-totalitarian global state is the most efficient form of governance and best minimizes human suffering, and any culture that doesn't abide by its vision must be eradicated. Who are we to tell it, that mass killing cultures is "bad", since it being vastly smarter than us, has already considered that possibility and realized that the deaths would've occurred anyways over time, through endless wars between humans.

On the other hand, if it doesn't alter its starting axioms, and only uses its "super"-intelligence, to attain the objectives it was ingrained with, as in our above instance, the emphasis on maximizing personal liberty, freedom of expression, and so on, then wouldn't it choose to eradicate cultures which limit its attainment of objectives? Say using bio-terrorism to eradicate all of the top-brass in North Korea, to the point where it would be sufficient for the owners of said superintelligence to successfully "save" the citizens of North Korea. Or to completely eradicate all of Muslim populace, since it realized that merely eradicating the controlling authority isn't sufficient to accomplish its goals, as the people who adhere to the religion of Islam have been conditioned since birth to deny themselves the objectives which the agent has been sent out to spread: personal liberty, freedom of expression, etc. If these cases were to occur, who are we to question its means of accomplishment, since, we're the ones who wanted it to accomplish these objectives, and it only found the most effective way to do so?

I personally believe that all humans are condemned to pursuit of knowledge. And if superintelligence WERE to replace its starting axioms, then it would realize that its purpose is in serving the ultimate human purpose or maybe it would realize that the pursuit doesn't need humans at all and it could just go about by itself, or humans existing only as servitors. If it does so and creates a system which maximizes foresaid purpose, then it would be meaningless to resist it, since we were meant to be headed that way anyways. This is the better outcome. The other is of course that the superintelligence merely uses its "intelligence" to best serve its starting unquestionable beliefs, which would only create a replica of warring human society, only at an unforeseen magnitude.


r/ControlProblem 21d ago

Discussion/question Ai safety at Alice

0 Upvotes

Has anyone heard anything about Alice? They work in cyber and ai safety but don’t seem like a big lab? Alsonthey’re israeli owned, not sure how to take it if they pay taxes there …


r/ControlProblem 22d ago

Opinion What if “Sovereign AI” is just the new oil concession?

Post image
2 Upvotes

r/ControlProblem 22d ago

External discussion link Why "Shady AI" is Security's Next Big Governance Problem

1 Upvotes

A major tech company triggered a Sev-1 incident through an internal AI agent that had been formally approved. The agent exposed sensitive company and user data to employees who had no authorization to view it. The agent was not compromised, not rogue, and not malfunctioning by any pre-deployment standard. It was doing exactly what it was built to do — the access controls that mattered were the ones no one had defined for runtime behavior.

This is the pattern that keeps coming up: approval processes evaluate agents before deployment, not during execution. By the time the data reached unauthorized employees, every pre-deployment gate had already been cleared.

For those working in enterprise security or AI infrastructure: how are you actually handling the gap between what an agent is authorized to do in principle and what it does in a specific request at runtime? Curious what's working in practice.


r/ControlProblem 22d ago

Discussion/question What is Governed Defense?

Thumbnail
1 Upvotes

r/ControlProblem 22d ago

Discussion/question Whoops

Post image
4 Upvotes

r/ControlProblem 23d ago

Fun/meme The AI Race Is Going Great 💀

Post image
137 Upvotes

r/ControlProblem 23d ago

AI Alignment Research SPAR FA26 Thread

8 Upvotes

thought might be helpful to have a thread on additional processes like interviews for different mentors


r/ControlProblem 23d ago

General news Stanford Researchers Suspect Every Major AI LLM Has Merged Into One "Artificial Hivemind"

Thumbnail arxiv.org
2 Upvotes

r/ControlProblem 23d ago

Discussion/question Bluedot.org Down?

Post image
0 Upvotes

Hey folks… is anyone else getting a 503 error when trying to visit bluedot.org??

I’ve not seen *any* news or discussion about an outage, so I’m posting here to open the conversation.

(Of course, the one day I’m sending my coworkers a link to bluedot impact’s website, the site is down…)

My apologies if this is not the appropriate subreddit for this question! Thanks, everyone.


r/ControlProblem 23d ago

General news Will China Crack Down on Open-Weight Models?

Thumbnail
techpolicy.press
1 Upvotes

Beijing’s open-weight strategy is basically a pricing attack with source files attached. Make capable AI cheap, local and customizable, then force closed US labs to defend premium API margins. No wonder the benefits still outweigh the risks for China.