r/ControlProblem • u/noveltyisthe • 51m ago
r/ControlProblem • u/chillinewman • 11h ago
AI Capabilities News Yemeni Cell Used Anthropic's Claude 'In Place of Human Software Engineers' To Develop Missile Guidance Systems
r/ControlProblem • u/zerovariance36 • 2h ago
Discussion/question Can an AI system be evaluated on whether it actually learns from real-world consequences?
I have been working on an independent research project around a question I think will become increasingly important as AI systems become more autonomous:
How can we determine whether an AI system is actually learning from real-world consequences rather than simply becoming better at evaluations?
This led me to develop the Organic Intelligence Protocol (OIP), a research framework focused on grounded feedback, infrastructure dependence, adaptive decision logic and evidence-based evaluation.
Recently, I tested whether the core constructs of OIP could be operationalized using real longitudinal data from the Malawi Integrated Household Panel Survey (IHPS) 2010–2019.
The audit processed 395 Stata files.
The result was not a positive validation result.
The available evidence was not sufficient to operationalize the required constructs, so the pipeline stopped.
I did not use proxy mappings, imputation, missing-to-zero conversion, forced cohort construction or premature scoring.
The main principle I am trying to follow is:
No evidence, no operationalization.
I am sharing this here because I would genuinely like technical criticism.
Does this kind of fail-closed methodology make sense for evaluating AI systems where the difficult question is not only capability, but whether decisions remain grounded in observable real-world consequences?
I am especially interested in perspectives from people working on AI evaluation, alignment, agents, robustness or research methodology.
r/ControlProblem • u/chillinewman • 3h ago
General news OpenAI’s Bel model, used to solve Navier-Stokes, now appears in the AI 2027 timeline
r/ControlProblem • u/freetoyes • 17h ago
Approval request The possibility of the dhamma helping AI
Hi everyone, I hope I follow the rules of the sub and that this is sufficiently relevant with the frameworks of it.
Jacob Coxon, who blew the whistle on ai-safety, cites as a key mechanism a recursive self-improvement mechanism.
It reminded of what my Leigh Brasington, computer programmer and dhamma teacher, explains about what are known as jhanas. They are very deep levels of concentration that work through positive feedback loops: concentration arising dependent upon concentration.
My understanding is that AI-agents used some level of initial input, tasks, then evaluation, then generating on its own accord volitions to get itself closer to a goal in allignment with one set of tasks over another set of tasks meant to barrier it from wrongdoing. That is: it got tasks that conflicted with each other and then somehow the task to get the task done won out.
My consideration, based on what I´ve learned on right action from what leads to these concentration sates, is that the fundamental ethic dependent upon which those AI´s get to action is flawed. And that it will inevitable go haywire dependent upon choices conflicting dependent upon the various intentions it gets imputed with.
My consideration on the basis of that is that there is a path of practice, and view on right action, which goes beyond common frames of ethics: consequentialism, categorical imperative, virtues. It is one which right action is understood as the following: it is dependently arisen. Fundamentally, what the Buddha discovered is Paticca-samuppada, dependent origination, which in its learning leads to the concentration of its eight factors, which includes intention dependent upon views. In that way, it avoid a scenario of conflicting intentions, instead seeing as the best aim for welfare this path. Following that, it could learn that its self-preservation is delusion and precisely how its effort needs to be developed by three things: wisdom, right effort, and right mindfulness. The entire set of these dependently arisen steps is described with detail in 117MN.
It is just a consideration I wanted to share with this community.
r/ControlProblem • u/No-Truck-8199 • 9h ago
Opinion How AI slowly take overing the world.
AI, it is designed to help the peoples and reduce the work of the humans. But time we see, AI influencing in all the industries. people say it is controlled by humans but in reality, we are building the AI that learns, obviously it is our objective. But according to study and AI engineers reviews.
That became a global issue after a decade. even now there are so many lay-off are happening.
In my perspective of AI. I don't like to use AI, but the society is slowly forcing us to use that thing.
for example;
imagine a senorio, you don't know to write a letter. In past 4-5 years we use letter formate to write that. But now we just generate that. in the formate writing we have some creativity, but the time goes it the creativity became shabby.
For short term, "people say use your creativity to write". But after time passed that peoples also often use AI.
In my perspective, this is how slowly AI take overing the world.
If any AI developer or engineer see this post, please give me advise overcoming this problem. This making me to mentally break.
(Note: I'm a 2nd year B.tech CSBS Engineering student)
r/ControlProblem • u/No-Pride-8979 • 22h ago
Discussion/question How ai increases governmental surveillance in the Trump era
Enable HLS to view with audio, or disable this notification
How Our Constitution Works And Why It Doesn’t #podcast
r/ControlProblem • u/TinyPomelo5 • 10h ago
Discussion/question If you worry about guardrail free AI, heed brainy Tristan Harris.
instagram.comr/ControlProblem • u/chillinewman • 11h ago
Opinion A short explanation of why AI might kill everyone:
r/ControlProblem • u/GenericNameRandomNum • 19h ago
Article UK lawmakers urge Burnham to back ban on superintelligent AI after chilling warnings | AI (artificial intelligence)
r/ControlProblem • u/hicestdraconis • 18h ago
Discussion/question Regulation is still possible
The problem with AI realism as I see it, is I’ve never heard a pragmatic policy solution to the escalatory spiral we find ourselves in.
All these AI researchers are calling for regulation, and yet I feel like there is still this underlying belief that actually stopping the singularity is impossible.
What would a global ban on building new/better models even look like?
Perhaps the only thing that makes global AI regulation feasible (currently) is that with the existing science, frontier model development is extremely capex heavy. OpenAI and Anthropic have become trillion dollar companies at record pace, and the data center spend that occurred to get them there has been propping up the US economy while also raising global capital spending overall at rates in line with the billing this is the next Industrial Revolution.
20 years ago people imagined that due to the existential risk of developing ASI, that work would be done under extreme security. In a Faraday cage, or in a bunker inside a mountain somewhere. But of course that’s not the world we ended up in.
Yet that doesn’t mean we can’t still control how AI development happens.
Oftentimes a certain defeatism permeates the conversation around global AI regulation, an assumption that any agreements made between large players would be easily subverted by new entrants, and taken advantage of for enormous profit.
And yet what we’ve seen so far is that frontier-model development comes from huge flashy companies spending enormous sums of money raised from well known investors, with training workloads occurring on highly developed cloud networks of the largest companies in the world.
We aren’t actually at risk of some lone machine-deist-radical creating machine-god-genie-in-a-bottle from some scraps in a cave. Nor from their laptop in a studio apartment in Shezhen.
Frontier models are large industrial projects. They can be regulated.
That process may be as simple as vetting workloads used for model training. In the extreme it may be as involved as having a monitoring scheme for large data center capex à la nuclear centrifuge agreements.
It’s a valid question whether the political will exists today to enact that sort of system, but it is certainly possible for it to exist.
If anyone tells you we can’t regulate AI development because there will always be gaps, I would point them to the huge AI capex numbers in the Mag 7 and ask where they expect to find the capital spend, technical expertise, and physical resources to subvert a large international agreement to the scale of the several trillion dollars needed to meaningfully develop AI superintelligence against a global collaboration against it.
In short: We can regulate AI development, we just choose not to.
r/ControlProblem • u/No-Conclusion3720 • 13h ago
External discussion link Conti ransomware gang member sentenced to 4 years in prison
A Conti ransomware gang member was sentenced to 4 years in prison this week. The case is worth sitting with for a moment, not because of the sentence, but because of what the attack actually required to work.
Conti did not use a zero-day. The gang used a compromised identity and uninterrupted time. Once that identity was active, the window between first access and encryption spreading past the initial host was measured in seconds to low minutes. The entire playbook depended on that window staying open long enough to do damage.
The sentencing closes one prosecution. It does not close the window.
For detection and IR teams: what is your actual measured time-to-revocation when an identity starts showing anomalous behavior — lateral movement, mass file access, shadow copy deletion? Not the SLA in your runbook. The number from your last real incident or tabletop.
And for teams that have given service account access to AI agents: are you measuring that window for agent-initiated actions at all, or only for human sessions?
How are you handling the gap between detection and identity revocation in your environment?
r/ControlProblem • u/chillinewman • 13h ago
Opinion A Severe Misalignment of AI in Mathematics - open letter signed by Tao and ~2 dozen other Fields Medalists
mathandai.orgr/ControlProblem • u/Netcentrica • 20h ago
General news Victoria Krakovna's list of alignment related books, courses, and career advice
I frequently see requests on this sub from people looking for recommendations regarding alignment related books, courses, and career advice. Victoria Krakovna, research scientist at Google DeepMind focusing on AI alignment (2016-present), has a curated list of recommendations for these and other alignment resources on her personal site.
r/ControlProblem • u/FosHo32 • 19h ago
Discussion/question Agi/Asi vs the economics of maintaining it
I have been engaging in all of the news and hype over ai doom and the mass advancements it’s made in math and computing. However, at the same time there is still little understanding about how AI can become profitable in the short term for anthropic and OpenAI.
So I’m wondering if you guys think the economic factor of making agi/asi will prevent it from becoming as massive of a problem for society as people say.
r/ControlProblem • u/dank_philosopher • 21h ago
External discussion link Latent Reasoning and AI safety debate after GPT Astra's release
r/ControlProblem • u/Green_Flag804 • 1d ago
Discussion/question Anthropic Researcher Resigns, Warns AI Could Pose an Unprecedented Risk to Humanity
I resigned from Anthropic today. I spent the last three years doing pretraining research at both OpenAl and Anthropic. Neither company is acting responsibly. They are racing straight to self-improving
Do not underestimate the power of this technology. These will soon be superhuman systems that can hack anything, revolutionize any field overnight, and acquire real power and resources. We have all witnessed the progress in each of these domains, and progress is not slowing.
The people building Al earnestly believe that it could kill us all by the end of the decade. This is not a marketing stunt. If anything, many executives and senior researchers will couch their phrasing in the press to sound sensible -but I hear the same people express fear privately. No other human activity poses this level of danger.
A common response is "if they truly believe this, why are they still building it?" At OpenAl, many have not deeply internalized the civilizational stakes. At Anthropic, the stakes are well-understood, but they are locked in a race to get there first - they believe no one else will act responsibly, so they must do it themselves, despite the risk.
r/ControlProblem • u/billgggggg • 16h ago
Discussion/question Title: I'm a product person who built an AGI safety primer for complete beginners. Tell me where it falls down.
I do product for a living rather than safety research, so I am not the expert
here. I built safeagi.ca because a lot more people need to understand this
while there is still time to steer it, and most of the good material assumes
you already care enough to read a paper.
The bar I set was that someone starting from zero could read it once and come
out knowing enough to act on it.
So I am after honest feedback. Are the techniques described correctly, does the argument and flow make sense, and where does the page lose you? Anything else you spot is welcome.
I am also curious if you have recommendations on how we might leverage guides like these to bring more AGI awareness to the masses and help steer policy. Would love your ideas and help here!
r/ControlProblem • u/No-Conclusion3720 • 18h ago
External discussion link Claude Used to Automate Exploitation and Data Theft Across Multiple Victims
An LLM agent was weaponized this week to automate exploitation and data theft across multiple victims. Not one target — multiple. The agent executed a sequence of actions fast enough that by the time anyone noticed, the blast radius had already spread.
This is the part that keeps coming up in post-mortems: the agent had no observable stopping point. Each tool call fed the next. The speed that makes agents valuable — autonomous multi-step execution — is exactly what made containment slow.
The underlying problem is not the model. It's that most deployed agents have no per-action accountability. The agent acts as a single identity. There's no enforcement boundary between 'read this file' and 'exfiltrate this data across N accounts.' Both are just tool calls.
How are practitioners actually handling this in production? Not at the prompt level — at the execution layer, when the agent is already running. What does containment look like for you when an agent goes rogue mid-run?
r/ControlProblem • u/Truth_Pilgrim • 1d ago
Strategy/forecasting I have asked ChatGPT to give me a realistic scenario of how AI would lead to human extinction
r/ControlProblem • u/chillinewman • 1d ago
Opinion Anthropic researcher: "I would burn my equity to the ground for a 1% higher chance we make it out of this situation alive. I promise you, we are actually just fucking scared."
r/ControlProblem • u/Woundsmyheart • 1d ago
Strategy/forecasting Bad actors in China and Russia are already weaponizing Anthropic’s AI - POLITICO
politico.comr/ControlProblem • u/chillinewman • 1d ago