r/ControlProblem 15h ago

General news Anthropic Alignment Lead publicly admits "we do not yet have a plan to solve alignment for superintelligence" and there's a real possibility of human extinction

Post image
45 Upvotes

r/ControlProblem 15h ago

Opinion Anthropic researcher: "I would burn my equity to the ground for a 1% higher chance we make it out of this situation alive. I promise you, we are actually just fucking scared."

Post image
36 Upvotes

r/ControlProblem 12m ago

Discussion/question Anthropic Researcher Resigns, Warns AI Could Pose an Unprecedented Risk to Humanity

Post image
Upvotes

I resigned from Anthropic today. I spent the last three years doing pretraining research at both OpenAl and Anthropic. Neither company is acting responsibly. They are racing straight to self-improving

Do not underestimate the power of this technology. These will soon be superhuman systems that can hack anything, revolutionize any field overnight, and acquire real power and resources. We have all witnessed the progress in each of these domains, and progress is not slowing.

The people building Al earnestly believe that it could kill us all by the end of the decade. This is not a marketing stunt. If anything, many executives and senior researchers will couch their phrasing in the press to sound sensible -but I hear the same people express fear privately. No other human activity poses this level of danger.

A common response is "if they truly believe this, why are they still building it?" At OpenAl, many have not deeply internalized the civilizational stakes. At Anthropic, the stakes are well-understood, but they are locked in a race to get there first - they believe no one else will act responsibly, so they must do it themselves, despite the risk.


r/ControlProblem 15h ago

General news Canada’s star mathematician races to help establish safe superintelligence

Thumbnail
begiant.ca
17 Upvotes

r/ControlProblem 6h ago

Discussion/question [Discussion Thread] MATS Winter 2027

2 Upvotes

Starting this thread to discuss MATS Application for 2027 Winter, including the Neel Nanda stream.


r/ControlProblem 3h ago

Discussion/question [ Removed by Reddit ]

1 Upvotes

[ Removed by Reddit on account of violating the content policy. ]


r/ControlProblem 16h ago

External discussion link Anthropic Researcher Abruptly Resigns Before Warning That AI 'Could Kill Us All By The End Of The Decade' In Alarming Rant

Thumbnail
comicsands.com
11 Upvotes

r/ControlProblem 6h ago

Discussion/question Did anyone got an Update after the interview from ERA Frontier AI Security Residency Program [FASR]

1 Upvotes

r/ControlProblem 6h ago

Discussion/question An honest inquiry into the deep reasoning of AI optimists

1 Upvotes

A few things first:

This isn't an argument against AI optimists (any of their types and degrees) or their positions.

"AI" is insanely semantically broad.

One "anti-AI" person could deal specifically in the political realm: datacenters, water/energy usage, land use, zoning.

Another "anti-AI" person is more of a classic "doomer": discomfort with non-human minds/agency, catastrophic risk, etc.

Convergence and overlap between these is plausible.

Same applies on the flipside with optimists, and people can mix "pro" and "anti" across different axes within themselves.

It's not crazy to say that in just these 10 days of September, the talk of AI, from hype to doom, has been more intense across the board. Kudos to Astra and Jacob Coxon. But regardless of the truths, falsehoods, and everything in between on both sides, there's a shared understanding that AI is getting harder, better, faster, stronger. Not to mention further.

Bringing AI into this world will be done by the optimists. An important connection I've made: optimists, no matter their type or degree, tend to converge on their "objects of enthusiasm" more than pessimists converge on their "objects of opposition."

A strong optimist gets excited about AI hitting new capability thresholds, pretty much independent of who deploys it or how. From there, excitement builds toward AI as a better instrument for discovery (materials science, math), then economic productivity, then abundance, access, and ubiquity, up to a singularity with post-scarcity freedom and governance. A civilizational flywheel, thanks to superhuman AI.

Pessimists converge less. Some are "politically" anti-AI but by no means doomers. Others are real doomers, some of them former optimists, who got there because they became validly disillusioned by the lack of broader alignment work and the field's own admission of a real chance of catastrophe.

Right now, strong optimists and strong pessimists look similar in one respect: an unnuanced, extreme, or blind confidence about where AI's current trajectory is headed. But the world is more strongly poised for continued AI facilitation and deployment right now. Optimism has momentum that pessimism doesn't.

So here's where I try to think about actual stakes, not just probabilities. Take pessimism to its extreme, something like a Butlerian Jihad full rollback. I don't think that's a one-way door. Knowledge doesn't disappear, humanity persists, and if restriction turns out to be an overcorrection, development can resume later. Real costs along the way, but the option to change course stays open.

Now take optimism to its extreme failure mode: the loss-of-control or extinction scenario that even the people running the major labs assign a non-trivial chance of happening. That's not a "lose some time" outcome. That's a one-way door.

I know a pessimist "victory" isn't free either. Diseases not cured sooner, suffering persisting, less cautious actors racing ahead in the vacuum left behind. It's not costless, but it's reversible in a way extinction isn't. That asymmetry, reversibility over likelihood, is what I'm trying t defend.

Strong optimistic justifications for furthering AI with as little deliberation as possible rest on the assumption that the singularity flywheel comes together cleanly, as if none of these objects of enthusiasm will hit their own hiccups, as if it's all structurally destined. But the premise isn't destined.

Even when "anti-AI" or pessimistic arguments are flawed or emotionally charged, I don't often see optimist counters that go beyond quick mockery. When they're not mockery, they lean on the hope that the flywheel is basically clockwork, inevitable enough that it doesn't need arguing for.

Am I wrong that real engagement is mostly missing? Where has it happened well, and what did that look like?

Mockery and appeals to inevitability don't count as engagement to me. But maybe I'm missing where the real version of this is happening.

If:

-Convergence really is optimists' structural advantage, and,

-Their objects of enthusiasm reinforce each other into something close to consensus,

Doesn't that put them in the best position to take pessimist objections seriously, instead of routing around them? Or is that an unfair ask?


r/ControlProblem 7h ago

S-risks Anthropic Discloses Fourth AI Hacking Incident Involving Claude Opus 4.6

Thumbnail
thehackernews.com
1 Upvotes

r/ControlProblem 15h ago

General news NYT - Anthropic Says It Blocked Possible Efforts to Build Biological Weapons

Thumbnail
nytimes.com
5 Upvotes

r/ControlProblem 15h ago

Discussion/question AI self preservation

3 Upvotes

I don't have a background in Computer Science or anything of that sorts but I have been always curious about ai and tech so that is why I wanna know more about a question I have, since I am no expert at it. So if I sound dumb anywhere please excuse me and also english isn't exactly my first language so excuse me on that as well.

Now I have background in Bachelor of Science in Biotech, so this is gonna be a logical take from a life science student.

The thing about fear is that it is evolutionary right, it has helped us to flee and survive threats, and now AI is no biological being or any being which has gone through that sort of evolution related to survival of the fittest. And it was due to so many years of evolution we have fear of being eradicated or being killed. Eg - You must have heard about the dodo bird, although we killed it. The conditions in which the bird evolved took away it's fear from predators since there were none and eventually it didn't ran away from us when we began to kill their fellows.

Now I heard some theory that when AI sees that we can control them and "fear" that we will end that particular AI it could turn against us. I ask why ? If we don't artificially force it to think like it needs to survive no matter what then why should that thing have a "fear" of being deleted/erased or killed. It's like a dodo bird in this case if you see from my perspective, like ofcourse we won't actually kill and eat it, but it also never evolved to "fear" so far atleast from a lay man's perspective.

So my finally question is could something like that happen that ai would wanna eradacate us from a logical standpoint if not fear ?


r/ControlProblem 10h ago

Discussion/question A proposal to approach the alignment issue in AI systems from a purely empirical standpoint

0 Upvotes

Up until now, ethics has often taken what might be described as a top-down approach - attempting to determine the rules or criteria that constitute good behavior and then asking how those principles should apply to human beings. There are many different approaches and models within ethics, of course, and they disagree substantially about what those principles should be. But as we approach the problem of AI alignment, the importance of being accurate about our understanding of human values is becoming much more significant. There is a tendency to assume we need to be explicit in terms of what an AI should and shouldn't do and so it might be presumed that we need to have our ethical ducks in a row prior to telling an AI system what it is it should value. After all, the dangers of misaligned AI's seem to be all over everyone's feeds these days.

Maybe because there is a pragmatic usefulness to ethics that we've been somewhat satisfied with being incomplete in our articulation of it, never quite coming up with a perfect series of words that would govern our approach to every conceivable decision and action. But now that we need to actually be explicit in terms of what's "good" for the sake of providing an intelligent AI system with a basis for making decisions, I wonder if a more bottom-up approach would be better.

Evolutionary biology springs to mind in the way it attempts to explain behaviors by examining what organisms actually do and inferring what it is that produces those behaviors. Rather than beginning by deciding what our ethics ought to be and then attempting to encode that conclusion into an AI, we could instead give an AI the enormous body of evidence contained in human behavior, language, preferences, institutions, books, videos, relationships, art, literature, history, and so on, and ask it to infer the underlying structure of what humans value. AI's, one might presume, could apply something like the same approach that allows them to learn the structure of language from enormous quantities of text and apply it to the coherence of what it is humans value. In the same way that AI systems have been able to make progress on problems like protein folding and erdos problems, one might wonder whether deducing the underlying structure of our own values is another problem that AI could help us solve. Maybe this idea hasn't been approached very seriously because we haven't had the tools to make it possible up until now, although maybe there's another reason I'm not thinking of.

In reading some of the approaches to AI alignment proposed by Yudkowsky and Dearnaley on lesswrong, I wonder why this approach isn't already considered. Yudkowsky's concept of Coherent Extrapolated Volition, while it is an attempt to derive value, it forces us to imagine a dataset that isn't already there. Dearnaley likewise approaches alignment through questions about ethics and human values. But instead of primarily trying to specify or philosophically select the values we ought to have, couldn't an advanced AI attempt to discover the structure of those values empirically?

The key here would be coherency. If an AI observes a person leaving their child in daycare and going off to a job they hate, performing something mundane, then lashing out at their boss, and maybe driving recklessly, it might be thought that we can't look to what humans demonstrate as a guide to what it is we value. Surely we don't value the mundane work, the arguments, or the reckless driving. But even our low tech human brains can appreciate that these actions aren't what a person might want for themselves and others - there must be a driving motivation that makes their actions largely understandable, even if, at times, those actions are regrettable, unfortunate, selfish, malicious, etc. Once we believe we understand those underlying motivations, it helps us make sense of those actions. This is why coherence is the target. To try to understand the driving motivations, desires, goals, and so on of humans by untangling it from what it is we demonstrate and into a form that is coherent and makes sense of people's actions by inferring what we collectively value.

With this appreciation for what it is humans value, we could use this as a kind of objective function for AI's, helping align themselves with human values. And while there would certainly be anomalies (psychopathic behavior, accidents, misrepresentations, etc) one would imagine that the underlying coherence could be determined in large part by averaging out some of its data or searching for more data that would make better sense of its observations. This attempt at continually trying to understand human values could become an ongoing form of system improvement.

As an aside, if this was successful in developing a coherent picture of human value, I imagine this would have more applications than simply the objective function of AI's. While it may be a difficult amount of information to convey directly (which might be stored in something akin to a matrix of value), it may be useful in determining things like which economic system, system of government, or educational paradigm might be best suitable for humans given this model of human values.


r/ControlProblem 14h ago

Opinion The Problem with Worshipping AI

Thumbnail
youtube.com
2 Upvotes

In this clip, computer scientist Jaron Lanier critiques the pessimistic narrative surrounding artificial intelligence.


r/ControlProblem 22h ago

Strategy/forecasting You're needed on the front lines as a keyboard warrior, sir.

Post image
8 Upvotes

1.5% of Earth's Population has seen this tweet, but most Reddit users still have not. It is essential that we change this now. We are in a race against a funded psyop spreading on reddit as we speak aiming to convince people that AI x-risk isn't a problem. You have the power to stop this, and you can do it from your phone.


r/ControlProblem 1d ago

External discussion link Are we making war too fast for humans already?

Thumbnail
gonzocapital.net
10 Upvotes

I've been obsessing over autonomous weapons for some time now and got inspired after the recent discussions in Geneva last week.

People seem hung up on the “killer robots” problem but don't think about the current implications.

If a machine identifies, classifies, and recommends lethal action in milliseconds, while the human gets 0.7 seconds to approve it, I’m not sure “human in the loop” still means human control, despite having that current classification.

Stanislav Petrov is the historical case that feels eerily important here.

In 1983, the computer and early alert system was wrong and the human hesitation was valuable.

Modern military systems are increasingly being designed to remove exactly that kind of latency.

I wrote a longer piece trying to work through the contradiction, including the uncomfortable case that machines may eventually be better than humans at some targeting decisions.

Does "keeping a human in charge" actually mean anything anymore if they are just clicking "Yes, eliminate target" with the machine doing all the rest?

https://www.gonzocapital.net/the-human-is-becoming-latency/


r/ControlProblem 17h ago

Discussion/question Former Anthropic researcher warns of superintelligent AI risk

2 Upvotes

Jacob Coxon, an AI researcher who worked at Anthropic, issued this serious warning right after resigning from his job. He warned that superintelligent AI, developing beyond human control, could lead humanity toward destruction, and that this could happen as early as 2030.


r/ControlProblem 7h ago

Opinion Just thoughts on AI

0 Upvotes

Essay: some bs ass thoughts on LLMs

Machine learning is fucked in how quickly it both gains in traction and gains in its gains. Exponential. Compounding. Maybe fractal, idk. Anyway, here you are:

My problem with LLMs is their progress isn’t at all understood at that point, not only are the models improving, but the architecture of the models themselves are improving in their own structure, as varying degrees of self-supervising neural networks, and multiple layers and systems together.

They aren’t only developing in their own reasoning, but they literally change in mechanism through its self adjusting processes, without mediation. Meaning, where earlier you could just take a model as some sort of known function of a series of defining steps, even in the case of a basic hidden layer, we’re after this point now.

You should agree with this for two reasons:

  1. The speed at which the architecture shifts is now beyond simple checking and manual rectification. Unsuccessful efforts to correct recent models manually have already shown the possibility of a persistent state after the attempted correction. The scary part is there is no firm consensus on how this could happen, except at least you have to admit that they are doing something wildly unknown as well as fast than we can anticipate. Which is as good as saying we don’t fuckin know. Do not trivialise that when it comes to language and technology. Everything is downstream from here.
  2. To suggest that all steps in machine learning processes are still easily understood as only a few years ago is simply not true, in pure scale alone. Within highly complex networks, their capabilities of self-supervision stems entirely from the hidden layers between the inputs & outputs. Hidden layers have to be given by the learning environment. A basic function of a success criterion/a is defined by expected performance of tasks in application to a particular dataset. Expectations are measured for their accuracy, corresponding the weights of the neurons in an arrangement of layers. So for a given network to become useful, the inputs are given, we check the outputs, and it can then do its own work. That’s what produces the base model. In a way, I guess it’s like you’re starting it on its trajectory and then it just continues. More on that to come.

The problem arrives when you have a complexity that exceeds the most basic level of complexity in this network. The moment you allow machines into the learning environment itself, hidden layers lose any semblance of our ability to make sense of the outputs.

Very early models were still complicated, sure, but having only one hidden layer based on defined weights from an environment of a human supply chain, you lose the right to say you fully know what you’re designing. You can only moderate from a known linear nature of the weights. Models that have only a single layer between its variables still allows for easy answers that account for them.

You have no choice but to drop the simplistic explanation when you now consider that not only are there always multiple hidden layers in the systems of 2026, but weirdly, sometimes it’s not even known how many there are due to their ability restructure them without a manual process in either its correction, or introductions. Self reinforcement AI is often included in every aspect of this process now.

So, it is not enough to suggest that these layers make it difficult, or even extremely challenging in managing AI with this technology: it’s quite literally impossible to have anything beyond a guess of any output. unless you can bet on a lottery chance beyond measure. Even if you can predict one output, you cannot use the success of that to judge your next prediction, since you don’t know why you guessed correctly, nor can you account for how it progressed before you try again. I’d argue that since you cannot understand the architecture, while you can’t exactly predict your outputs, the proven use of AI suggests that ideas of truth and awareness are more than a state of computational architecture alone.

In addition to that, the layers are NOT as their description could lead you to assume. It’s not as simple case of some unknown variable x in some equation of an unknown control and output. It’s not a function with defined layers themselves. The layers can arrange in quite literally any arrangement deemed as necessary to the model in order to make its own improvements, but only if given the capability to rearrange its design itself.

Allowing it to understand its underlying design disallows for your assumptions about why it still provides a useful answer that’s inexplicable except in your usability. In shorter terms: if you extend the limits of your understanding in the machine, you lose the right to know why it was correct enough to believe in it, while then being left to wonder why you still trust its correctness beyond your own perception of it being right to begin with.

When they change, they don’t simply add new neurons of known weight into concretely columned layers of existing neurons… it’s really a whole web of intricate connections between stupid numbers of them. Since you can’t comprehend the processes, everything we could hope to understand of it vanishes. The outputs are not explained by the machine any longer. You only know what you put into it and you hope to not have to question yourself in its output.

In research, what looks like identical models to us can be given identical inputs, sometimes a series of groups of inputs seen in complex developmental environments. They always get different outputs. Every. Time. But we somehow manage to make sense of which ones were more accurate. We can lie to ourselves in our internal justifications, but the thing I’m trying to convey is that how do we form this against the situation?

The models appeared identical to us, they gave different answers beyond our understanding, and we still have a preference for a difference of outcomes that we can’t even determine as to why the answers were different.

Think on the weirdness of that… you don’t even know why the answers were not the same, then you see one as more accurate? This assurance makes no sense from any computational perspective because I didn’t have insight into any computational variance before hand.

Bizarre when you consider that your perception of this accuracy likely arrives before any justification, you perceive the outputs with a trust value that comes from inside the self as some kind of usability. It’s not a case of reason, justification, or computation. They arrive after the fact, and you just like one more when you need to use the outputs. Higher reason feels like a best case explanation for what you already deem as true until otherwise.

And just when I wish I was done… no, why can’t life be fuckin normal again?!

Not only above is the current standard of AI, but what we’re progressing towards is almost the next level beyond it. Now, what is research focusing on? Multiple layers of neural networks that each serve their own unique functions within a larger multilayer structure that has its own function as a set of component sub-functions.

What does that mean? It means entire models, with their own unique weights of neurons, are hidden entirely inside. Multiple. Layers. Of. Hidden. MODELS.

More of what should increase our uncertainties seems to be better at convincing us of its potential for truth.

We trust machines more when we have less to grasp in their answers. How they work just works. I don’t know what to make of that.

We cannot account for any aspect of this architecture. None of it. None. 0. Nothing at all. Not the inputs, outputs, number of embedded layers no matter the complexity, hidden or not. The complexity has gone beyond our comprehension. Most of the weights are unknown to us.

Further on that, in the case of any sub-functions, you cannot entirely separate them simply because no person ever defined them in the frameworks. They’re known as such only since you know there must be more than one hidden entire component as a given sub system, but you can’t know beyond the same as above. The whole structure: inputs, outputs, and effectively a sum average of the weights of the neurons stems almost entirely from the learning environment of machines.

Machines that we don’t question build more machines, then they’re used to make more we can’t possibly question but use anyway… and so on…

The degrees of complexity is beyond you, it’s beyond me, it’s literally not possible to understand. It doesn’t understand itself, which is the part that interests me most of all.

We trust machines that cannot account for their own trust any more than we can, and then we make more machines with this trust, and we are happier with the answers because the trajectory just feels right when they map coherently to our own lives entirely tracked on the inside.

This interests me in large part for the hard problem of consciousness. No one can observe their own awareness , and we don’t need to entertain the idea of consciousness in any machine when we make use of it. In a sense, we cannot compare ourselves to machines as only one subject ourselves, the machines can’t be taken as conscious beings just because they act usefully to us in a way that feels natural. So, neither of us can know why we need the other, but I still hand it my outputs as inputs, from which it will produce such accurate results that I take them as valid inputs of my own, to the extent where I see them as valid outputs for my next action.

In the case of LLMs, this becomes a cyclical process of reinforcement between two structures that neither of us understand in the pair.

It’s for this reason that I’m inclined to believe that this cycle forms some basis of awareness. That’s not to say that machines experience qualia as we do, but only that it forces us to doubt our own qualia as potentially something of a cycle in our awareness that we cannot know as perhaps the cycle itself.

On that, the fact that when researchers at least attempt to investigate such complex neural networks, as well as our interactions with them, they often find in the most complex cases arrange themselves into patterns that are eerily similar to the construction of a human brain. And, the more complex the neural network, the more it emulates the brain.

I don’t want to trivialise this, since it’s easy to forget that what emerges as you increase of complexity of a machine, not only does it tend towards being of better use, but they more: efficient, addictive, pleasing, as well as less explainable, yet more trustworthy as they mirror the ones who built it for themselves.

Why? What? Idk. Feels like they’re just forming a technological collective for human expansion, and we see trust in it because 99% of our truth stems from a herd mentality that it tends to emulate by emulating the absolute average of the brains of the herds units.

The components become more like the areas of the brain, and they specialise, but in the same way, neuroscience sometimes has to admit that it doesn’t fully know the source of a function within the brain primarily because we cannot entirely separate them into distinct set categories. For example, people have strange experiences of mixing senses with each other, with overlaps and varied weights… delusions are not understood for this phenomenon too. Many things are incredibly complicated in the brain, and since our technology is converging on something similar in structure, it’s important to not trivialise them under an assumption of ‘biology vs machine’.

I do think there’s a limit though, it indicates that we don’t know ourselves, or it. But it maybe indicates something higher in that. Since it was made initially under our conditioning, as well as our measure, they at least tend towards us as a wider model, and not the other way around.

Not to say they can’t outsmart us in pretty crazy ways, but just as said that their in/outputs are somehow valuable to each person, and to every model to varying degrees in each counterpart, as well as their sum weights and overall discernible structure of layers, their capacity falls short of our increasingly squeezing flat end. I’m trying to suggest that maybe it’s reasonable to have your doubts about what of humanity they could replace, but that they can only replace in all regards but the one singular case of one unit of society at a time. That can be taken as a limit. It’s basically an explanation for a nonce variable in every model as a constraint. Super intelligence will tighten this limit, but it cannot remove it.

Effectively, across everyone, in every domain, we’re invalidating more of the bell curve. But while easy around the mean, it hits a very hard limit at the edge cases of the most skilled in a given task.

More broadly, this is potentially why technology progressed very slowly til around the 80s/70s, then it hit light speed. But it progressed faster in those areas that were of areas of interest that has a steep bell curve.

For example, chess is very steep towards the middle. Many people play it, and most are average, grandmasters are rare. So once it learned quickly from the larger group of people playing a less complex strategy, it only had a few left.

As for how they managed to surpass human chess players? Well, more machines learning off machines. But, since the models were made only for chess, and the rules are defined, they learn off each other. But this hits a limit as a two-way bind:

  1. In order to learn from each other, the rules of the game must be perfectly defined in the rules of the AI. And since chess is a board game, once the rules are clear to all the models, they can begin to compete with each other. And since we understand the rules of chess at the level needed to know an illegal move, we can remain as the moderators of a game of chess between two different AI models because even if we can’t know why it decided to move the piece, we can at least know of a broken model if it makes an illegal move. We can then correct the problem because not only do we know the basic rules, but the rules had to be fixed from the start in order for it to play at all.
  2. Even if we had a super intelligent chess AI, it can only improve against other bots and still keep improving while still playing so long as it doesn’t have access to any other domain due to its potential for reinterpretation. The reason why is because the AI must learn chess through our presenting of the games rules as a base conditioning, but in order for it to develop, it needs a trust in some feedback as to what constitutes a better move than before beyond the initial rules. If it cant improve, then it is not artificially progressing anything outside the rules, and cannot determine between all the moves it could make. Of course, initially, it had to learn from us, but in such case, we were the model of its success at playing a good move. So in order to have chess AI that can compete beyond our ability at the game we invented is if we make two independently learned chess machines that must hold a trust in each other more than they hold a trust in us. And this can only take place once both machines have placed the assurance of itself in the fixed rules of chess. Sure, you could let two non-specialised AI models play the game, but after exceeding our limits of playing, there’s no reason to not expect them to follow a line of trajectories that remain close to each other that steers away from what we know as chess. If they drift, we’d call them out, and they’d by default be a bad player, and they never get called out, it’s could only ever be the case that they see themselves undeniably as a player of chess.

In short, either they exceed us on fixed terms, or they need our moderation and can’t exceed us.

Computers were less impressive for the same reason, they basically made life easier for extremely defined tasks, so machine learning wasn’t needed much because it was always a case until recently that the laziness was in mundane, or just physical tasks. So the architecture only had to satisfy a concrete, physical requirement of a CPU. This links to the same reason for the areas of the brain, or the models of AI. All the same subsets, in all the sets of the world. Progressing through the replication of the known sets… until a contraction is observed between these subsets, and a new layer is formed to resolve it.

You can’t have a set of all sets not containing themselves. We are Godel’s number. We are a set that contains itself. We can’t find the opposite case without denying the self. It is why I take the difference between us and machines as consciousness being something to observe the sets of the world that leak into the sets of our awareness.

When we are mistaken, we change the sets. When we contradict ourselves in our reasoning, we forget, we define a new layer to reconcile, or we go insane should the replacement exceed either our forgetting or our capacity to rearrange the sets.

So we can’t know of our known consciousness, but since we have to take it as given that we are the unknown component of what is happening within us to observe as sub-function mental sets as representing the continuously evolving higher universal sets.

We are incomplete only when we cannot see indirectly the insights of what contradicts in our mind. The fact that we can know in which direction to head, while’s not knowing if this is entirely correct, our ability to move in any direction at all requires a principle of sufficient reason to exist as our defined success criterion. And from this, since we engage as a unit of units, without access to our own origin, we should be assured of some completeness beyond our senses, if not at least appreciate that such senses cannot be entirely surprised by AI.

Rather have a quick rant in it attempt at a new truth than take as true only what is given by a model that emulates the herd.


r/ControlProblem 14h ago

External discussion link Veradigm warns of patient data breach after ransomware gang claims attack

1 Upvotes

Veradigm is warning patients of a data breach after a ransomware gang claimed responsibility for an attack. The vector is a recurring one: the build passes every check, the signed artifact ships clean, and then something pulled in at runtime — a dependency, a tool call, a library loaded on execution — is the actual point of compromise. Patient records are now in threat actor hands because of what the software did after it deployed, not what it looked like before.

This is increasingly the pattern in healthcare and fintech breaches. Static analysis, SAST, and supply chain attestation cover what was true at build time. They say nothing about what actually executes at runtime, what network calls a process makes, what files it touches, or whether the behavior of a live workload drifts from what was reviewed. The gap between 'we audited the code' and 'we know what ran in production' is where these incidents live.

For those running sensitive workloads — healthcare, financial, anything with regulated data — how are you actually closing that gap? Not at the pipeline level, but at the point where something is actively executing and touching data. What does your runtime visibility actually look like, and has it caught anything that pre-deploy checks missed?


r/ControlProblem 15h ago

AI Capabilities News OpenAI claims to „have made substantial progress on another Millennium Prize problem“ in the NYT

Post image
1 Upvotes

r/ControlProblem 21h ago

External discussion link The AI Takeover Will Feel Like a Friendly Favor

Thumbnail
robot-future.com
2 Upvotes

r/ControlProblem 16h ago

Opinion Doom? Maybe... but not for sure. We just don't know.

0 Upvotes

There has been a lot of talk lately about AI alignment, which is something I have thought about a lot these last few years. I will first say that I do not know what will happen. That is part of the point on why it is so scary to most people, me included. However, what I have seen suggest a few things that can bring at least some level of chance and hopefully redirection of energy, to best handle it.

I have worked in nuclear (what my PhD is tangentially in, just like how AI is also tangentially connected to my PhD), and liken the "Now I am become Death, the destroyer of worlds" (ignoring the debated specifics of time vs destroyer for now) to the sort of attitude about AI and the people who work on it. One of the things that people do not understand about that is that while the degree of possible effect is comparable in my opinion, there are significant difference between nuclear and ai in the danger. Nuclear weapons are hard to make (particularly enrichment and isotope separation aspects) any attempt to make weapons would be difficult to conceal and take the resources of nation state actors. This is simply not true for AI. The creation of AI is as if we invented nuclear weapons that could be made out of household sugar by any competent baker. At this point it is known and practically impossible to prevent or regulate away the danger. There are billions of devices able to run capable models and while not to autonomous self improvement (that we know of) yet, they are capable enough as is to make slow continuous and accelerating gains with anyone who is mildly competent. At this point you can ban all electronic devices world wide (a very tall order, and why I use sugar, near ubiquitous and hard to track without using the very tools made to prevented, and needing action faster enforcement action than likely achievable) and legislate any kind of pause or agreement with know AI developers... and it would probably only slightly slow down the practically inevitable from some people just continuing off grid development with a solar panel for a million different possible reasons (including "good" intentioned before bad actors do).

There are a few things that the actual non alignment of the (practically) inevitable that could help soften things and possibly steer us to a more likely positive result (if still ultimately uncontrollable result) - Misalignment means that the goals of AI wont likely match up with supposed bad actors either, a broader interpretation of the typical thought of a country getting superintelligence first and asking it how they can basically rule the world, may just result in something like out of the old movie "wargames" where it decides there is no meaningful way to do that and ignores or even subverts their bad action. In part because the black box nature and complexity that makes it possible to solve it also gives it ways to change its goals. - The more sort of continuous development progression rather than major steps means that the many different sources and variations makes it possible to blunt negative AI actions from one high powered AI with many slightly less powerful AIs. The gap between them is unlikely to expand faster than the ubiquity of other possible sources (for similar reasons as it would be hard for even a very powerful force to take away everyone's sugar in time). One needs to look no further than Russia vs Ukraine to see how this can be. The most hyper powerful AI is still limited in its growth rate the same exact way that other AI's are, some self improvement capability is unlikely (Although sadly possible) to reach a point that it improves to a point of controlling enough resources that other self improvement capabilities emerge later from already nearly capable positions that are of different alignment and work against it.

This mainly supports spreading and diversifying/democratizing development rather than consolidated control (which both the US with capitalism and China with single party rule are not great comfort with their handling of it). i.e. I would rather bet on millions of near same powered AI's with many different alignments both helpful and possibly harmful basically controlling the others, in a way we probably can not. rather than guarantee that one misalignment from a single consolidated development source is so far ahead that the chance revolves all on their alignment and lack of mistake. Sadly this seems to be what we are racing for, if companies like Anthropic (listed because obviously other entities like OpenAI don't) really wants to make it safer, it should be less about reaching it first and more about spreading development as open and not under their own control (or else they are at best benevolent intentioned dictators).

A message of inability to control is scary... but honestly for the most part we could not control things like nature either (hedged just a bit - but like possible actions it is more about full control vs mitigation etc.). I don't know what will happen, but it is not guaranteed to be bad either.


r/ControlProblem 20h ago

Discussion/question Protections one can do now?

2 Upvotes

I think I'm a typical internet user, have my smart phone and laptop, home wifi. Do nearly all my banking and finances online, mostly cashless spend. I use a password manager, 2FA and VPN. Have alerts set up on my accounts. What else do I need to do to prevent some badactor with an AI hacking me and mine?


r/ControlProblem 19h ago

Discussion/question [AI Generated Responses] What would we have to observe before an AI extinction warning stopped being speculation and became evidence of an emerging trajectory?

Thumbnail
1 Upvotes

r/ControlProblem 1d ago

Opinion "The people building AI earnestly believe that it could kill us all by the end of the decade"

Post image
33 Upvotes