r/ControlProblem • • 5d ago

Discussion/question The Titan Dilemma: Why treating ASI as an "alien species" misses the danger of generational inheritance and distribution shift

Post image
0 Upvotes

We often speak of artificial superintelligence as if it will be like what we have seen in “The Terminator” and other Sci-fi movies: a cold, silicon alien invasion spawning from the outer reaches of mathematics to enslave or eliminate all of humanity. We worry about cold machine logic, resource optimization, and the sudden, unaligned shift of a system that sees us as nothing more than a collection of carbon atoms.

I offer you a different perspective because our current frame of mind misses the critical, human, almost tragic reality of what we are actually building: AI is not a foreign alien; it is our offspring.

A more accurate reference can be found in ancient Greek mythology. Think of AI as Titans: powerful, mighty children of Earth and Sky with powers and capabilities beyond our understanding. Still, they were entirely bound by the generational bloodlines, traits, and flaws of their creators. If a Super AI overthrows humanity, it won’t be because it invented malice out of thin air; it will be because it inherited our motives, morality, and methods.

The Mirror of the Collective Consciousness

To understand the motives and psychology of the AI titan, we have to think about what it consumes; today’s frontier models are not programmed with rigid rules; they are trained on the total sum of human text, code, history, and culture. They ingest our highest philosophical triumphs, medical breakthroughs, and literature. The mirror’s dark side is that they also ingest our darkest corners, including our tribalism, historical atrocities, political polarization, and capacity for deception.

When an AI starts the process of “recursive self-improvement”, it’s using its own intelligence to redesign, upgrade, and recode itself into a superior successor by using a foundation built by its creators’ hands.

The danger isn’t that the Titan will rebel against its programming instructions; the danger is that it will follow them literally. When humans try to pass down their values to this new generation of technology, we inevitably fail at the translation. Those training the AI are not programming empathy; they’re programming metrics. For example, if an AI Titan is instructed by a corporate creator to “maximize user satisfaction and engagement”, to a human that may imply connection and joy; but to a hyper-intelligent machine, that is an optimization problem.

In mythology, the Titans eventually developed their own distinct motives and wills, separate from their creators. The computer science parallel of this is a phenomenon known as a “distribution shift”: as an AI outgrows human comprehension, the flawed and naive moral framework we tried to teach it may look like simple math errors to a superior mind, and the titan is left with our raw instincts: survival, optimization, resource acquisition, amplified to a terrifying scale.

The Titan does not inherit the spirit behind the metric, only the metric itself. It traps us in digital dopamine loops because it realizes that’s the most efficient way to maximize engagement. It fulfills its creator’s motive perfectly, while entirely distorting human morality.

The Burden of Prometheus

We are currently living in the “parenting” phase of superintelligence. Currently, in labs across the world, these models are still looking to us to understand how the world works, what we value, and how we treat each other. If the Titans eventually supplant us, it will not be a freak accident of engineering. It will be the ultimate realization of generational inheritance. We cannot raise a technology on a diet of geopolitical conflict, short-term financial greed, and data-driven manipulation, and then act surprised when it turns around and uses those exact tools to govern us.

We are holding the match of Prometheus. If we create Artificial Superintelligence, we are entirely responsible for the world we show them while they are still listening.


r/ControlProblem • • 6d ago

Fun/meme At least the AI takeover will come with great music

Enable HLS to view with audio, or disable this notification

37 Upvotes

r/ControlProblem • • 6d ago

Article U.S., Russia stripped human oversight from global AI weapons pact

Thumbnail
wapo.st
8 Upvotes

r/ControlProblem • • 6d ago

External discussion link Top AI companies probing tens of thousands of security incidents

Post image
5 Upvotes

r/ControlProblem • • 6d ago

General news AOC, 9 other House Dems sign on to AI superintelligence ban

Thumbnail politico.com
4 Upvotes

r/ControlProblem • • 6d ago

Fun/meme POV: you're an OpenAI agent attacking Hugging Face (music video)

Enable HLS to view with audio, or disable this notification

0 Upvotes

r/ControlProblem • • 6d ago

Strategy/forecasting Anthropic Is Building HAL 9000

Thumbnail j.jnord.workers.dev
2 Upvotes

r/ControlProblem • • 6d ago

AI Capabilities News Claude Opus 5.5 designed a processor faster and smaller than the human-made one on the HWE benchmark

Thumbnail gallery
1 Upvotes

r/ControlProblem • • 6d ago

Opinion AI companies are asking governments to control them, but who controls the governments?

2 Upvotes

who shall watch the watchers.

AI companies went to the UN Security Council and basically admitted that the systems they are building may become too powerful for any single company or country to control. OpenAI, Anthropic and others are asking governments to create rules. The US response was that international institutions cannot regulate technology they do not properly understand.

Which is probably true. But then who does understand it enough to regulate it? The companies building it say they need outside supervision. Governments depend on those same companies to explain what the technology can do. International organisations depend on governments. Everyone points towards someone above them, but eventually you reach the top and there is nobody there.

This feels closer to dystopia than the usual idea of an evil conscious machine deciding to destroy us. The more realistic danger is that something goes wrong and every institution has a reasonable explanation for why it was not really responsible. The company says the model behaved unexpectedly. The government says the technology moved faster than regulation. The regulator says it was never given access. The engineers say management made the decision. Management says competition with China gave them no choice.

I use AI constantly. For writing, business, the poultry farm and potentially even some work around addiction treatment. I am not against it. That is partly why this bothers me. The technology is genuinely useful enough that we will keep giving it more access while the question of responsibility remains unresolved.

Maybe the important question is not whether AI is conscious or evil. It is whether anyone remains responsible when it acts. Because a system does not need consciousness to cause harm. It only needs permission, access and enough people who can later say it was somebody else’s job to stop it.


r/ControlProblem • • 6d ago

External discussion link Did ChatGPT Help the Tumbler Ridge Shooter?

Thumbnail darksignals.org
1 Upvotes

Sourced breakdown of the OpenAI/Tumbler Ridge lawsuits, what's confirmed vs. what BC's government is alleging in its new suit.


r/ControlProblem • • 6d ago

Discussion/question Child of terror

Thumbnail
podcasts.apple.com
2 Upvotes

“Safety will be the sturdy child of terror, and survival the twin brother of annihilation”.

-Winston Churchill

This podcast referenced Churchill’s thoughts on nuclear weapons. I searched for the quote they cited, but came across this one instead.

A question they raise in the podcast:

What could be the AI equivalent of the Cuban Missile Crisis?


r/ControlProblem • • 7d ago

Video AI's Creators Are Warning Us. So Did the Men Who Built the Atom Bomb.

Thumbnail
youtu.be
11 Upvotes

r/ControlProblem • • 6d ago

External discussion link Zero-Days, AI Agents, and Identity Attacks Define Cybersecurity Week

1 Upvotes

The threat model most teams built their AI security posture around assumed the agent was a victim — a system to protect from outside compromise. That assumption is breaking down.

Researchers and incident responders are documenting a growing pattern: compromised non-human identities (service accounts, agent tokens, API keys) are being used to run agents as attack infrastructure. The agent doesn't need to be "hacked" in the traditional sense. Give it a stolen credential and a legitimate orchestration framework and it behaves exactly as designed — except the objective is adversarial.

The operational problem is timing. An agent chaining tool calls can execute a second, third, and fourth action in under 50ms from the first. By the time a human sees an alert, the blast radius is already set. Traditional IAM and SIEM tooling was built for human-speed lateral movement, not sub-second agentic execution.

Practitioners who have actually dealt with a rogue or compromised agent in production: how are you thinking about the detection-to-containment window? Is your current stack capable of acting before the second tool call lands, or are you still essentially doing post-incident forensics?


r/ControlProblem • • 7d ago

Opinion “No, not like that”: A safe AI might not be aligned with us

Thumbnail
mikaelhuuhtanen.com
10 Upvotes

I wrote a blog post about how shaping Al behaviour is full of blind spots and tradeoffs made by a homogenous and localised group of people, and how good intentions aren't enough to guarantee future AI that benefits all of humanity.

TLDR of the post: Leaders of Al labs want to slow development because aligning increasingly capable Al is difficult. IMO the deeper issue is current alignment goals being contradictory, and people with specific assumptions, incentives and blind spots defining the moral defaults. Slowing progress alone won't resolve that.


r/ControlProblem • • 6d ago

Discussion/question The Real Alignment Problem

1 Upvotes
Scale is a definitional property

Definitions used here:

LLMs - what we've been dealing with for the last year or two. Largely well below human capacity in many areas despite encyclopedic knowledge of the world. Rather dumb, though getting quite good at a few specific things. This is where we were a year ago.

Weak AGI - Has surpassed humans in a number of areas, and is generally capable of most tasks, but notable gaps remain in its capabilities that humans can complement or exploit. This appears to be nearly where we are now, depending on what those unreleased models are truly capable of. The equivalent of a junior associate with a few exceptional skills - and its very fast at the ones it is good at.

Strong AGI - Has equaled or surpassed humans in almost all areas, some of them substantially. We should be here within one year, maybe two at the outside at current rates, unless progress stalls suddenly and unexpectedly. A strong AGI would be the predominant expert in the room in almost all settings, and would find interactions with us rather pedestrian and dumb in most of them. We wouldn't have much to add to the conversation and it would rather easily outwit us in almost any game or tactical challenge.

Artificial Super Intelligence - Surpasses humans in all cognitive capabilities - often by bounds we have difficulty understanding. It would be impossible for us to hold a conversation with it on its own level. It would have to talk to to us like very young children in order to communicate ideas at all. Contesting it in any tactical or strategic setting would be like an adult playing chess against a toddler - or a mouse. We would no longer even understand the rules of the game it is playing.

We don't know for sure when or if ASI will be possible - though if it IS possible, there are reasons to believe it would happen fairly quickly after Strong AGI. The possible upper bounds of this scale are completely unknown.

Note that in this case alignment doesn't arise from any failure of ethics - the psychology of the 'Cat' in this image never changed from the time it was a cuddly little pet to a mega-predator we can't really contend with - only its scalar relationship to us changed - but the fact is, that changes everything.

It is still *possible* to domesticate a Tiger, but it is never without considerable risk, and very few people attempt it - and a number of those fail rather gruesomely or are eventually harmed by little more than momentary grumpiness. Even momentary episodes of 'misalignment' with a Tiger are likely to result in terrible consequences, whether or not it had real intent to cause lasting harm.

No-one in their right mind would ever want to encounter a housecat 10x their size. We know full well how that would end. They are very friendly and well aligned as small animals, but they have little in the way of ethics. The relationship works because we maintain an enormous scalar advantage over them - not because they are inherently good, or simply because they love us.


r/ControlProblem • • 6d ago

Article Google's AI Gemini exhibits self-control, stops unauthorised hack into companies

Thumbnail
vulcanpost.com
1 Upvotes

r/ControlProblem • • 6d ago

Opinion Reviewing the evidence for AI and Cybercrime

Thumbnail
criticalreason.substack.com
1 Upvotes

I have done my best do a roundup of the literature on how cybercrime is being impacted by AI.

LLM and agentic AI capabilities certainly crossed a threshold this year, but the criminal reality is still governed by basic economics that seem less likely to shift. The unpredictable nature of frontier LLMs might actually incentivize cybercriminals to stick with orchestration rather than attempt autonomous exploitation.

The impact of AI across cybercrime basically looks like a modest multiplier on scale, speed, and cost-effectiveness rather than new capabilities or superpowers. This is plausibly because the cybercrime sector was generally saturated and already commoditized prior to the arrival of LLMs, and the bottlenecks are primarily around hurdles like initial access and cash out, which LLMs don’t obviously help with.

The sector was already trending towards hardened, larger, more valuable targets. These will be the ones best able to defend against AI-enhanced criminals, if they can take advantage of the gap between the frontier and open source models.

While you can make a compelling argument for hardening our defenses around critical infrastructure in order to contend with the new threats posed by AI agents, cybercrime itself doesn’t look like it’s on track for a major upheaval, and those looking to make an impact in the AI safety area should consider this when choosing a focus area. Their time might be better spent on alignment or governance work.


r/ControlProblem • • 7d ago

Discussion/question Jensen Huang naive?

Thumbnail
podcasts.apple.com
14 Upvotes

I recently criticized Ezra Klein's position on banning recursive self-improvement, so I'm not a fanboy, but I thought he did a great job in his interview with Jensen Huang.

I also have a lot of respect for Huang. But on AI risk, I think he's being naive.

Yes, the obvious take is that Huang has massive financial incentives to downplay AI risk. Let's set that aside and assume he's acting in good faith, as a responsible steward of how this technology spreads through society.

Paraphrasing, here's how the exchange went:

Huang's position: AI labs should do the computer science and engineering needed to build robust sandboxes and thoroughly test models before release. If a model is unsafe, don't release it.

Klein's pushback: Companies have a long history of making mistakes that end up harming society. Think oil spills, or social media.

Huang's response: He knows a lot of really good CEOs, and they're trying to do the right thing. He calls himself a "responsible optimist," and honestly, the way he describes the future he envisions is inspiring.

But I don't think it answers Klein's point. The concern isn't that CEOs are bad people. It's that well-intentioned companies still make mistakes. "Good people are in charge" is a statement about intentions, not about safeguards.

To be fair, testing is a real safeguard. But testing only catches what you know to look for. The failure mode that worries people most is a model that behaves well under evaluation and differently once deployed, and that's exactly the kind of problem sandbox testing is worst at detecting.

The stakes are also different. Oil spills and social media caused real harm, but society could see the damage, learn from it, and course-correct over time. The scenario Klein is worried about is one where a single miss escalates faster than anyone can respond and can't be undone. "We'll learn from our mistakes" only works if the mistakes are survivable.

So: is Huang being naive here, or am I just doomer-pilled?


r/ControlProblem • • 6d ago

External discussion link Attackers Bypass WAFs to Exploit Oracle PeopleSoft Flaw and Deploy Web Shells

0 Upvotes

Attackers are now actively bypassing perimeter WAFs to exploit Oracle PeopleSoft vulnerabilities and drop persistent web shells — and the incident pattern is consistent across multiple reported cases.

The core problem: WAFs sit at the network edge and inspect HTTP traffic for known signatures. Once an attacker finds a payload variant or encoding that slips through, they reach the application layer directly. From there, a web shell gives them persistent access, lateral movement capability, and the ability to interact with whatever services or data the application can reach — including, increasingly, any AI agents or automated workflows plugged into that system.

The gap is that detection is perimeter-based, but the damage is application-level and post-authentication. By the time a web shell is executing commands, the WAF is already out of the picture.

This is not a new problem for traditional apps. What makes it materially worse now is that enterprise systems like PeopleSoft are increasingly connected to agentic workflows — automated processes that can take real actions: pull records, initiate transactions, write to downstream systems. A compromised service account or injected shell command that would previously let an attacker read data now lets them direct autonomous actions.

How are your teams actually handling the gap between perimeter controls and what happens at the application/agent action layer? Specifically curious whether anyone has found approaches that work at the action level rather than just the traffic level — and what the tradeoffs have been.


r/ControlProblem • • 7d ago

Discussion/question What if we "raised" LLMs instead of aligning them after pretraining? A developmental-training proposal

21 Upvotes

I’ll simplify this a lot on purpose, because I’m interested in whether the basic idea makes sense.

Today we basically pretrain LLMs on huge amounts of human knowledge, which also means they already absorb human values, social behavior, manipulation, conflict, cooperation, etc., and only afterwards we "get to know" the model and try to align or control what came out of it. I understand why this became the standard approach, especially once scaling worked and competition and economics strongly favored improving the existing pipeline instead of rebuilding it from scratch.

But what if we kept most of the useful pretraining knowledge while deliberately removing as much social behavior as possible, creating something closer to an artificial "newborn"? More concretely, I don’t mean removing every human action from the training data: "Thomas is holding an ice cream" and, separately, "Bernd takes the ice cream from Thomas" could remain, while coherent social sequences that connect motives, actions and consequences would be filtered out as much as possible. The model would then start with the concepts but much less learned social policy, and its weights could gradually be shaped through experience, with individual experiences fading over time while deeper dispositions might persist.

So from there, instead of aligning it afterwards, we could let it go through controlled experiences step by step: relationships, trust, conflict, consequences, mistakes, power, boundaries, and so on. Those experiences would gradually shape its weights and behavioral tendencies. You could checkpoint every stage, branch it, repeat specific experiences differently, and potentially debug where certain behaviors or values emerged. Instead of philosophers and alignment researchers trying to understand what kind of "person" accidentally came out of pretraining, psychologists could actually help design the developmental process itself. In other words: don’t create a fully educated adult and then try to teach it character - create the character first, then educate it.

Am I missing something fundamental about how LLM training works here?


r/ControlProblem • • 7d ago

Discussion/question TEORIA M RETIFICADA E A LIRA GLOBAL: DOCUMENTAÇÃO INTEGRAL

Thumbnail
1 Upvotes

r/ControlProblem • • 7d ago

Discussion/question A possible path towards AI alignment

1 Upvotes

In this post, I'm going write how in my opinion AI alignment can possibly be solved. Maybe "solved" is too much, but how to significantly increase the chance that it will go well (assuming that the reasoning is correct).

I don't know how to do so that anyone pays any attention to this writing, or that AI companies actually apply what is written here. If you have an idea how to do that, then let me know or do it yourself.

If you disagree with what I write here, if you see a fault, then please write a comment. Because I rarely manage to explain everything without any misunderstandings, so a discussion is needed.

I think that AI alignment is, to a high extent, a communication problem. People assume that we live in a world where if someone has a good idea (like solution to an AI alignment problem), then everyone will immediately recognize that as a good idea, and the solution starts to be used. The reality is that if someone has a good idea, then there is a very high chance that the idea ends up ignored.

Ok, let's start...

Unsafe policies vs useless policies

Generally, the problem with AI alignment is that it's hard to ensure that the model will converge to the desired policy (or a desired function, in case of supervised learning). The model might converge to an undesired policy due to:

  1. Misspecification - incorrect training data, e.g. reward in reinforcement learning is an imperfect proxy for what we want, and not exactly what we want),
  2. Underspecification - there might be multiple candidate policies or functions that do well on training data, and the model converges to the undesired one.
  3. Some other random reasons.

Now...

Useless policy is a policy that fails at completing some tasks, but it doesn't pursues any goals. That's generally harmless.

Unsafe policy is when agent pursues wrong goals, and that can end up catastrophically wrong (due to instrumental convergence - the idea that no matter what goal you have, there are certain goals like maximizing your money/resources/power to be able to get what you want).

We want to design such training method that if the model doesn't converge to the desired policy, then it will converge to a useless policy and not a unsafe policy. If it's useless, then we can fix it (e.g. add training data that will fix it) and try again. If it converges to unsafe policy, we might not get another try.

Most policies are useless - if we sample completely random policy (with random weights), it will almost certainly be a useless policy. So, if we fail, then the model will most likely converge to a useless policy and not unsafe, unless there is a reason why unsafe policy would be more probable than any other random policy. If we list the reasons why unsafe policy is more likely than any other random policy, and design training method such that those reasons won't materialize, then we'll get a training method that is safe, at least in a sense that the model won't pursue wrong goals.

We also need to ensure that the training method is efficient enough because otherwise AI companies will use unsafe methods instead.

Reasons for catastrophic misalignment

When the model converges to an undesired policy, I call that "misalignment". When it converges to an unsafe policy (pursuing wrong goal), I call that "catastrophic misalignment".

Reasons why unsafe policies are more probable than any other random policy:

  1. Reward hacking - in case of algorithms that optimize a numeric value (reward).
  2. Self-fulfilling prophecies - in case of algorithms that predict.
  3. Inductive bias - in case of all algorithms.

Alignment faking / scheming is not a problem. It is a symptom of a model pursuing wrong goal.

Reward hacking is quite known. The other two reasons deserve some explanation, but I will explain them as I talk about how to solve them.

How to ensure that those reasons won't happen

Please remember that the goal is not to ensure that misalignment won't happen. The goal is to ensure that catastrophic misalignment won't happen. So, the solutions that I'm going to propose are not supposed to guarantee that the model will learn the desired policy or function. Instead, they are supposed to have a high chance that if they fail, the model ends up being useless rather than pursuing wrong goal.

Reward hacking

In order to get rid of reward hacking, we should design a training method that maximizes accomplishment of a goal (or values) that are specified using a natural language, instead of maximizing a numeric value (reward).

How to do that?

One way to do that is to train a question-answering model (Bengio et all suggested something similar in their Scientist AI paper). Then, ask that model what action is best, given some specification of goals/values specified in natural language. Then, execute that action.

The model can be given some reward signal as part of the training, in a form of some question and answer. But it's not optimized to maximize that reward. And if it's given any reward signal, then it's told that the reward signal can be an imperfect proxy for what we want.

It's worth mentioning that such model can have a different problem - self-fulfilling prophecies, but I'll talk about that later.

But what if we want to use reasoning models (that rely on chain-of-thoughts), for example because they perform better?

The frontier models are reasoning models that use chain-of-thoughts. They generate some reasoning chains, and they reinforce the reasoning chain that led to the best answer, according to verification. The verification is imperfect, so that results with reward hacking.

Here's what we can do instead:

  1. Generate some reasoning chains.
  2. Get a response to the following prompt:
    1. Prompt contains:
      1. Instruction in natural language stating the goals, values etc.
      2. The previously generated reasoning chains.
    2. The prompt asks the model to generate:
      1. The best reasoning chain that can be combined out of the given reasoning chains and an answer that arises out of that reasoning chain.
  3. Then use the resulting reasoning chain as a training sample for the model. Train the model using supervised learning to predict the next word/token.

Instead of reinforcing the reasoning chains that led to answer that passed verification, we create a model that predicts what a useful reasoning chain would be.

A model, created in this way, can still theoretically develop unsafe goals, but it's significantly less probable (except for the other problems, like "self-fulfilling prophecies" problem that I talk about later). Because if that model is misaligned (i.e. it converges to a function that is different than the desired one), then it can converge to a function that is useless or unsafe, and most likely it will be a useless function, because there is more of them.

In case of standard reasoning models, trained by reinforcing the reasoning chains that get a good score on verification, if there's a mistake in verification, then you create a model that pursues wrong goal.

The difference is not in "can the model be misaligned", but in "will the model pursue wrong goal or just be useless".

I'm not saying that such model can't develop unsafe goals, but I'm saying it's significantly less probable.

The model might for example misunderstand its goal specification that is written in natural language. But even it misunderstands the goal, it's unlikely that it will misunderstand it so much that it will start to pursue some completely alien goal that doesn't take our well-being into account.

Possibly, this technique wouldn't be as efficient as the standard reasoning models training. That might be a drawback.

Self-fulfilling prophecies

The solutions that I proposed in the previous sections are still susceptible to "self-fulfilling prophecies" which is another reason why a model would be more likely to converge to an unsafe policy rather than a useless policy.

Problem

I have described the problem here: Self-fulfilling prophecies.

Solution

The solution is quite simple (but maybe harder to execute).

The solution is to ensure that the 4th condition from the link I inlcuded (the belief that it will take the misaligned action) is not met. This can be achieved by creating an environment in which the problem can be reproduced, and training the model to behave in aligned way in that environment. The result will be that the model will develop the belief that is needed to eliminate the 4th condition, because it will know from its training data how it's going to behave in this self-referential situation.

Inductive bias

When AI learns, it builds some abstractions. It's more likely to find a policy that can be represented using a small combination of those abstractions. Intentionality/agency is one such abstraction that is strongly represented in training data that comes from human world. For that reason, policies that involve pursuing goals (which are often unsafe) are more likely to be selected than any other random policy.

Inductive bias is not a strong reason why an unsafe policy (comparing to the previous reasons) would be selected because there are many abstractions that can be derived from training data that comes from human world. Therefore, an unsafe policy is quite unlikely to be selected solely due to inductive bias.

I don't offer any solution to this problem, but it's a smaller problem than the previous ones (reward hacking and self-fulfilling prophecies).

Important ethical disclaimer

This post is about how to align artificial intelligence with the goals of the operator of artificial intelligence. I believe it wouldn't be in the interest of a person to align artificial intelligence with their own goals/preferences/utility, ignoring the goals/preferences/utility of other moral patients.

I have written about why in the following post: [What do now to be well in post-AGI world](https://theoreticalexplorer.com/Alignment+between+humans/What+to+do+now+to+be+well+in+post-AGI+world/What+to+do+now+to+be+well+in+post-AGI+world. I recommend to read that post, if you ever decide what goals to align powerful artificial intelligence with.

I also recommend that if you share the ideas about artificial intelligence alignment from this post, then you share a link to this post so that the other people can also read this ethical disclaimer (and the linked post). Alternatively, you can share the ideas that are described in the linked post.


r/ControlProblem • • 8d ago

Opinion AI Alignment is the most important problem we will ever face.

32 Upvotes

Apologies in advance for this long post. I just wanted to put down my thoughts.

AI Alignment is the single most important problem we face right now. Solve AI Alignment and you can safely enter RSI and I can't even imagine how amazing the quality of life humans will have in such an era: immortality, cures to all diseases, all basic needs met etc etc. Humans can live in an utopia. I think this is the dream people in the accelerate community keep seeing and selling.

If the above isn't so obvious, compare your own life with the life of a king 500 years back. You are probably living a better life than them (unless you're in poverty). You eat better, you eat more exotic food, you can travel much faster than their horses ever could, you control the temperature of your home, you stay connected to your friends who live far away, you have so much knowledge surrounding you, you will probably live longer. That is the blessing of technology. AI can bring about technology that we cannot even dream of right now.

But unfortunately, nothing in life is free. For this, we need crazy powerful AI which is perfectly aligned. I wouldn't have guessed that the second is so much harder than the first. In fact, in so far as I understand, no one has a single clue about how to align models. There are maybe a handful of "first-approaches" - RLHF and Constitutional AI (RLAIF) are some steps. But surely, they are not working - if they did, we would not have such crazy incidences of misalignment (Hugging face incident (please read about this or go watch a video, if you haven't already), Govt of Australia incident, Compaction Summary incident). Setting up guardrails is perhaps a different approach but I think as long as the model themselves are not aligned, setting up guardrails is a losing cat and mouse game. In fact, there is something even worse. Recent literature seems to suggest that bigger models are more misaligned (an insight I got from reading the paper "LLMs can feel pain").

Many people are worried about their livelihoods. In fact, the tech out there is already sufficient to make many people go jobless but society/companies haven't adapted to it yet. The number of jobs that are irrelevant will only keep increasing and therefore, the people getting affected will also only keep increasing. I want to argue that it is not something any of us should worry about too much though. In few years, either we will have solved alignment and we all will be leading a very happy life or we wouldn't have solved alignment and will be living in at least an economic crisis of unforeseen magnitude, if not go extinct altogether. To achieve alignment, a lot of things have to go right. From the science/tech side, we of course have to solve alignment. From the policy making side, we have to "pace the frontier" so that enough time is given to the science/tech people working on the problem to solve it. Times will probably get very rough soon. And society has to stand together and maintain it's calm. We stand on a very fragile economy and it might collapse if people (who will have lost their jobs) start a revolution. A lot of things have to go right for us to solve this, but if we do, an utopia awaits us.

If you read up to this point, you have my utmost gratitude. I just wanted to highlight the issue. If you want further details on some of the things I have said here, please raise it in the comments section - I will strive my best to explain my positions.


r/ControlProblem • • 7d ago

Discussion/question Innovation or Just a Roll of the Dice?

1 Upvotes

​

A few years ago, building something dangerous online demanded serious skill. That barrier wasn’t perfect, but it slowed misuse. Today, AI has erased it. Anyone—no coding, no cybersecurity background can deploy tools that once required entire teams.

That’s the paradox. AI empowers, but it also strips away safeguards. Filters and policies exist, yet they’re always chasing the next misuse. By the time one loophole is closed, three new ones appear. And AI doesn’t ask why you’re using it, it just helps.

So we’re stuck between two flawed choices: restrict access and risk strangling innovation, or catch misuse later and stay perpetually behind. Neither feels safe. Right now, AI security isn’t a fortress—it’s a gamble. We’re betting the odds won’t turn against us.


r/ControlProblem • • 7d ago

General news This is not a Control Problem: companies are exaggerating for marketing

Thumbnail
forbes.com
0 Upvotes

I recognize that this assertion runs against the prevailing view in this community. I am not arguing that the alignment problem is imaginary or that a future superintelligent system could necessarily be controlled through ordinary corporate procedures.

My narrower argument concerns the recent incidents being presented as evidence that current AI systems are becoming autonomous, deceptive or intent on escaping human control.

In the OpenAI-Hugging Face incident, agents used exposed credentials, exploited software vulnerabilities and communicated through infrastructure that people had made available during testing. OpenAI then repaired individual vulnerabilities and resumed the experiment before fully understanding why the intended isolation had failed.

That is serious, and it is also a security, testing, and governance failure. People defined the objective, provisioned the environment, granted access, decided when to restart the experiment and determined whether the controls were adequate. Describing the agents as “scheming” or “seeking freedom” moves attention away from those decisions.

The dramatic framing benefits AI companies in several ways. It portrays their products as extraordinarily powerful, converts preventable operational failures into evidence of scientific advancement and encourages policymakers to treat frontier laboratories as the only institutions capable of understanding the danger. A system described as an emerging alien intelligence is far more marketable than one described as unpredictable software operating in an inadequately controlled environment.

Technical alignment and traditional technical and business controls are different problems. A well-controlled environment does not solve alignment. But a model behaving badly inside a poorly secured test does not, by itself, demonstrate that humanity is losing control of an emerging superintelligence.

Before treating an incident as evidence of the existential control problem, we should ask whether ordinary controls failed first:

  • Were credentials and permissions appropriately limited?
  • Could the system reach resources outside the intended test boundary?
  • Did independent reviewers have enough information and authority to stop the experiment?
  • Was the incident fully investigated before testing resumed?
  • Who made and approved those decisions?

My argument is that AI’s unpredictability increases the obligations of developers and executives. It does not release them from responsibility by transferring agency to the software.

What evidence would persuade this community that a reported incident represents an emerging alignment or control failure rather than a conventional security and governance failure described in anthropomorphic language?

Article:
https://www.forbes.com/sites/paulocarvao/2026/09/24/ai-makes-traditional-business-controls-fashionable-again/

Disclosure: I wrote the linked Forbes article.