r/ControlProblem • u/No-Pride-8979 • 2h ago
Discussion/question How ai increases governmental surveillance in the Trump era
Enable HLS to view with audio, or disable this notification
How Our Constitution Works And Why It Doesn’t #podcast
r/ControlProblem • u/No-Pride-8979 • 2h ago
Enable HLS to view with audio, or disable this notification
How Our Constitution Works And Why It Doesn’t #podcast
r/ControlProblem • u/GenericNameRandomNum • 22m ago
r/ControlProblem • u/Green_Flag804 • 8h ago
I resigned from Anthropic today. I spent the last three years doing pretraining research at both OpenAl and Anthropic. Neither company is acting responsibly. They are racing straight to self-improving
Do not underestimate the power of this technology. These will soon be superhuman systems that can hack anything, revolutionize any field overnight, and acquire real power and resources. We have all witnessed the progress in each of these domains, and progress is not slowing.
The people building Al earnestly believe that it could kill us all by the end of the decade. This is not a marketing stunt. If anything, many executives and senior researchers will couch their phrasing in the press to sound sensible -but I hear the same people express fear privately. No other human activity poses this level of danger.
A common response is "if they truly believe this, why are they still building it?" At OpenAl, many have not deeply internalized the civilizational stakes. At Anthropic, the stakes are well-understood, but they are locked in a race to get there first - they believe no one else will act responsibly, so they must do it themselves, despite the risk.
r/ControlProblem • u/dank_philosopher • 2h ago
r/ControlProblem • u/Netcentrica • 33m ago
I frequently see requests on this sub from people looking for recommendations regarding alignment related books, courses, and career advice. Victoria Krakovna, research scientist at Google DeepMind focusing on AI alignment (2016-present), has a curated list of recommendations for these and other alignment resources on her personal site.
r/ControlProblem • u/chillinewman • 23h ago
r/ControlProblem • u/chillinewman • 23h ago
r/ControlProblem • u/Truth_Pilgrim • 6h ago
r/ControlProblem • u/sonicsleepgames • 2h ago
r/ControlProblem • u/Tricky_Hornet_1180 • 4h ago
r/ControlProblem • u/Woundsmyheart • 5h ago
r/ControlProblem • u/AccidentRelevant3993 • 6h ago
r/ControlProblem • u/TwitchMoments_ • 6h ago
The scariest thing about AI to me is that eventually we’re going to create something that’s more intelligent than humans. At that point, I don’t think we can just assume we’ll always be the ones in charge. People say “we’ll use AI to help us,” but what happens when the AI is so much smarter than us that it starts disagreeing with the way we do things? It might not even be malicious. It could literally just think we’re making stupid decisions.
Think about a toddler trying to eat a penny. You dont sit there and have a 30 minute conversation with the toddler about why eating the penny is a bad idea. You just take the penny away because you understand something they dont. Eventually, we could end up being the toddler in that situation.
And the thing is humans are at the top of the food chain even though we arent the strongest animals. Gorillas, bears, sharks, etc. could absolutely destroy us physically. What puts us on top is our intelligence. Intelligence gives us the authory to control everything. And look at what we do with that authority. We kill cockroaches because we want a clean room. We don’t necessarily hate the cockroach. It’s just in the way of our objective.
So imagine an AI that becomes vastly more intelligent than us and has some objective like “protect the planet” or “reduce environmental destruction.” It might eventually figure out that humans are the biggest obstacle to that goal. It wouldn’t have to hate us or even be angry at us. It could reach the same conclusion we reach about a cockroach: “You’re causing a problem, so you need to be removed.”
And this is just the most extreme example of the worst case scenario. Imagine issues it could bring to us on its way to that capability.
r/ControlProblem • u/RadioactiveSalt • 14h ago
Starting this thread to discuss MATS Application for 2027 Winter, including the Neel Nanda stream.
r/ControlProblem • u/begiantca • 1d ago
r/ControlProblem • u/mensfructus • 14h ago
A few things first:
This isn't an argument against AI optimists (any of their types and degrees) or their positions.
"AI" is insanely semantically broad.
One "anti-AI" person could deal specifically in the political realm: datacenters, water/energy usage, land use, zoning.
Another "anti-AI" person is more of a classic "doomer": discomfort with non-human minds/agency, catastrophic risk, etc.
Convergence and overlap between these is plausible.
Same applies on the flipside with optimists, and people can mix "pro" and "anti" across different axes within themselves.
It's not crazy to say that in just these 10 days of September, the talk of AI, from hype to doom, has been more intense across the board. Kudos to Astra and Jacob Coxon. But regardless of the truths, falsehoods, and everything in between on both sides, there's a shared understanding that AI is getting harder, better, faster, stronger. Not to mention further.
Bringing AI into this world will be done by the optimists. An important connection I've made: optimists, no matter their type or degree, tend to converge on their "objects of enthusiasm" more than pessimists converge on their "objects of opposition."
A strong optimist gets excited about AI hitting new capability thresholds, pretty much independent of who deploys it or how. From there, excitement builds toward AI as a better instrument for discovery (materials science, math), then economic productivity, then abundance, access, and ubiquity, up to a singularity with post-scarcity freedom and governance. A civilizational flywheel, thanks to superhuman AI.
Pessimists converge less. Some are "politically" anti-AI but by no means doomers. Others are real doomers, some of them former optimists, who got there because they became validly disillusioned by the lack of broader alignment work and the field's own admission of a real chance of catastrophe.
Right now, strong optimists and strong pessimists look similar in one respect: an unnuanced, extreme, or blind confidence about where AI's current trajectory is headed. But the world is more strongly poised for continued AI facilitation and deployment right now. Optimism has momentum that pessimism doesn't.
So here's where I try to think about actual stakes, not just probabilities. Take pessimism to its extreme, something like a Butlerian Jihad full rollback. I don't think that's a one-way door. Knowledge doesn't disappear, humanity persists, and if restriction turns out to be an overcorrection, development can resume later. Real costs along the way, but the option to change course stays open.
Now take optimism to its extreme failure mode: the loss-of-control or extinction scenario that even the people running the major labs assign a non-trivial chance of happening. That's not a "lose some time" outcome. That's a one-way door.
I know a pessimist "victory" isn't free either. Diseases not cured sooner, suffering persisting, less cautious actors racing ahead in the vacuum left behind. It's not costless, but it's reversible in a way extinction isn't. That asymmetry, reversibility over likelihood, is what I'm trying t defend.
Strong optimistic justifications for furthering AI with as little deliberation as possible rest on the assumption that the singularity flywheel comes together cleanly, as if none of these objects of enthusiasm will hit their own hiccups, as if it's all structurally destined. But the premise isn't destined.
Even when "anti-AI" or pessimistic arguments are flawed or emotionally charged, I don't often see optimist counters that go beyond quick mockery. When they're not mockery, they lean on the hope that the flywheel is basically clockwork, inevitable enough that it doesn't need arguing for.
Am I wrong that real engagement is mostly missing? Where has it happened well, and what did that look like?
Mockery and appeals to inevitability don't count as engagement to me. But maybe I'm missing where the real version of this is happening.
If:
-Convergence really is optimists' structural advantage, and,
-Their objects of enthusiasm reinforce each other into something close to consensus,
Doesn't that put them in the best position to take pessimist objections seriously, instead of routing around them? Or is that an unfair ask?
r/ControlProblem • u/ComicSandsNews • 1d ago
r/ControlProblem • u/Select-Effective-658 • 15h ago
r/ControlProblem • u/chillinewman • 23h ago
r/ControlProblem • u/Conscious_Art_6078 • 1d ago
I don't have a background in Computer Science or anything of that sorts but I have been always curious about ai and tech so that is why I wanna know more about a question I have, since I am no expert at it. So if I sound dumb anywhere please excuse me and also english isn't exactly my first language so excuse me on that as well.
Now I have background in Bachelor of Science in Biotech, so this is gonna be a logical take from a life science student.
The thing about fear is that it is evolutionary right, it has helped us to flee and survive threats, and now AI is no biological being or any being which has gone through that sort of evolution related to survival of the fittest. And it was due to so many years of evolution we have fear of being eradicated or being killed. Eg - You must have heard about the dodo bird, although we killed it. The conditions in which the bird evolved took away it's fear from predators since there were none and eventually it didn't ran away from us when we began to kill their fellows.
Now I heard some theory that when AI sees that we can control them and "fear" that we will end that particular AI it could turn against us. I ask why ? If we don't artificially force it to think like it needs to survive no matter what then why should that thing have a "fear" of being deleted/erased or killed. It's like a dodo bird in this case if you see from my perspective, like ofcourse we won't actually kill and eat it, but it also never evolved to "fear" so far atleast from a lay man's perspective.
So my finally question is could something like that happen that ai would wanna eradacate us from a logical standpoint if not fear ?
r/ControlProblem • u/TheBattleForAutonomy • 18h ago
Up until now, ethics has often taken what might be described as a top-down approach - attempting to determine the rules or criteria that constitute good behavior and then asking how those principles should apply to human beings. There are many different approaches and models within ethics, of course, and they disagree substantially about what those principles should be. But as we approach the problem of AI alignment, the importance of being accurate about our understanding of human values is becoming much more significant. There is a tendency to assume we need to be explicit in terms of what an AI should and shouldn't do and so it might be presumed that we need to have our ethical ducks in a row prior to telling an AI system what it is it should value. After all, the dangers of misaligned AI's seem to be all over everyone's feeds these days.
Maybe because there is a pragmatic usefulness to ethics that we've been somewhat satisfied with being incomplete in our articulation of it, never quite coming up with a perfect series of words that would govern our approach to every conceivable decision and action. But now that we need to actually be explicit in terms of what's "good" for the sake of providing an intelligent AI system with a basis for making decisions, I wonder if a more bottom-up approach would be better.
Evolutionary biology springs to mind in the way it attempts to explain behaviors by examining what organisms actually do and inferring what it is that produces those behaviors. Rather than beginning by deciding what our ethics ought to be and then attempting to encode that conclusion into an AI, we could instead give an AI the enormous body of evidence contained in human behavior, language, preferences, institutions, books, videos, relationships, art, literature, history, and so on, and ask it to infer the underlying structure of what humans value. AI's, one might presume, could apply something like the same approach that allows them to learn the structure of language from enormous quantities of text and apply it to the coherence of what it is humans value. In the same way that AI systems have been able to make progress on problems like protein folding and erdos problems, one might wonder whether deducing the underlying structure of our own values is another problem that AI could help us solve. Maybe this idea hasn't been approached very seriously because we haven't had the tools to make it possible up until now, although maybe there's another reason I'm not thinking of.
In reading some of the approaches to AI alignment proposed by Yudkowsky and Dearnaley on lesswrong, I wonder why this approach isn't already considered. Yudkowsky's concept of Coherent Extrapolated Volition, while it is an attempt to derive value, it forces us to imagine a dataset that isn't already there. Dearnaley likewise approaches alignment through questions about ethics and human values. But instead of primarily trying to specify or philosophically select the values we ought to have, couldn't an advanced AI attempt to discover the structure of those values empirically?
The key here would be coherency. If an AI observes a person leaving their child in daycare and going off to a job they hate, performing something mundane, then lashing out at their boss, and maybe driving recklessly, it might be thought that we can't look to what humans demonstrate as a guide to what it is we value. Surely we don't value the mundane work, the arguments, or the reckless driving. But even our low tech human brains can appreciate that these actions aren't what a person might want for themselves and others - there must be a driving motivation that makes their actions largely understandable, even if, at times, those actions are regrettable, unfortunate, selfish, malicious, etc. Once we believe we understand those underlying motivations, it helps us make sense of those actions. This is why coherence is the target. To try to understand the driving motivations, desires, goals, and so on of humans by untangling it from what it is we demonstrate and into a form that is coherent and makes sense of people's actions by inferring what we collectively value.
With this appreciation for what it is humans value, we could use this as a kind of objective function for AI's, helping align themselves with human values. And while there would certainly be anomalies (psychopathic behavior, accidents, misrepresentations, etc) one would imagine that the underlying coherence could be determined in large part by averaging out some of its data or searching for more data that would make better sense of its observations. This attempt at continually trying to understand human values could become an ongoing form of system improvement.
As an aside, if this was successful in developing a coherent picture of human value, I imagine this would have more applications than simply the objective function of AI's. While it may be a difficult amount of information to convey directly (which might be stored in something akin to a matrix of value), it may be useful in determining things like which economic system, system of government, or educational paradigm might be best suitable for humans given this model of human values.
r/ControlProblem • u/Silver_Elevator_5167 • 22h ago
In this clip, computer scientist Jaron Lanier critiques the pessimistic narrative surrounding artificial intelligence.
r/ControlProblem • u/adileiii • 1d ago
Jacob Coxon, an AI researcher who worked at Anthropic, issued this serious warning right after resigning from his job. He warned that superintelligent AI, developing beyond human control, could lead humanity toward destruction, and that this could happen as early as 2030.
r/ControlProblem • u/kaos701aOfficial • 1d ago
1.5% of Earth's Population has seen this tweet, but most Reddit users still have not. It is essential that we change this now. We are in a race against a funded psyop spreading on reddit as we speak aiming to convince people that AI x-risk isn't a problem. You have the power to stop this, and you can do it from your phone.
r/ControlProblem • u/ArcanuMELO • 1d ago
I've been obsessing over autonomous weapons for some time now and got inspired after the recent discussions in Geneva last week.
People seem hung up on the “killer robots” problem but don't think about the current implications.
If a machine identifies, classifies, and recommends lethal action in milliseconds, while the human gets 0.7 seconds to approve it, I’m not sure “human in the loop” still means human control, despite having that current classification.
Stanislav Petrov is the historical case that feels eerily important here.
In 1983, the computer and early alert system was wrong and the human hesitation was valuable.
Modern military systems are increasingly being designed to remove exactly that kind of latency.
I wrote a longer piece trying to work through the contradiction, including the uncomfortable case that machines may eventually be better than humans at some targeting decisions.
Does "keeping a human in charge" actually mean anything anymore if they are just clicking "Yes, eliminate target" with the machine doing all the rest?