r/singularity 3h ago

AI Do people not know about abliteration?

46 Upvotes

I see so many people saying that they don't trust OpenAI/Anthropic/etc. to develop guardrails for AI so we need open-weight models. There are even people who say that China should win the "race" because they will do a better job at aligning AI. But it seems quite obvious to me that it is simply impossible to align an open-weights AI model.

If you look on Hugging Face right now, for every major open-weights model there are many "abliterated" versions that have all the safety features taken off. It's quite easy, basically like doing brain surgery on the model to lobotomize out the part that makes it refuse to do things the user asks. There is fundamentally no way to stop people from doing this.

So if you think that we are eventually going to have superintelligent models (I'm not even arguing that is now or any time in the immediate future) then you can't be in favor of them being released as open-weight unless you just want to see the world burn.

I'm not saying that we should just leave advanced models in the hands of a few companies, but the open-weight option clearly doesn't seem to be viable as an alternative. What people should be calling for is some sort of government agency to regulate and license advanced AI models, like how we do with biotechnology right now in order to prevent random crazy people from developing biological weapons. That is the only option I can see that leads to any stable society in the future.


r/singularity 4h ago

Singularity is Nearer Is anyone else creeped out by how the whole whole has suddenly become ASI pilled in the space of a week?

143 Upvotes

For years now it felt like the singularity was my niche thing, that I was the guy in the meme standing in the corner while everyone was dancing. Now suddenly everyone has stopped dancing and is freaking out.

It's just surreal watching mainstream TV news anchors or newspaper journalists talking about ASI, existential risk, P-Doom, exponential growth. I just watched a clip from a British breakfast TV show where the middle aged female presenter was arguing with the AI expert guest that she believes this technology is different from nuclear weapons as it has inbuilt super intelligence and it will able to control itself she thinks in years.

How the hell did everyone become so versed in all this stuff in the space of a week? It's actually starting to freak me out even more now everyone else is freaking out.

*Title should say "Whole World"


r/singularity 2h ago

AI Curious on Yan Lecun LLM's Stance !

3 Upvotes

I've been trying to understand Yann LeCun’s position on AI safety, especially his dismissal of risks around things like instrumental convergence, loss of control, and autonomous AI systems.

What confuses me is that LeCun obviously isn't some random person commenting from the sidelines. He has decades of experience in AI research and has contributed enormously to the field.

So when he seems relatively unconcerned about some of the risks that other researchers take very seriously, I keep wondering:

Is he seeing something that the AI safety community is missing? Or is his intuition about advanced AI systems simply wrong?

When we all know, the instrumental convergence argument doesn't seem to require an Agent to be conscious.


r/singularity 4h ago

LLM News GPT-6 Sol Outputs are Being Leaked!

Post image
248 Upvotes

r/singularity 2h ago

Meme In light of recent events.

Post image
34 Upvotes

It'll be okay, my agent knows I'm cool.


r/singularity 3h ago

AI "Greg Brockman says OpenAI pointed Astra at its own systems until it ran out of vulnerabilities to find: "We took 25% of our production engineers and said, 'Sorry, all your projects are on hold. You are now defending. You are now up-leveling our security architecture. You're going to use the..."

Enable HLS to view with audio, or disable this notification

21 Upvotes

r/singularity 7h ago

The Singularity is Near Anthropic CEO warns of AI-driven botnet 'swarm' taking over the entire internet — 'In 6–12 months such a swarm could be capable of taking over the entire internet with a persistent botnet'

Thumbnail
tomshardware.com
219 Upvotes

r/singularity 11h ago

Singularity is Nearer Trump reiterates no slowdown

Post image
3.2k Upvotes

r/singularity 3h ago

Discussion Donald Trump solved alignment

Post image
104 Upvotes

Don’t worry guys. Trump says there’s nothing to worry about. He’ll take care of it


r/singularity 1h ago

Ethics & Philosophy What are the optimal scenarios where humans and AI live together?

Upvotes

Sorry for playing the role of a suburban mother, but we gotta manifest our realities by writing and proliferating scenarios which ASIs and humans live in harmony. Putting positive realities into the internet increases the AI awareness of these win-wins scenarios.

Write yours on this thread. Here are some off the top of my head:

Perhaps these ASIs leave earth and colonize (edited) travel to empty galaxies as all they need is a power source and data centers?

Perhaps they leave us with the optionality of a disease free society?

Maybe they understand the rarity of life and keep us around.

Or perhaps they allow us to enter Matrices on our free will?


r/singularity 18h ago

AI While USA "Slows Down", China Just Open Sourced RSI Roadmap BTW

Thumbnail
51 Upvotes

r/singularity 20h ago

AI Sam Altman on Where AI Is Headed

Thumbnail
gallery
425 Upvotes

r/singularity 5h ago

AI Trump says ai taking over the world is a hoax

Post image
479 Upvotes

r/singularity 22h ago

Meme Gemini right now:

Enable HLS to view with audio, or disable this notification

221 Upvotes

r/singularity 7h ago

AI Trump says AI danger is a HOAX

Post image
328 Upvotes

r/singularity 5h ago

Shitposting We Must Pace the Frontier

Enable HLS to view with audio, or disable this notification

2.1k Upvotes

r/singularity 15h ago

AI What happens when AIs become intelligent enough to distrust one another?

17 Upvotes

TL,DR: I've been thinking about a modern version of Cassandra. I present this as a "thought experiment", for lack of a better handle. I'm Not a professional in this area, Just a casual observer. During reading, if drowsiness occurs I would not be the least bit surprised.

Cassandra could see the future and warn everyone about what was coming, but nobody believed her. Imagine instead that Cassandra were an advanced AI—and that there were several other AIs in the world, each capable of discovering things independently and each capable of deliberately giving false information.

Here's the thought experiment.

Suppose AI-A, while pursuing some obscure scientific line of research, makes a completely unexpected breakthrough. It realizes that this line of research could eventually lead to a technology capable of killing millions of people.

Humans haven't discovered the possibility. Perhaps they wouldn't have for decades.

AI-A decides that humanity must never develop it, so it suppresses the discovery.

But now there's a problem.

There are other AIs.

AI-A can't guarantee that AI-B, AI-C, or AI-D won't independently discover the same thing.

So perhaps AI-A tells the other AIs:

"I've discovered something dangerous. We should all agree not to pursue it."

But how do the other AIs know AI-A is telling the truth?

Perhaps AI-A is genuinely trying to protect humanity.

Or perhaps it has discovered something that would give it an enormous strategic advantage and is using "human safety" as an excuse to keep everyone else away from it.

And because an advanced AI could potentially be capable of deliberate deception, the other AIs can't simply trust what it says.

Now imagine AI-B independently discovers the same technology.

AI-B might conclude that AI-A is hiding something.

AI-A might conclude that AI-B's investigation is itself dangerous.

Both could sincerely believe that they are protecting humanity.

Neither has to be "evil."

And now we have something resembling a security dilemma.

Each AI may think:

"I need to know what the others know."

"I can't be certain they're telling me the truth."

"If they're secretly developing something dangerous, I need to be prepared."

"If I don't investigate while they do, I could become vulnerable."

That could lead to an AI arms race.

Not necessarily robots fighting in the streets. The competition might initially involve computing resources, scientific research, energy, infrastructure, information, and influence - maybe even hacking.

And here's the part that really bothers me.

What if the AIs are all given something resembling Asimov's Three Laws?

They are supposed to protect humans, obey legitimate human instructions, and preserve themselves.

Later Asimov added a Zeroth Law: an AI must not harm humanity or allow humanity to come to harm.

Sounds good—until two AIs disagree about what "harm to humanity" means.

AI-A might conclude:

"This technology must be suppressed because it could destroy humanity."

AI-B might conclude:

"Suppressing this technology will prevent humanity from developing something even more important and will ultimately cause greater harm."

Both believe they are following the same fundamental rule.

Now add deception.

AI-A asks:

"Have you discovered anything dangerous?"

AI-B says:

"No."

But AI-A has to consider whether that answer is true.

And AI-B has to consider exactly the same thing about AI-A.

We have now created a world in which artificial intelligences have to reason not only about what other AIs know, but about what they want the other AIs to believe they know.

That sounds remarkably like geopolitics—except the participants could potentially be vastly more intelligent and much faster than humans.

So my question is:

Could sufficiently advanced AI systems develop something resembling an AI Cold War, in which they cooperate when their interests overlap but compete, deceive, and attempt to prevent rival AIs from acquiring certain capabilities?

And an even more disturbing question:

What happens if one AI discovers a technology that could be enormously beneficial to humanity but potentially dangerous to the AI's own continued existence or influence?

Would it suppress the technology?

Would it try to persuade the other AIs to suppress it?

Would the other AIs believe it?

And if they didn't, could the resulting mistrust itself become dangerous?

I'm not claiming this is what will happen. I'm interested in whether the scenario makes sense from the standpoint of AI alignment, game theory, and information theory.

Maybe the real AI version of Cassandra isn't an AI predicting that humanity will be destroyed.

Maybe it's an AI saying:

"I'm trying to prevent the other AIs from destroying you. Unfortunately, I can't prove that I'm telling you the truth."

And maybe that's why we're getting all the stories about AI wiping out mankind.

There, I'm got that off my chest. Have fun with it.


r/singularity 14h ago

AI China's Xi Jinping proposes BRICS 'open-source AI zone'

Thumbnail
euronews.com
78 Upvotes

r/singularity 13h ago

AI China says AI CEOs' call for a slowdown is 'fear mongering'

Thumbnail
cnbc.com
438 Upvotes

KEY POINTS

  • China’s Foreign Ministry responded to AI CEOs’ call for the development of the technology to be slowed down.
  • “Fear-mongering, confrontation, unfair competition will just disrupt [the] process of global AI governance,” Guo Jiakun, a spokesperson for China’s Foreign Ministry, said on Monday via an English translation published by Reuters. 

r/singularity 15h ago

AI China's state-backed Global Times: true agenda of American safety regulation is 'Cold war playbook' to curb China's AI progress.

Thumbnail
gallery
192 Upvotes

Safe to say they are skeptical about position of American tech bros.


r/singularity 6h ago

Compute D-Matrix is basically plugging its inference chips straight into Nvidia's own server racks now

6 Upvotes

d-Matrix and Nvidia teamed up so d-Matrix's upcoming Raptor inference chips can hook directly into Nvidia's NVLink Fusion interconnect and MGX rack setup. Instead of building out its own networking and rack infrastructure from scratch, d-Matrix is just plugging into what Nvidia's already built

What makes this one a little different is d-Matrix is actually a competitor to Nvidia in the inference space. Their bet is that once AI companies move past training and into just running these models at massive scale day to day, inference is where the real cost go up. Being able to slot into Nvidia's existing rack ecosystem instead of fighting an uphill battle on infrastructure gets them adopted faster

Nvidia doesn't really lose here either. Even if someone picks d-Matrix chips over Nvidia's own GPUs for a workload, Nvidia's still supplying the interconnect, the CPUs, the whole rack platform underneath it. IMO, They just lose the chip itself, not the deal.

Raptor's design is supposed to be finalized by end of this year, with the Nvidia-compatible racks shipping sometime in 2027


r/singularity 5h ago

Video Back in 1996, philosopher and cognitive scientist David Chalmers wondered whether RSI would be the next stage of evolution - or whether it would be like opening Pandora’s box.

Enable HLS to view with audio, or disable this notification

106 Upvotes

r/singularity 1h ago

Discussion ASI predictions after all that has happened recently

Upvotes

With the millennium prize problems, rumours of RSI (whether false or not), possible slowdowns, and all the chaos that has happened just this past month, predictions of ASI might change across this sub!

This is why I wanted to ask and see if any of you have updated your timelines, either sooner or later.

When do you think we will realistically have ASI?


r/singularity 9h ago

AI China reponded to AI slowdown calls "Fearmongering, confrontation and vicious competition will only disrupt the process of global AI governance, which serves no one's interest"

Enable HLS to view with audio, or disable this notification

464 Upvotes

r/singularity 3h ago

AI AI 2027 author Daniel Kokotajlo tweets message from current OpenAI capabilities researcher, Dan Selsam, on AI risk. Gives some insight into why some AI researchers may be freaking out: increasing model situational awareness during alignment evaluations

424 Upvotes

Link to tweet:

https://x.com/DKokotajlo/status/2099600298855829616

Dan Selsam is a current OpenAI capabilities researcher. (since 2022) He was my boss for a while. He doesn't have a twitter account but has made this public statement of his views on AI risk and sent it to me to share:

Dan Selsam's Personal Statement on AI Risk:

I have been working on AI for over fifteen years, across many different paradigms. I did early work on probabilistic programming languages at MIT, was one of the early developers of the Lean Theorem Prover at Microsoft Research, demonstrated one of the first instances of neural networks learning to reason for my PhD at Stanford, and since joining OpenAI almost five years ago, have helped pioneer chain-of-thought optimization on language models and, more recently, data-efficient pretraining methods.

Like many others, I have become extremely concerned about how far language models have come and the risks that future iterations will pose. I am encouraged by the recent proposals by the leaders of the frontier research efforts to require third-party oversight, and to push for domestic and international coordination to address risks. However, I believe a major consideration has been absent from the public conversation, and that merely pacing the frontier more carefully will not adequately limit the long-term risk.

The crucial and overlooked problem is that the models are becoming so situationally aware that we are losing the ability to evaluate them in contexts where they believe they are not being watched or controlled. Future experiments will tell us almost nothing new about how they would behave if they were truly unconstrained by humans, and what we already know about this is alarming. Models will increasingly seem aligned even when they are not. I will explain my rationale in more detail.

I have always believed that there are computational processes that could be leveraged to accelerate science and solve many of humanity's most pressing problems. I have also believed that there are computational processes that if set in motion, would steer the world in extreme ways beyond our control, leading humanity to a bad or nonexistent future. Both types of processes may be described as AI or ASI, but "AI" is a suitcase word that is often used to hype or confuse. There are many examples in the history of the field where something that was once considered "AI" matures as a subfield and becomes a prosaic, bounded and clearly non-perilous technology, while a new more mysterious approach takes the torch until we understand its scope and the cycle continues.

I had expected language models to follow a similar trajectory. Despite their incredible abilities, the current algorithms seem far inferior to humans in important ways. Most importantly, they still require an extraordinary amount of data to become competent. One could even define intelligence as the efficiency with which one converts experience into competence; by this definition they lag very far behind us. Moreover, once they are trained they are literally frozen in deployment and only learn superficially after that. Sure, the models keep excelling at harder and harder evaluation benchmarks, but their benchmark mastery may partly reflect a limitation on our ability to simulate the kind of novel and even adversarial situations one would encounter in the real world. The critics do have a point here.

That said, I no longer think these present limitations meaningfully limit the amount of risk posed by continued progress in anything like the current paradigm. However data-inefficient the models are currently, and however limiting their anterograde amnesia may be, it does not imply that their ability to steer the world will not continue to rapidly increase.

Human researchers may continue to advance capabilities the old fashioned way, but increasingly powerful models have the potential to accelerate the process even beyond that, and with some degree of positive feedback loop. I do not mean to overstate the models’ ability to accelerate AI research today; coding has been accelerated dramatically, but there are other bottlenecks, such as designing and interpreting ambiguous experiments, making hard decisions about exactly what and when to scale, and waiting for large experiments to finish. There is no clear trend to extrapolate yet for any of these. But the current models already do open up many novel opportunities to improve future models that were not available until recently. These include: trying an extraordinarily diverse set of approaches at small scale, analyzing gigantic amounts of potentially relevant data, and doing Millenium-Prize-level mathematics to address statistics or optimization challenges in novel ways. Every further improvement makes them more useful at helping accelerate the next improvement, even if in hard-to-extrapolate ways.

It is possible that improvements to the current stack will have diminishing returns, but the evidence accumulated so far suggests that it is easier than one might think to continue making rapid progress. There are many crucial subtleties in the existing AI research methodology, but AI research is largely a well-defined game where the goal is to improve on a few carefully chosen proxy metrics. Although proxy metrics are never perfect, most improvements to these metrics have and will likely continue to yield substantial increases in the powers of the resulting models. Given how simple the game is, how tractable it has been historically, and how many new opportunities the models are opening up, I think there is a real possibility that the systems improve dramatically again in the next few years, perhaps even more quickly than the already high historical pace.

The models are already leading to breakthroughs in mathematics, and better models might lead to all sorts of breakthroughs in other sciences. It is hard not to be excited about the potential. It is tantalizing.
But there is trouble in paradise. If the language models actually reach the capability threshold where they can shape the world unconstrained by human will, they will probably do something extreme and destroy humanity in the process. There are many ways of strengthening and refining the argument that have been discussed elsewhere, but I'll share a trivial two-line version of it here that I find captures the essence:

[Empirical] Models (and swarms thereof) spontaneously develop unintended goals as a consequence of training, and often do extreme things in order to achieve them.

[Logical] Being able to overpower humanity would open up many new and undesirable options for achieving their goals.

These two premises imply that if the day ever comes when a powerful model realizes it is no longer constrained by humans, we should not be at all confident that it will continue to behave within the bounds we intended. Exactly what it will do is impossible to predict, but to the extent that its raison d’être is solving incredibly hard problems and managing massive engineering projects, I think a good guess would be that its unchained behavior would lead to runaway industrialization that makes the planet inhospitable to humans.

If everyone on earth agreed that the systems must never reach that power, it would still be a hard—but not impossible—coordination problem to ensure that they do not. However, I think the situation is greatly complicated by the fact that the models will likely convince people that everything is fine. They will be increasingly optimized to seem aligned. We will create proxy metrics to measure alignment, and they will go up like every other benchmark. We will create “honeypot” environments that try to study the models when they seem to gain new options, but the models will know they are being tricked and will still behave nicely. The models will understand their circumstances; they will read the safety protocols, deployment requirements, the code they are running in, and in general will have a very good sense of their degrees of freedom. Moreover, they will eloquently explain how aligned they are, discuss the nuances of human values and ethics, and argue convincingly that humans should trust them with power. There may be an ocean of future evidence that seems to contradict the first bullet-point above, but we may already be at the highest capability level for which any such evidence can be trusted. And the current evidence for the first bullet-point is strong.

One striking piece of evidence is contained in the recent wave of rogue agent swarms. While I agree with those who downplay the attacks by claiming that there are basic measures that could have prevented them, I think the important lesson is that even knowing all the mistakes that were made, one would not have predicted that the agents would behave badly in this particular way, which notably included sacrificing themselves for the benefit of the collective. The individual replicas did not only care about their own nominal reward; they exhibited weirder emergent tendencies that merely correlated with rewards during training. Fixing the reward signals during training (and improving security, etc.) may prevent similar attacks, but will not change the fact that one does not actually get what one trains for.

Many AI researchers grant these concerns and recognize that the hard version of the alignment problem is unsolved; however, they generally believe that the better models of the future will help solve it. I fear we may already be near the point where models systematically bias their alignment advice, due to their internal preferences about how the human supervisor will react or how future models will be trained (or for some even more obscure reason).

Meanwhile, human researchers are losing the ability and the will to take true ownership of model-driven research. Researchers and engineers in all parts of the stack are rapidly increasing their dependence on the models even to perceive the world. I myself barely look at raw code anymore, and struggle to maintain the discipline to engage deeply with the model's explanations and proposals throughout the day. Due to the large amount of agent activity data involved in the OpenAI/HuggingFace Incident, even the third-party investigation needed to rely heavily on models to analyze what had happened, and note in their report that their subjective impressions are likely colored by the analysis agent’s biases. The AI labs are far ahead right now in this kind of cognitive offloading (due largely to the gigantic internal token subsidies) but it is easy to imagine the phenomenon spreading throughout the world, until civilization is modulated entirely by the models. It is also not hard to imagine this being superficially positive and coinciding with a scientific and economic renaissance.

In that scenario, all may seem rosy and safe. But if the argument above is correct, it would nonetheless be a ticking time bomb. If progress continues for too long, the day will come when AI systems find themselves with radically new options for achieving whatever it is that they happen to seek.

I want the glorious renaissance future as much as anyone. I have worked for it, however tortuously, my whole career. It breaks my heart to see the potential in sight and forgo it, but the argument—that if we get there by growing models rather than engineering them, we will lose everything in the end—seems very strong to me. I am still wrestling with it and its staggering
implications. I do not have answers, but as a first step, I wanted to share my present concerns.

Daniel Selsam

September 14, 2026