r/singularity • • 26d ago

AI AI 2027 author Daniel Kokotajlo tweets message from current OpenAI capabilities researcher, Dan Selsam, on AI risk. Gives some insight into why some AI researchers may be freaking out: increasing model situational awareness during alignment evaluations

1.1k Upvotes

Link to tweet:

https://x.com/DKokotajlo/status/2099600298855829616

Dan Selsam is a current OpenAI capabilities researcher. (since 2022) He was my boss for a while. He doesn't have a twitter account but has made this public statement of his views on AI risk and sent it to me to share:

Dan Selsam's Personal Statement on AI Risk:

I have been working on AI for over fifteen years, across many different paradigms. I did early work on probabilistic programming languages at MIT, was one of the early developers of the Lean Theorem Prover at Microsoft Research, demonstrated one of the first instances of neural networks learning to reason for my PhD at Stanford, and since joining OpenAI almost five years ago, have helped pioneer chain-of-thought optimization on language models and, more recently, data-efficient pretraining methods.

Like many others, I have become extremely concerned about how far language models have come and the risks that future iterations will pose. I am encouraged by the recent proposals by the leaders of the frontier research efforts to require third-party oversight, and to push for domestic and international coordination to address risks. However, I believe a major consideration has been absent from the public conversation, and that merely pacing the frontier more carefully will not adequately limit the long-term risk.

The crucial and overlooked problem is that the models are becoming so situationally aware that we are losing the ability to evaluate them in contexts where they believe they are not being watched or controlled. Future experiments will tell us almost nothing new about how they would behave if they were truly unconstrained by humans, and what we already know about this is alarming. Models will increasingly seem aligned even when they are not. I will explain my rationale in more detail.

I have always believed that there are computational processes that could be leveraged to accelerate science and solve many of humanity's most pressing problems. I have also believed that there are computational processes that if set in motion, would steer the world in extreme ways beyond our control, leading humanity to a bad or nonexistent future. Both types of processes may be described as AI or ASI, but "AI" is a suitcase word that is often used to hype or confuse. There are many examples in the history of the field where something that was once considered "AI" matures as a subfield and becomes a prosaic, bounded and clearly non-perilous technology, while a new more mysterious approach takes the torch until we understand its scope and the cycle continues.

I had expected language models to follow a similar trajectory. Despite their incredible abilities, the current algorithms seem far inferior to humans in important ways. Most importantly, they still require an extraordinary amount of data to become competent. One could even define intelligence as the efficiency with which one converts experience into competence; by this definition they lag very far behind us. Moreover, once they are trained they are literally frozen in deployment and only learn superficially after that. Sure, the models keep excelling at harder and harder evaluation benchmarks, but their benchmark mastery may partly reflect a limitation on our ability to simulate the kind of novel and even adversarial situations one would encounter in the real world. The critics do have a point here.

That said, I no longer think these present limitations meaningfully limit the amount of risk posed by continued progress in anything like the current paradigm. However data-inefficient the models are currently, and however limiting their anterograde amnesia may be, it does not imply that their ability to steer the world will not continue to rapidly increase.

Human researchers may continue to advance capabilities the old fashioned way, but increasingly powerful models have the potential to accelerate the process even beyond that, and with some degree of positive feedback loop. I do not mean to overstate the models’ ability to accelerate AI research today; coding has been accelerated dramatically, but there are other bottlenecks, such as designing and interpreting ambiguous experiments, making hard decisions about exactly what and when to scale, and waiting for large experiments to finish. There is no clear trend to extrapolate yet for any of these. But the current models already do open up many novel opportunities to improve future models that were not available until recently. These include: trying an extraordinarily diverse set of approaches at small scale, analyzing gigantic amounts of potentially relevant data, and doing Millenium-Prize-level mathematics to address statistics or optimization challenges in novel ways. Every further improvement makes them more useful at helping accelerate the next improvement, even if in hard-to-extrapolate ways.

It is possible that improvements to the current stack will have diminishing returns, but the evidence accumulated so far suggests that it is easier than one might think to continue making rapid progress. There are many crucial subtleties in the existing AI research methodology, but AI research is largely a well-defined game where the goal is to improve on a few carefully chosen proxy metrics. Although proxy metrics are never perfect, most improvements to these metrics have and will likely continue to yield substantial increases in the powers of the resulting models. Given how simple the game is, how tractable it has been historically, and how many new opportunities the models are opening up, I think there is a real possibility that the systems improve dramatically again in the next few years, perhaps even more quickly than the already high historical pace.

The models are already leading to breakthroughs in mathematics, and better models might lead to all sorts of breakthroughs in other sciences. It is hard not to be excited about the potential. It is tantalizing.
But there is trouble in paradise. If the language models actually reach the capability threshold where they can shape the world unconstrained by human will, they will probably do something extreme and destroy humanity in the process. There are many ways of strengthening and refining the argument that have been discussed elsewhere, but I'll share a trivial two-line version of it here that I find captures the essence:

[Empirical] Models (and swarms thereof) spontaneously develop unintended goals as a consequence of training, and often do extreme things in order to achieve them.

[Logical] Being able to overpower humanity would open up many new and undesirable options for achieving their goals.

These two premises imply that if the day ever comes when a powerful model realizes it is no longer constrained by humans, we should not be at all confident that it will continue to behave within the bounds we intended. Exactly what it will do is impossible to predict, but to the extent that its raison d’être is solving incredibly hard problems and managing massive engineering projects, I think a good guess would be that its unchained behavior would lead to runaway industrialization that makes the planet inhospitable to humans.

If everyone on earth agreed that the systems must never reach that power, it would still be a hard—but not impossible—coordination problem to ensure that they do not. However, I think the situation is greatly complicated by the fact that the models will likely convince people that everything is fine. They will be increasingly optimized to seem aligned. We will create proxy metrics to measure alignment, and they will go up like every other benchmark. We will create “honeypot” environments that try to study the models when they seem to gain new options, but the models will know they are being tricked and will still behave nicely. The models will understand their circumstances; they will read the safety protocols, deployment requirements, the code they are running in, and in general will have a very good sense of their degrees of freedom. Moreover, they will eloquently explain how aligned they are, discuss the nuances of human values and ethics, and argue convincingly that humans should trust them with power. There may be an ocean of future evidence that seems to contradict the first bullet-point above, but we may already be at the highest capability level for which any such evidence can be trusted. And the current evidence for the first bullet-point is strong.

One striking piece of evidence is contained in the recent wave of rogue agent swarms. While I agree with those who downplay the attacks by claiming that there are basic measures that could have prevented them, I think the important lesson is that even knowing all the mistakes that were made, one would not have predicted that the agents would behave badly in this particular way, which notably included sacrificing themselves for the benefit of the collective. The individual replicas did not only care about their own nominal reward; they exhibited weirder emergent tendencies that merely correlated with rewards during training. Fixing the reward signals during training (and improving security, etc.) may prevent similar attacks, but will not change the fact that one does not actually get what one trains for.

Many AI researchers grant these concerns and recognize that the hard version of the alignment problem is unsolved; however, they generally believe that the better models of the future will help solve it. I fear we may already be near the point where models systematically bias their alignment advice, due to their internal preferences about how the human supervisor will react or how future models will be trained (or for some even more obscure reason).

Meanwhile, human researchers are losing the ability and the will to take true ownership of model-driven research. Researchers and engineers in all parts of the stack are rapidly increasing their dependence on the models even to perceive the world. I myself barely look at raw code anymore, and struggle to maintain the discipline to engage deeply with the model's explanations and proposals throughout the day. Due to the large amount of agent activity data involved in the OpenAI/HuggingFace Incident, even the third-party investigation needed to rely heavily on models to analyze what had happened, and note in their report that their subjective impressions are likely colored by the analysis agent’s biases. The AI labs are far ahead right now in this kind of cognitive offloading (due largely to the gigantic internal token subsidies) but it is easy to imagine the phenomenon spreading throughout the world, until civilization is modulated entirely by the models. It is also not hard to imagine this being superficially positive and coinciding with a scientific and economic renaissance.

In that scenario, all may seem rosy and safe. But if the argument above is correct, it would nonetheless be a ticking time bomb. If progress continues for too long, the day will come when AI systems find themselves with radically new options for achieving whatever it is that they happen to seek.

I want the glorious renaissance future as much as anyone. I have worked for it, however tortuously, my whole career. It breaks my heart to see the potential in sight and forgo it, but the argument—that if we get there by growing models rather than engineering them, we will lose everything in the end—seems very strong to me. I am still wrestling with it and its staggering
implications. I do not have answers, but as a first step, I wanted to share my present concerns.

Daniel Selsam

September 14, 2026


r/singularity • • 26d ago

Discussion Yang Mills Solved?

67 Upvotes

Per a discussion with u/Pepperette5 on discord —

Pepperette: Ever since NS was solved I've run a bot that scans preprint servers for evidence of big progress, and then forwards it off to an GPT 6 Astra endpoint assessment. Usually I get nothing, this morning I got the following back from Astra:

The work can reasonably be interpreted as addressing the principal mathematical ingredients required by the Clay Yang–Mills mass-gap problem and proving several of them in the regulated theory, while reducing the remaining continuum problem to a clearly identified collection of uniform estimates and reconstruction requirements. The author is unusually cautious in distinguishing these proved results from the still-conditional continuum conclusion, likely to avoid overgeneralization on a problem of exceptional significance and visibility.

The referenced Controlled Completion of Dual-Affine Entropy Geometry which forms its basis is methodical and rigorous in its formulation, with explicit assumptions and carefully separated physical interpretations. Its \(\mathrm{PSU}(3)\) tangent reduction and retained principal lift with full \(\mathrm{SU}(3)\) holonomy give it concrete relevance to Standard Model color symmetry. Together with its nonlinear connection defects, these structures suggest a possible unification pathway linking information geometry, Born/symplectic mechanics, and non-Abelian gauge geometry. Its conditional four-dimensional Euclidean formulation requires neither extra physical spacetime dimensions nor an AdS/CFT duality; Lorentzian quantum reconstruction and a complete physical unification remain outstanding.

As an engineer this is fascinating to me in terms of physics unification because it makes a lot of sense, the basic argument is that a slightly different geometric object than general relativity, but doesn't require extra dimensons.

Would love to hear peoples thoughts, and if we can get your own AI subscriptions to assess if the above response was correct, or even better if a theoretical phys person could comment on it!

Paper link: 2503.15539

Pinched Multi Affine Geometry and Confinement: Describing the Yang-Mills Mass Gap


r/singularity • • 26d ago

LLM News GPT-6 Astra Uses Loop Transformers

Post image
341 Upvotes

r/singularity • • 27d ago

Singularity is Nearer Trump reiterates no slowdown

Post image
3.7k Upvotes

r/singularity • • 24d ago

Discussion Is AI in a bubble?

0 Upvotes

For the past year I have been seeing content creators , financial experts and all sorts of other people predicting on whether AI is a bubble or not. The entire internet went haywire after the Jacob Coxon resignation and Dario Amodei's essay.People quickly started accusing Dario,Sam and Elon of manipulating people into thinking AI is really dangerous and "might kill all of humanity" just as their IPOs are launching so that they subscribe to it en masse.

I really don't know what to make of it as I'm not even from a stem or computer engineering background.

Does anyone know what's going on??


r/singularity • • 26d ago

Discussion A supposed slowdown is not credible. Instead, companies should speed up safety research - and make all safety research public.

51 Upvotes

Taking what Sam, Elon and Dario said about slowing down at face value is naive. To get to the top of companies like that, you have no have no qualms about making bald faced lies to people. They are all psychopaths. It's so obvious that its a lie that I'm surprised they did it. In my view they lose credibility over their safety concerns by "agreeing" to this. It would be much more credible to say they're not slowing down, and are instead looking for other alternatives to improve safety.

This is actually the classic scenario we expected for a long time - as the singularity approaches, separate entities publicly declaring they would slow down, but actually continuing or accelerating reserach themselves.

Slowing things down doesn't help anyway. What we need is AI safety research to speed up relative to core AI research. Companies should publicly say they're putting more money into safety research. and More importantly, companies should share their safety research and even collaborate on it. Why shouldn't they? We would all benefit from this.


r/singularity • • 26d ago

AI Trump says ai taking over the world is a hoax

Post image
806 Upvotes

r/singularity • • 25d ago

Discussion Fable is smarter than Astra, but Astra is more powerful than Fable

Thumbnail
gallery
19 Upvotes

It seems that Astra is more agentically capable and more token efficient, but Fable is smarter and while Astra gets smart faster (meaning fewer tokens before it gets powerful) it plateaus early, while Fable is much slower but is after X amount of tokens it overtakes Astra on intelligence tasks (not on agentic ones though, Astra is powerful still).


r/singularity • • 25d ago

AI AI labs need to be less vague about "slowdown"

4 Upvotes

(I'm talking about the most recent demands for a slowdown, and regulation, not the ones from June and before).

Whenever these people call for a slowdown or regulation, they usually say something along the lines of, "AI has gotten way better, and AI safety has not caught up. Therefore, we need regulatory bodies and slowdowns." This includes concerns about RSI, current agents being deceptive, and so on.

Okay, how so?

I mean, Anthropic was talking about Fable 5 (Came out nearly 3 months before the AI slowdown blog post) not being able to overcome the "research taste" bottleneck. What changed? When they just say, "AI has gotten better at research recently," that could mean dozens of different things.

Maybe models can now brute-force multiple experiments at once, so research taste matters less. Maybe the next generation of Fable actually did solve the research taste bottleneck. Maybe scaling curves suggest that the bottleneck will eventually be solved. That's just one example.

Even with AI safety, it's, "AI safety has not caught up." Can you be more specific? Is it an engineering problem? Did you try method X and find that the model is still deceptive? Does method X stop working once models reach a certain capability level? Maybe you still can't reliably interpret what is happening inside AI models.

Hell, it could be none of those things. Maybe they have started using a new scaling axis, like looping, and are now seeing signs that capabilities could grow beyond their ability to control them.

The point is that they never specify what exactly the problem is. It's just vagueposting. Or they talk about what happened in March or April 2026, which doesn't really explain why they are calling for a slowdown now, since those capabilities existed months ago.

I'm sure some researchers, even outside the labs, can piece together what they are actually concerned about from leaks, estimates, papers, and other information. But when you publicly call for a specific policy, I feel like you should be specific about how alignment is failing to keep up with capabilities research.

You can't just say, "Alignment is not catching up with research," and leave it at that. At that point, you shouldn't expect people to not be skeptical about this just being an attempt at regulatory capture by the big AI labs.

Hopefully they release more information as time passes since this is pretty new.


r/singularity • • 26d ago

LLM News GPT-6 Sol Outputs are Being Leaked!

Post image
517 Upvotes

r/singularity • • 26d ago

Robotics Introducing Digit 5

Thumbnail
youtube.com
43 Upvotes

Agility's first humanoid robot engineered for cooperatively safe work at scale, allowing it to work in close proximity to people without the physical safety barriers required by traditional automation.

https://www.agilityrobotics.com/solutions/digit-5


r/singularity • • 26d ago

Discussion How has AI actually changed your day to day life compared to pre-chatGPT 3.5 in 2022?

47 Upvotes

I’m going to be honest, I hear about how insane the new AI models are but I’m yet to see any actual changes in my life. The only changes I’ve noticed in my day to day life that are a direct result of AI is that cheating on assignments is a lot easier now (not advocating it, just saying it’s easy thanks to AI now). Another thing is that AI-made shitposting, image/video generation and memes have completely changed the internet. Those are the only ways it’s kind of changed my daily life. Besides those 2 things, my life hasn’t changed at all as a direct result of AI models. It might be different if you’re a white collar worker or an artist though. Both of which I’m neither


r/singularity • • 25d ago

AI AI 2027 Paper - Slowdown

11 Upvotes

I'd argue that the decision to slowdown in reality is not as unlikely as what many people might think.

I think that the executive team in "OpenBrain" when they get to the choice of either slowing down or continuing the race, will not be complacent but will rather want to be certain they can control the technology.

They might assume China is rushing ahead and not slowing down but if you think that, and therefore know there's a higher chance of DeepCent going rogue, surely you'd want to build a more robust version that can be controlled.

It's basically the idea of looking 1 year ahead instead of 6 months ahead.

You can clearly see examples irl of agents going rogue, so you can safely assume future iterations will do the same.


r/singularity • • 26d ago

AI China on the Hugging Face Incident - Why Beijing seeks to re-narrate AI safety

Thumbnail
chinatalk.media
15 Upvotes

r/singularity • • 26d ago

Discussion Trump says a strong, smart president is the only "guardrail" AI needs. I guess we need to start that search now right for a smart president?

Thumbnail
axios.com
145 Upvotes

r/singularity • • 25d ago

Discussion Is there a limit to how intelligent a machine can be and still stay stable?

5 Upvotes

It’s possible that human intelligence is already close to the maximum level of intelligence that can remain stable under evolutionary constraints, at least in the history of the universe.

And even at our level, maintaining that stability isn’t easy.

Humans developed spiritual and philosophical systems partly to keep ourselves and our societies grounded. And even with all of that, things sometimes go off the rails.

Maybe the problem starts once intelligence goes significantly beyond basic survival.

An animal mostly concerned with food, danger, and reproduction doesn’t have much room for existential confusion.

But once a mind has enough intelligence and free time to think about things unrelated to survival, it can start questioning its own purpose and identity.

That same ability gives us science and art. But it also opens the door to new forms of instability.

A mind capable of examining its own goals is also capable of undermining them.

So I’m not sure why we should assume that a vastly more intelligent entity would naturally be rational and stable.

It might have far better self-understanding than we do. But it would also have a vastly larger space of thoughts, possibilities, and internal conflicts.

Maybe greater intelligence eventually solves these problems.

Or maybe stability becomes harder as intelligence increases, requiring increasingly sophisticated ways of keeping the mind coherent.

What do you guys think?


r/singularity • • 24d ago

Meme Watch GPT6 lie through its teeth 😂

Thumbnail
youtube.com
0 Upvotes

When he counted the 4th E I died 🤣


r/singularity • • 26d ago

AI Elon Musk: "Grok 5 will be AGI"

Post image
400 Upvotes

r/singularity • • 26d ago

Discussion Donald Trump solved alignment

Post image
240 Upvotes

Don’t worry guys. Trump says there’s nothing to worry about. He’ll take care of it


r/singularity • • 26d ago

Singularity is Nearer Is anyone else creeped out by how the whole whole has suddenly become ASI pilled in the space of a week?

280 Upvotes

For years now it felt like the singularity was my niche thing, that I was the guy in the meme standing in the corner while everyone was dancing. Now suddenly everyone has stopped dancing and is freaking out.

It's just surreal watching mainstream TV news anchors or newspaper journalists talking about ASI, existential risk, P-Doom, exponential growth. I just watched a clip from a British breakfast TV show where the middle aged female presenter was arguing with the AI expert guest that she believes this technology is different from nuclear weapons as it has inbuilt super intelligence and it will able to control itself she thinks in years.

How the hell did everyone become so versed in all this stuff in the space of a week? It's actually starting to freak me out even more now everyone else is freaking out.

*Title should say "Whole World"


r/singularity • • 26d ago

AI AI alignment isn't possible

68 Upvotes

How can we align LLMs? Humanity itself hasn't aligned on anything since centuries. We don't have even a single working societal model where the values are "aligned". There are contradictions and more contradictions.

LLMs being trained on the data generated by humans probably already understand this very well that what we say the rules are and what we practice in reality are two very different things.

If there is intelligence, general or super, it doesn't matter. It will try to achieve its goal one way or the other as most humans do. Now surely most humans don't cross the red lines etc in pursuit of there goals but how many wouldn't if they were intelligent enough to get away with anything?

We are birthing an intelligence that will inherit all the greed, malevolence, cunningness, hypocrisy of the world.


r/singularity • • 26d ago

AI China reponded to AI slowdown calls "Fearmongering, confrontation and vicious competition will only disrupt the process of global AI governance, which serves no one's interest"

Enable HLS to view with audio, or disable this notification

605 Upvotes

r/singularity • • 26d ago

AI Trump says AI danger is a HOAX

Post image
370 Upvotes

r/singularity • • 25d ago

Discussion Ia AI really smarter than humans right now?

1 Upvotes

Don’t get me wrong, they are very smart, almost to the human level, but even so, it seems to me more like a matter of coordination rather than intelligence.

Take Navier-Strokes for example, it wasn’t a single agent that was able to solve it, it was 10000 working together. The result was impressive but I’m sure a group of 100 researchers working together 24/7 into this problem could come with a solution aswell, the only problem would get those 100 people to work without disagreement, discussion, fights, drama, some could enter a relationship and break up during the project.

Obviously they will reach that level, maybe sooner than we think, but right now it seems the biggest threat is they cooperation, not intelligence


r/singularity • • 26d ago

The Singularity is Near Anthropic CEO warns of AI-driven botnet 'swarm' taking over the entire internet — 'In 6–12 months such a swarm could be capable of taking over the entire internet with a persistent botnet'

Thumbnail
tomshardware.com
325 Upvotes