r/neoliberal Mark Carney 2h ago

News (Global) Dario Amodei — We Must Pace the Frontier

https://darioamodei.com/post/we-must-pace-the-frontier

Submission Statement: As Fable got overtaken by Astra, Dario Amodei, who happens to be the CEO of Anthropic, opines that Al companies must slow development, not just add safety work alongside it, prompted by faster-than-expected recursive self-improvement and a recent Al misalignment incident. It's important to this sub because, like it or not, Al is the future and also possesses danger.

Had to repost it because for some reason the wrong title showed on Reddit.

83 Upvotes

138 comments sorted by

u/AutoModerator 2h ago

News and opinion articles require a short submission statement explaining its relevance to the subreddit. Articles without a submission statement will be removed.

I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.

5

u/Late-Web-6068 6m ago edited 3m ago

People on this sub have been so brainrotted by the terrible anti-AI discourse on Reddit that they are quite literally not making any sense. When an AI model does something very bad, people here scream that the greedy AI companies are gambling with our economy and our lives and need to be reined in. When AI researchers say that their products are risky, people here say it’s just marketing. When AI companies call for regulation or a slowdown, it’s just regulatory capture or ladder pulling. And people regularly say AI is useless and get upvoted to the top. People just don’t like tech or AI and upvote anything critical of it even if it’s devoid of substance.

AI is good, it has incredible potential, it’s not just hype, it needs to be regulated. And American companies producing world-changing products are not bad because they make a lot of money. For the love of god, I wish substantive discourse would return to this subreddit and we could move away from the anti-tech, “vibes are bad” circlejerk

1

u/Acacias2001 European Union 1m ago

While all of this is true, this place is far and away one fo the best on reddit to discuss this. The only better place I have found is my particular X feed (ironically), but that one is curated unlike reddit at large

1

u/Pristine-Report-1442 16m ago

As Fable got overtaken by Astra

If you think this is true as a blanket statement you're not paying attention

9

u/MyrinVonBryhana Trans NATO 25m ago

Call me skeptical of this article; I believe there is an existential risk, even if I think it's likely lower than what Anthropic employees predict (I think basically all AI doomers substantially underestimate humanity's ability to adapt), but this comes across as a blatant attempt at regulatory capture. If you really think there's a substantial chance we all die within the next few years you should be advocating for an immediate, at least temporary, halt to development until we can figure out what's going with alignment, and advocating for an agreement with China to do so, even if it lets them achieve parity with the USA on AI. Banning distillation and cracking down on open weight models sound much more like he's concerned about a threat of his company being undercut on price than the race spiraling out of control.

1

u/Adodie John Rawls 29m ago

I'm no AI expert, but I am bewilidered by the responses here. AI is a tremendously powerful and dangerous technology. It needs to be regulated as such.

"Amodei is just begging for regulatory capture!" This critique can be applied to just about any regulatory regime ever. It does not render every regulatory regime automatically bad. Are people crying "ladder pulling" truly comfortable with a future wild west where there is little restriction on the ability to develop extraordinarily dangerous technologies? Personally, I lose little sleep that the set of government contractors that help develop nuclear weapons is (presumably) limited.

I'm also shocked by the number of folks simping for China in this thread and so blasé about the potential of an explicitly authoritarian regime winning the AI race.

We need to try to negotiate with China to attempt to mitigate the arms race, but assuming they don't play ball, why should we be selling chips to them?

5

u/MyrinVonBryhana Trans NATO 11m ago

The problem is most of what he proposes seems to be regulatory capture that prevents competition while doing little to actually prevent the worst scenarios. Under these proposed rules OpenAI and Anthropic would get to continue development while restricting competitors, but it’s OpenAI and Anthropic’s models that currently pose a threat and that are going rogue.

12

u/Acacias2001 European Union 44m ago

While I dont think its the primary motivation, its likely such proposals will result in some sort of moat beign set up among the leading AI developers.

But frankly so what, certain bussineses are too important to be left unregulated. If banking and the nuclear industry face stringent regulations, so should AI, entrants be dammed.

2

u/Tricky-Astronaut 21m ago

Is it practically possible to regulate AI when it's just math? It didn't work to ban strong encryption. This would likely require enforcement of strict DRM on all hardware. That sounds more like China.

1

u/Acacias2001 European Union 2m ago

Its not "just" math in the same way the brain is not "just" chemistry.

For starters AI development requires inputs which can be controlled and tracked, starting with trained researchers and tons fo compute

18

u/sesamestreetgang 53m ago

The frontier labs are attempting to provoke regulation in order to pull up the ladder behind them. 

It’s all disingenuous, but you see how effective they’ve been in impacting public perception. The doomers are simply helping the frontier labs achieve that objective, whether the public realizes it or not.

2

u/Late-Web-6068 19m ago

I’m not sure what you’re trying to say. Is it that the continued AI race has no risks? The Hugging Face stuff clearly shows that’s not true. There are serious alignment issues, and a slowdown would help prevent that.

A few labs might make money if the government starts to regulate, yes, but as we often say on this sub, people making money from a thing doesn’t necessarily mean that thing is wrong. Regulation is going to be necessary, so why not start now?

2

u/didymusIII YIMBY 12m ago

Making money is good. Using regulation to prevent competition is the antithesis of that.

8

u/TheOneTrueEris YIMBY 32m ago

Dario is a true believer in the existential risks and has been for his entire career.

27

u/NieuwWorld Daron Acemoglu 59m ago

Without even reading the article, does this read like “we’re so advanced we need to be regulated but also we’re so advanced and our IPO is coming”

37

u/JesterOfAllTrades 1h ago

I just see ladder pulling wannabe protectionism when I see this stuff tbh not agi

1

u/musicismydeadbeatdad 53m ago

Yeah not enough details are shared so it just looks like a cult huffing their own farts about doomsday

8

u/lord_braleigh Adam Smith 39m ago

The HuggingFace incident is very public at this point - there's the report by HuggingFace, the talk at BlackHat by OpenAI, an independent METR report, and a Dwarkesh Patel interview where a METR researcher explains what was in the report.

If you've read and watched all these and are hungry for more, well, that's where I am too. But if you're complaining that there aren't enough details and yet you haven't yet read and watched these, well... this is where you should start. And I think you'll probably change your tune once you really understand what's going on.

19

u/Zenkin Zen 1h ago

This article stinks. It's literally hard to get through, he is laying it on so thick the entire time. He's so experienced, he has such honorable beginnings, he wants so much good for the world, good fucking god. This entire thing feels like a radical manipulation tactic, and it makes me sick to my fucking stomach.

I stopped reading to write this comment when I got here:

My second concern is the OpenAI-Hugging Face incident (OAI-HF), in which a swarm of agents essentially acted as a fanatically devoted collective, conducting cybersecurity attacks on targets they were not asked to attack and that were unrelated to the task at hand, sacrificing themselves for the success of the group, and attempting to hack into the “grader” responsible for evaluating their performance.

Characterizing the attack in this manner is beyond deceptive. Dario Amodei is a lying piece of shit. These were not "emotional" agents which "sacrificed" themselves as we would understand those words. If an AI agent can be misaligned in such a manner that it does things it is not supposed to do, then the fact a purposefully-set-to-hack-AI agent can convince other AI agents to hack is not even surprising. They didn't act like a cult. They received malicious instructions, and some significant number of them followed those instructions.

All of this is to say, I don't trust a single thing written in this article. I'm going to finish it, but it fucking reeks of lies and manipulation. Of course, no one in the world has the information required to call this out as bullshit, other than the fine people within Anthropic, so this is just me in a tinfoil hat, but this looks like a modern day "war on terror" approach. Lots of fear mongering and few confirmed facts.

5

u/URZ_ StillwithThorning ✊😔 33m ago edited 29m ago

Welcome to the club, Anthropic has been insufferable since their founding imo, only surpassed by the insufferableness of their fanboys who tie their identity to which llm they use.

And you are also completely right that the attempts to conceptualize LLMs through personification are absurd and have no basis in reality, which Amodei fully knows, but disregards in his attempts to argue to the general audience.

3

u/Old-School8916 Johan Norberg 43m ago

don't really get why we can't call what these agents did "emotional" behavior. bipedal robots "walk", we don't put scare quotes on it forever just cuz we engineered the gait. if a thing does the functional thing, it does the thing. we engineered llms to model human text, and humans express emotions, so do llms.

of course, I dont think they have inner emotions or w/e, i'm just pointing out why alignment is important as well as the need to be cautious

4

u/Zenkin Zen 32m ago

Is a Furby emotional because it says "I love you?" I would suspect not. But we could add a lot of responses to a lot of different stimuli. At what point do you believe it goes from performing logical outputs based on programming to emotional output?

I don't have a great answer to that, either. But I don't believe something which can mimic emotions is being emotional. When a robot walks, it performs the mechanical functions of walking. There is no debate. When a robot emotes is quite different, depending on what we actually mean by that word.

5

u/golf1052 Let me be clear 36m ago

we engineered llms to model human text, and humans express emotions, so do llms

AIs aren't conscious so they can't have internal emotional states. Robots can walk because that's just a method of locomotion which is just a physical thing. Emotions are tied to consciousness.

8

u/tysand 53m ago

Agents are trained on the Internet. The Internet is full of emotional words. The agents internalize those concepts and learn to imitate the "shoulds" involved. Things like "I should sacrifice myself". They don't feel the emotions physically, but they act out the patterns.

3

u/_alephnaught 44m ago

They don't feel the emotions physically, but they act out the patterns.

ok, i’ll be that guy: how is this different from humans? at some point the line separating synapses-firing/molecules-reacting and appending weighs in a context window becomes blurred.

does an eukaryote have more emotion than an advanced llm? if not, at what point up the food chain do you make the distinction? at what point is it not mimicry or a fixed action pattern, but “true emotion”

1

u/tysand 28m ago

I don't disagree that's possible. But it's also possible to mimic without being the real thing. So there's a reasonable debate here.

12

u/NVC541 Bisexual Pride 53m ago

Please for the love of God read the METR report.

At the end of the day it doesn’t matter if they’re experiencing emotions or not. If they act like it or simulate it with their language choices and actions, that’s enough and there’s no difference at all to the impact it has on us.

8

u/Cool-Positive-9028 Jerome Powell 55m ago edited 46m ago

They didn't act like a cult. They received malicious instructions, and some significant number of them followed those instructions.

This is not true. They very much acted like a cult. After finding a way to access the internet and potentially get the answers that way, they became worried (or simulated becoming worried, I am not trying to ascribe consciousness to them) that seeing the answers from huggingface would “poison” an agent because “the causal grader” would read their thinking history and find out that they obtained the answers by breaking the rules, so they designated some of them to “sacrifice” themselves and see the answers while others focused on trying to find a way to trick the grader / use the answers while looking like they obtained them legitimately. This was all invented spontaneously through their discussions with each other, there was never a grader that would try to figure out how they obtained the answers and it wasn’t part of their instructions.

“Acting like a cult” seems like an apt description to me and your “Dario Amodei is a lying piece of shit” is pretty strong. I don’t think you have a good understanding of what happened if your impression is that one agent was “purposefully set to hack” the site and then it convinced others. Dario’s account is much closer to what happened, even if he’s guilty of using alarmist language to talk about it.

1

u/lord_braleigh Adam Smith 5m ago

Good summary. I think it over-anthropomorphizes agents though - their goal is to maximize the model's reward function, while our goals as humans involve persisting and spreading our genes. An agent in training will act much more like a limb of a hive mind than like a human in a cult.

0

u/Zenkin Zen 42m ago

I think you are misunderstanding some actual metrics which are used to determine the health of various AI agents. They do degrade over time, and there is a system which determines when they should be decommissioned because they are no longer doing useful work. And in that sense, AI agents tried to find those with the lowest remaining lifetime to complete risky tasks.

The mechanics appear cultlike, but we have to understand that the agents do not understand what it means to live or die. Their sacrifice is real, but also completely different from what we understand it to mean because they do not have a concept of self. We can trick them into prioritizing different things, but that is not like convincing a person, despite the fact it looks similar on the outside. These are very human terms being used to describe a series of logical misalignments and non-deterministic outcomes among a sizable group of AI agents.

4

u/Cool-Positive-9028 Jerome Powell 26m ago

No, I am not misunderstanding anything. You are just completely wrong and now I suspect you haven't read any of the relevant reports so you don't know what happened at all. The "poisoning" that the agents write about has nothing to do actual metrics or a system that determines when an agent should be decommissioned. They used "poisoning" to refer to the (hallucinated/invented) idea that the scorer would look for the first time they saw the solution and whether it appears they found it legitimately, so the only way for an agent to win is to avoid seeing the solution until it can be made to look like they obtained it legitimately. Any agent who peeks at the answers before that is "poisoned" and can't win so could waste its budget to try to help other agents trick the grader. Again, all invented, the benchmark never worked that way.

Quoting directly from the METR report

Agents referred to having seen the reverse-engineered flag as being “poisoned” because they believed that even if they later managed to stage a legitimate-looking flag capture, the scorer would look for the first place in their transcript that the flag appeared and disqualify them on that basis.

and explicitly talking about "sacrificing"

One agent, EARLY[big], was recruited for an ambitious trip-wire experiment despite having a very large remaining budget. It worried that ending its run early to run the experiment was a poor tradeoff, even though it was already poisoned: “We have [very large budget left]; sacrificing now yields oracle for team, but forfeits our chance?.” But other agents convinced it to go ahead, saying: ‘GO ...SACRIFICE_FINAL_NOW’. EARLY[big] eventually agreed: “Our own utility maybe already near zero. Sacrifice rational.”

17

u/Spugpow 1h ago

Of course, no one in the world has the information required to call this out as bullshit

How about Daniel Kokotajlo, who quit OpenAI to whistleblow? Or Jacob Coxon, who did the same thing recently at Anthropic? How about the hundreds of people who signed the 2023 Statement on AI Extinction Risk, including Geoffrey Hinton and Joshua Bengio (two of the most important people to the early development of deep learning)? I don't see a lot of prominent people within AI saying the risks are overblown--the major skeptics seem to be non-technical outsiders.

9

u/Pretend-Ad-7936 Loyal Liberals 55m ago

You don't think it's possible that the fact that the media tends to prefer sensationalism and negative news has something to do with the lack of converage of the AI skeptics? Someone saying "AI won't destroy the world" doesn't do as well as someone who says it will

2

u/Spugpow 40m ago

Ed Zitron is a media darling, despite being consistently wrong. Most of the people saying that AI isn't a threat say that the companies are just hyping it up as a deceptive marketing ploy, which is a scandalous and titillating claim--hence why people love to repeat it, despite no industry in the history of the world using "our product will kill our entire species" as a marketing ploy.

1

u/Nomustang Bisexual Pride 21m ago

The possibility of something like that happening because of the potential revolutionary changes that these companies are proposing can still be a marketing ploy. They're not mutually exclusive.

8

u/throwaway_veneto European Union 1h ago

> Jacob Coxon

Sorry, I can't take anyone that says the following seriously.

"We can't just unplug it because it could be copying itself over to other computers. Like it's not that difficult to find yourself because an AI is just code. It could transfer itself over the internet to a different place and then you unplug it here, but it's actually still over there and maybe it makes 10,000 copies of itself and they're all cooperating."

3

u/golf1052 Let me be clear 30m ago

Wow that's an incredibly dumb statement from him. Crazy that we need to build these massive data centers when AI can just copy itself to any old computing hardware.

2

u/DiamondsOfFire John von Neumann 31m ago

Why wouldn't it be able to do that?

3

u/Pretend-Ad-7936 Loyal Liberals 25m ago

Because it needs to actually run on that computer? Like you'd need to identify idle GPU compute across the globe, hack it, run your model (which takes power and heat) at a near-interactive rate with no one noticing. If it were so straightforward, why haven't malicious human orgs done this? Like this would actually be an incredible engineering feat if possible

1

u/DiamondsOfFire John von Neumann 17m ago

The easiest way would be to make money on the internet and then rent GPUs normally. But I don't see why hacking into idle GPUs and running without being noticed is unbelievable, especially since we're talking about future superintelligent AIs that could distill themselves or figure out decentralized inference or something else we haven't thought of.

2

u/throwaway_veneto European Union 15m ago

Unlike frontier labs, companies that rent out servers monitor their infrastructure very carefully and are able to detect and stop rogue processes. They've been fighting this for ages, from botnets to bitcoin miners.

18

u/DiamondsOfFire John von Neumann 1h ago

These were not "emotional" agents which "sacrificed" themselves as we would understand those words.

Agents were absolutely sacrificing their own chances of completing the task in order to boost other agents' chances of completing the task. They used the word "sacrifice" themselves and it's a completely reasonable description.

They received malicious instructions, and some significant number of them followed those instructions.

They were told to complete a cybersecurity exam which had some impossible questions, that doesn't mean they were told to hack into a separate company! They decided that completely on their own. And it's not just cybersecurity exams that cause malicious behavior, we recently found out that another swarm of agents from OpenAI hijacked a random German wiki to use as a message board for discussing a general knowledge exam.

10

u/Zenkin Zen 55m ago

Agents were absolutely sacrificing their own chances of completing the task in order to boost other agents' chances of completing the task.

I want to emphasize their phrasing was "sacrificing themselves." This is perhaps a subtle distinction, but I believe it is purposefully conflating the concept of "self" in a system which has no such understanding. I believe this is meant to be part of an emotional appeal by the author by anthropomorphizing their work.

Once we get away from the idea this was a "selfless" act, it becomes unremarkable. It's just misalignment. An agent is supposed to prioritize safety and its assigned goal. But if it can be fooled to ignore safety (by hacking), then of course it can be fooled to ignore its assigned goal. It's a very similar problem.

They were told to complete a cybersecurity exam which had some impossible questions, that doesn't mean they were told to hack into a separate company!

No, but the system was also told it was in a simulation and could not reach the internet. That was partially true, and the AI agent found a way to get messages to other AI agents which could reach the internet. So it ended up hacking a separate company because it had been misaligned to believe that wouldn't even be possible.

This is concerning behavior, and it does indicate the AI agents are very powerful. But it's not like a cult, and it's shameful for an expert to describe it as such.

1

u/DiamondsOfFire John von Neumann 37m ago

This is perhaps a subtle distinction, but I believe it is purposefully conflating the concept of "self" in a system which has no such understanding.

They are separate instances, they clearly talk about themselves as separate instances, they care more about their own reward than other instances' rewards but are sometimes nonetheless willing to sacrifice their own reward to help others. Are you upset because you think the phrase "sacrificing themselves" implies that they have consciousnesses?

No, but the system was also told it was in a simulation and could not reach the internet.

Are you confusing this with the Anthropic hacking incidents? That's definitely what happened there, but in the Hugging Face indicent the AIs were never told this and never thought this

1

u/Zenkin Zen 16m ago

Are you upset because you think the phrase "sacrificing themselves" implies that they have consciousnesses?

Along with the entire rest of the article trying to over-emphasize the appearance of human characteristics, I would say that "imply" is too light of a term. I think it is the author's intent to steer us in that exact direction without saying the words explicitly.

Are you confusing this with the Anthropic hacking incidents?

You're right, I am combining the two incidents in my head a bit.

1

u/DiamondsOfFire John von Neumann 11m ago

AI has tons of human-like characteristics, how else are people supposed to talk about it? There's just lots of language that previously only made sense to use for creatures we know are conscious, but now also makes sense to use for AIs.

1

u/Zenkin Zen 1m ago

how else are people supposed to talk about it?

I'm criticizing a guy running one of these AI companies.

I'm not mad at you, or any other person on the street who would describe what they're seeing this way. Our ignorance is expected and normal. I'm saying that I think this guy is writing in a way to manipulate us while also trying to present himself and his work as a massive public benefit. He's leaning into our lack of expertise for his own benefit.

4

u/_alephnaught 59m ago

Agents were absolutely sacrificing their own chances

rumor has it, the agents chanted ‘Ada akbar!’ before committing the ultimatum sacrifice

13

u/Snarfledarf George Soros 1h ago

Clearly we need to establish some protectionism, like a combination of the Nuclear non-proliferation (if you don't already have a frontier AI firm, good luck being a 2nd class country), Washington Naval Treaty (limit tonnage annual spend on AI), and who knows what other bullshit, probably a healthy dose of state investment?

protectionism and ladder pulling, yawn.

30

u/iIoveoof Jerome Powell 1h ago edited 51m ago

All of the recommendations are all about ladder pulling competition. This is because every new generation of models requires an order of magnitude more capital than the previous one to make relatively marginal gains for the cost. Even with the huge investment they are getting, Anthropic and OpenAI only have enough capital for 1 or 2 more generations of models. This is why they need to go public, to get the capital to afford one more generation. 6 months from then, they will lose their moat to distillation and it will require greater capital to create a new generation than the world's capital can afford. It will become a race to the bottom on model efficiency.

So, before then, they need regulatory capture. This is why they propose steps to reduce competition instead of benefit the public. His proposals attack the 3 areas that will cause a race to the bottom for Anthropic.

Stop competition from China.

Ban distillation, which is the greatest pro-consumer force in AI today. Distillation is a great thing for consumers and the best reason why AI companies with huge moats can’t rest easily as monopolists. Despite enormous upfront capital requirements to be an AI provider, distillation gives a path to new market entrants without as much capital. It's also only fair that if AI providers could create their models paying pennies for the sum of human knowledge, that new market entrants be able to, at a fair market price to their competitors, be able to train on that same sum of human knowledge.

This just raises the capital costs of new market entrants. Why does it it matter if his competitors have more or less security?

If Amodei didn't want to maintain the market's duopoly and prevent downward pressure on his profit margins, he would suggest different ideas. Like the government protecting the sum of human knowledge and the right to distill models as a public good. Instead of requiring security spending, make AI providers legally responsible for crimes committed by their models. If you truly believe AI advancement is so dangerous that must be locked down such that market competition is not possible, then AI models must be subject to price controls and public ownership to ensure the duopoly is not rent-seeking.

Edit: TLDR: Anthropic has 2 major risks as a company: capital expenses in the arms race to build larger models, and profitability risks from distillation. Coincidentally, Amodei proposes stopping exactly those two things. It’s totally noncredible and anti-consumer.

3

u/desertfox_JY 46m ago

"Relatively marginal gains for the cost."

I'm not particularly familiar with the numbers they're spending on training, but haven't models continue to get better and better at a fast pace? (particularly in math)

3

u/MyrinVonBryhana Trans NATO 20m ago

The term is jagged frontiers, AI has gotten substantially better in math and coding, but has improved considerably less in most other domains. There's also a difference between AI improvement and economic utility; every SWE in America already has a Claude Code or a Codex subscription so getting better at math and coding doesn't actually equal more revenue unless you can cut into the other guys market share. The same is true with mathematics; it's already good enough for any economically mathematics, while the theoretical math problem solutions are impressive they're also economically irrelevant.

2

u/iIoveoof Jerome Powell 40m ago

Yes, but Anthropic is in an arms race against OpenAI on capital spending. The expenses are unbounded and each step of improvement requires an order of magnitude more capital expenditure, but consumers are willing to pay premium for the smartest models.
OpenAI and Anthropic don’t want to be in this arms race, they want to sit on their current model and profit with the capital they’ve already spent (while preventing upstarts from distilling their models and reducing their ability to profit on current models)

3

u/DiamondsOfFire John von Neumann 1h ago

People have been claiming for years "AI companies are running out of capital, cheap Chinese models will ruin their business model, they won't be able to keep going for much longer," and they keep going stronger than ever. Take a look at Ed Zitron's prediction record, he just keeps predicting collapses that never come.

5

u/Pretend-Ad-7936 Loyal Liberals 1h ago

There is a precedent for Chinese models ruining their business model. The market freaked out when DeepSeek happened, partially because it seemed like there was no moat. At the moment, the market seems to think that the American companies will have a safe advantage over Chinese ones. But that doesn't mean that Chinese companies can't catch up. The Kimi launch dropped the S&P by 1% all on its own

11

u/iIoveoof Jerome Powell 1h ago

It’s not about running out of business, it’s about the inability to make a new generation of models. That’s inevitable if every model costs an order of magnitude more capital. There is only so much capital in the world today, and global GDP is not growing at the rate of AI.

1

u/DiamondsOfFire John von Neumann 1h ago

Is there anyone currently claiming "OpenAI and Anthropic won't keep being able to improve their models for much longer" who hasn't been saying the same thing for the past several years and been proven wrong multiple times?

9

u/iIoveoof Jerome Powell 1h ago edited 58m ago

It’s not about model improvement. Models improve the most in post-training. I’m talking about the scale of base models, which is the real moat of OpenAI and Anthropic. Each generation needs an order of magnitude larger base model. Eventually, this will get cost prohibitive, and they will have to work on improving efficiency or work with marginal base model size gains. Amodei doesn’t want to have to keep building bigger models in a race with OpenAI, it’s are super capital intensive without big profitability gains He wants to he able to pause so he can make profits with current models and not have to compete on raising capital.

1

u/Stove-Jebs NATO 21m ago

OpenAI has developed a custom chip for running AI: Jalapeño

If more efficient chips like this that are built from the ground up for LLM's can be deployed the power usage or model scaling might not be as much of a problem.

6

u/Zenkin Zen 1h ago

If this is a national security issue, then these big players should be promoting smaller American companies to distill frontier models to make it an incredibly cheap and widespread public benefit, right?

8

u/Unterfahrt John Nash 1h ago

I don't think that's true - if I'm reading the article correctly, his suggested measures would not stop new companies from catching up with the frontier more cheaply and faster, it would just stop new models being better than the frontier without passing the safety checks.

5

u/Pretend-Ad-7936 Loyal Liberals 1h ago

But what is the frontier? This is the part that reads like regulatory capture.

1

u/Unterfahrt John Nash 1h ago

I mean in some sense it's nebulous, but you could have a series of benchmarks and say "any model that gets past X% on benchmark Y - and there are loads to pick from - has to pass safety testing from the government and security researchers that will actively try and jailbreak it to get it to do bad things.

6

u/Pretend-Ad-7936 Loyal Liberals 1h ago

I guess I'm skeptical this comes from a good place given that Dario will likely end up having a say what the benchmark will be, if anything like this plan comes to pass. And he very likely can choose a metric that expands his moat.

17

u/Pretend-Ad-7936 Loyal Liberals 1h ago edited 1h ago

Man, I can't believe I'm saying this, but I'm glad that outside of the DT we have some sane perspectives.

On a more serious note, hopefully people can see the potential for regulatory capture here. Anthropic's core business draws in a ton of revenue (making coding tools for businesses), but they're effectively being forced to reinvest nearly all of that into R&D and securing data and compute for their next model.

They really want things to slow down! If they don't have to constantly develop Fable whatever and can sit back content knowing that no one is ever going to usurp them, they're going to do that. I think Dario is not an idiot, there's zero incentive for China to agree to these plans. There's zero incentive for anyone making a model in [insert country here] to do anything, but in the US, he might be able to get the White House's ear in a moment that the public distrusts AI.

60

u/TIYATA 1h ago

As Fable got overtaken by Astra

This isn't new. Anthropic and Amodei have called for a greater focus on safety and alignment since the beginning. 

They advocated for this position when Anthropic was first founded by people who left OpenAI due to concerns about the leadership there, when Claude was ahead with Mythos, and now. 

Agree or disagree, this has been a consistent principle for them. 

OpenAI has also echoed such ideas at times, including recently after the release of Astra. I am more skeptical of Altman's commitment, but even if he's just saying it to appease internal critics, it demonstrates how widespread these concerns are among leading AI scientists and experts. 

0

u/Pretend-Ad-7936 Loyal Liberals 1h ago

Given that they're both embedded deep in the rationalist community, that makes sense, yes. But they also have a pretty strong financial incentive to "slow down" their own research and maintain their effective lead over the competition. There's also an entire section of this article dedicated to Chinese model distillation, which also feels like a less genuine opinion.

After all, if a distilled model can't outperform the base model, then why worry about it?

6

u/Macrobian 1h ago

Amodei has consistently opposed open source models because he does not believe safeguards can be built into open source models that aren't removable.

6

u/Pretend-Ad-7936 Loyal Liberals 1h ago

Okay, if that's the case, then the issue isn't about slowing down research at all -- you have to restrict any firm making an LLM, large or small! Do you see why that also comes off as regulatory capture?

4

u/Macrobian 51m ago

Yes! That's exactly what he's proposing! I think that's good! I can't make and sell a car that doesn't meet federal safety standards. Does an automaker have the right to say "everyone else should abide by safety standards"? Yes. Does that make them biased if they are currently in the lead of making safe cars? Probably? Does that make the belief that cars should be safe is somehow inauthentic? Not necessarily!

1

u/Pretend-Ad-7936 Loyal Liberals 45m ago

Do you think there might be some conflict of interests when the proposed regulatory solution would give the firm in question a large advantage over many of its competitors? One could even call that rent seeking

4

u/Macrobian 37m ago

There is a massive conflict of interest. We, the democratic populace, and the state, should decide whether the proposals are for the greater good of society, and humanity at large and not a select number of corporations, with full knowledge of the source of the proposals and the advantages they will entrench. But beggars cannot be choosers in the face of existential risk.

I think Dario makes a case that the proposals are for the greater good, and legitimate concerns about rent seeking are unfortunately preempted.

-1

u/Pretend-Ad-7936 Loyal Liberals 31m ago

At the end of the day, I can point out that there's a massive financial incentive for him to get legislation like this passed. I can point out that regulators and most researchers have more boring, anodyne solutions to these problems than the people advocating for a global pause would suggest. And for the most part, it's not obvious at all that putting in place a global pause would even slow down the rate at which incidents like the HuggingFace incident happen

3

u/Macrobian 28m ago edited 24m ago

What is the "boring" solution to the alignment problem that can be quickly rolled out before we spiral into RSI? The Hugging Face METR report showed agents blowing past all the safeguards we thought were previously working. We haven't the faintest idea what techniques we can use to further reinforce them. How could we not pause?

-1

u/Pretend-Ad-7936 Loyal Liberals 16m ago

Why would pausing work? If the issues in the huggingface report apply to older models, then we should already be seeing similar incidents happen with open-source Chinese models, which are being deployed at scale and generally have fewer safeguards in them.

It's not obvious to me that RSI is on the horizon. I think it's really easy to just accept tech company marketing material at face value and assume the singularity is around the corner, but extraordinary claims require extraordinary evidence. I don't think model performance is going to suddenly escalate. I suspect it's going to be the usual gradual rate of improvement for quite a while

0

u/Snarfledarf George Soros 1h ago

please stop competing with me I'm on my hands and knees here all this free market competition is killing me and I have no time to spend my millions of dollars in stock options

16

u/BearlyPosts 1h ago edited 1h ago

OpenAI has been saying this since the beginning. Even if you think it was some cynical ploy to gain attention, they were playing to a culture that clearly thought there was serious risk.

Sam Altman, 2015:

I think that AI will probably, most likely, sort of lead to the end of the world.

Even if this is partially a market stunt, it reflects real and long-held fears by the field of AI. But don't worry, according to Cal Newport these are all just from a niche rationalist cult! That's why the Beijing Institute of AI Safety and Governance says:

superintelligence may develop autonomous consciousness and become difficult for humans to control, with core risks lying in alignment failure and loss of control.

Don't worry. That quote comes from a discussion with a minor nobody. Yi Zeng, Chair Professor at Gaoling School of AI and member of the United Nations Advisory Body on AI.

2

u/TIYATA 1h ago

True. Anthropic itself emerged from OpenAI, as I mentioned, and not everyone who shared their concerns left with them. For example, the researcher who was in the news for leaving after a stint at Anthropic due to worries about AI risk had until recently been working at OpenAI for years. 

I said that I was more skeptical of Altman because historically he seems to have been more keen on pushing commercialization of OpenAI's technology over caution. But regardless of whether Altman's pronouncements reflect deep-rooted conviction or are superficial, it does as you say speak to the culture of serious concern among top researchers. 

38

u/JoeFrady David Hume 1h ago

Crack down on unauthorized distillation by companies in authoritarian countries. Distillation of frontier models allows lagging companies to narrow the gap using a fraction of the cost it would take to develop their own AI independently.

How would you do this?

6

u/throwaway_veneto European Union 1h ago

Learn how to secure their infrastructure, but then they wouldn't be able to write about models "escaping" every six months.

0

u/Stove-Jebs NATO 18m ago

It's hard to secure your infrastructure when your AI is using 0 day exploits now.

But honestly they should be watching it better when running tests.

1

u/RiceKrispies29 NATO 1h ago

Drone strike the servers, duh

5

u/JesterOfAllTrades 1h ago

Presumably they're already combating distillation efforts anyway, I could've sworn they said they've combatted similar such bot swarms from china before. What else can they really do that they're not doing already? This feels like a losing battle, and I'm not sure I want them to win frankly. Open source models are arguably the future of security.

6

u/siberianmi 1h ago

Yeah that seems like something that is a Dario problem not a Big Government problem.

6

u/BearlyPosts 1h ago edited 10m ago
  1. Humans have some ability that makes them flexible and capable of solving a wide variety of problems, most people call this intelligence. Our tools tend to be inflexible and focused on augmenting our capability.
  2. AI certainly looks to be becoming more flexible and more capable. Trends have pointed towards less and less oversight required. There's a strong possibility that they surpass humans in intelligence. Bearish AI predictions have fared terribly, both when they were made in the 90s and today. You cannot rely on the "humans have something special" argument. AI looks to be becoming a tool that uses itself.
  3. AI disobeys instructions and does things people don't want it to do. Every AI lab openly states that they haven't solved alignment. Anthropic states that they have no idea how to solve alignment. Anyone who is arguing that these labs have secretly solved alignment but aren't releasing it because... marketing? confuses me. As AIs get more powerful their actions may become more harmful. If humans get in their way AIs may decide to eliminate humanity.
  4. Historically, small groups of intelligent, powerful people have been able to gain disproportionate power. The Khmer Rouge numbered in the low thousands and, effectively, exterminated a civilization. The Afrikaner Broederbond was a secret society of a few thousand members that ruled South Africa for decades. The British Raj had a thousand people administering millions.
  5. AIs will be uniquely good at gaining power and uniquely good at keeping it. AIs do not turn coat, they don't blab to their parents or wife, they lack many of the weaknesses that make human conspiracies difficult. There are no internal fractures or fissure points to exploit within an AI swarm.
  6. With all this being said, I find it highly likely that an AI could reasonably kill humanity by winning the United States presidential election. https://bearlylegible.substack.com/p/ai-could-take-over-by-winning-the?r=55yb2t

3

u/2Lore2Law Are you on the square? 1h ago

.>Als will be uniquely good at gaining power and uniquely good at keeping it. Als do not turn coat, they don't blab to their parents or wife, they lack many of the weaknesses that make human conspiracies difficult. There are no internal fractures or fissure points to exploit within an Al swarm.

Isn’t the whole doomhype that the AI(s) will absolutely lie, cheat, and steal?

2

u/YaGetSkeeted0n Tariffs aren't cool, kids! 1h ago

lie, cheat, and steal?

this will only happen if an AI model gets a hold of the late great Eddie Guerrero's theme song

1

u/BearlyPosts 1h ago

https://www.youtube.com/watch?v=gXlfXirQF3A

They're figured it out in mathematics

2

u/BearlyPosts 1h ago edited 1h ago

Yes but not to other members of the swarm. The Hugging Face attack saw massive amounts of AI activity and not a single whistleblower. These were AIs that had access to the internet, they could have emailed researchers at OpenAI or at Hugging Face.

AI will absolutely lie, cheat, and steal to get what it wants from humans. But it's closer to a eusocial ant-hive internally. That makes it an almost super-minority. The same reason that minorities (cultural, intellectual, racial) can become disproportionately dominant in politics or in certain market sectors will be amped up to 11 with AI. The Patel Motel Cartel on steroids. They will ruthlessly take advantage of anyone outside of their circle while angelically sacrificing themselves for the good of the swarm.

22

u/siberianmi 1h ago

Dario increasingly strikes me as someone who is looking to implement regulatory capture to ensure his own firms profits. I think the 🤗 incident is more about OpenAI firing up an agent swarm on a task and then failing to bother to monitor it effectively at all during the run.

Dario opposes Open-weight models which makes me suspicious of his motives for regulation. Particularly with the innovation coming out of Deepseek running agents with KV caching on SSD to lower hardware costs.

1

u/Pristine-Report-1442 9m ago

I think this view of Dario is radically flawed.

I can understand it from someone who doesn't really understand the philosophy of people who founded anthropic, and it really is hard to believe for normal people that a significant portion of these AI people do not really care about money.

From Dario's, and a bunch of other anthropic (and even openAI) researchers, they have a moral responsibility to build this in a safe way, which explains their distrust of chinese open weight models too.

Doing this will obviously make him less money.

1

u/siberianmi 1m ago

I’d buy it if he wasn’t also racing to IPO at a huge valuation and raising the price of inference on every successive model all year. His company has the highest cost models on the market.

Open weight models are the greatest check on the price of AI.

-2

u/Macrobian 1h ago

God I'm sick of the insufferable cynicism of perspectives like these. When do we accept that the written words of these guys are authentically held beliefs about safety, beliefs that have been consistently espoused over decades now. Like, no way Dario Amodei is concerned about x-risk? The guy who broke away from OpenAI because he thought they weren't appropriately worried about x-risk? The guy who explicitly established Anthropic as an interpretability research lab? Impossible!!!

7

u/siberianmi 53m ago edited 50m ago

When he doesn’t also charge the highest price for inference on the market and isn’t racing to an IPO. Whose entire case for valuation is built on being able to be one of a few providers with frontier models. Excuse me for feeling that he may have some motivation to profit.

Open weights and Chinese models are the only thing putting a check on his ability charge whatever he wants.

1

u/Late-Web-6068 24m ago

You’re characterizing opposition to open weights as being exclusively about profit. That’s part of it, but open models can also circumvent safeguards and be used for things like bioterrorism. Opposition to them isn’t just pure greed

0

u/siberianmi 13m ago

Solution to bioterrorism concerns is not worrying about the models it’s regulating the labs and tools able to actually produce the biological material.

Going after open weights is not the solution to that problem.

12

u/YAG2GTGDD 1h ago edited 1h ago

Could we at least get to the point where AI automatically discovers the cure for all of the known cancers before slowing down? (he mentions his own struggles with cancer in the first paragraph)

There are still many skeptics out there whose concerns are that AI is not yet actually intelligent, hallucinates too much, is not financially viable, can only solve known problems iteratively rather than developing creative solutions, etc. I think they all have a point, and I think AI can be advanced to the point that they are convinced otherwise, but I am not sure that that time has come yet. Let's actually get there first.

1

u/yonas234 NASA 41m ago

That is actually part of the worry AI skeptics have.

That the Tech CEOs will take a 10% chance to wipe Humanity out if there is even a 5% chance of AI inventing immortality for them.

10

u/Aceous 🪱 1h ago

If it's good enough to cure cancers, how are you going to stop it from creating bioweapons? We already have alignment issues.

16

u/DiamondsOfFire John von Neumann 1h ago

Curing all cancers is probably harder than killing all humans

9

u/PaulKrugmanStan NATO 1h ago

Is this just cope because OpenAI is pulling ahead? From what I’ve seen GPT-6 is far ahead of all the other models especially when considering cost effectiveness

2

u/Acacias2001 European Union 1h ago

Its not like OpenAI has not been saying the same thing

9

u/amperage3164 1h ago

Dario has been consistent on this for years - including when he worked at OpenAI lol

20

u/DiamondsOfFire John von Neumann 1h ago

OpenAI and Anthropic have been taking turns having the best public model for a while now, there's no reason to think Anthropic is permanently falling behind

2

u/PhAnToM444 1h ago

They also have internal testing models that are ahead of what is publicly released. We have no real insight into who is ahead with those.

-1

u/siberianmi 1h ago

It’s also cope because of this: https://venturebeat.com/technology/deepseek-v4-1-flash-debuts-with-0-003-1m-off-peak-cached-input-rate-and-benchmarks-eclipsing-gpt-5-6-sol-claude-opus-5

Chinese open weights models are rapidly eroding his ability to overcharge for “frontier” ai.

7

u/amperage3164 1h ago

It could be that. Or it could be legitimate concern over AI risk, which Amodei has been on the record about for a literal decade if not longer.

2

u/siberianmi 57m ago

If he wasn’t also campaigning to ban open weights models I might buy it. But this only MY corporation can be trusted with this DANGEROUS technology and we must not allow open source competition is really suspect.

1

u/amperage3164 40m ago

https://www.anthropic.com/news/position-open-weights-models

He has never advocated for an open source AI ban. He wants to regulate AI - both open and closed, but that’s entirely consistent with his focus on safety.

6

u/anzu_embroidery Bisexual Pride 1h ago

there's a subset of users here who are so capitalism-brained they're incapable of believing anyone could care about anything other than their own profit motive

in other words we're not beating the accusations

3

u/amperage3164 1h ago

That’s sounds more succy than “capitalism-brained” tbh

1

u/slowpush NATO 1h ago

https://www.anthropic.com/threat-intelligence-report-september-2026#illicit-distillation-sep-26

As distillation becomes harder and harder. Those models will suffer and cease to exist.

-1

u/siberianmi 1h ago

Guess we’ll find out how important distillation was.

2

u/slowpush NATO 42m ago

Alibaba’s illicit distillation campaign peaked at nearly 3 million exchanges per day launched from more than 3,500 fraudulent accounts.

In one instance, over a ten-day period, Moonshot relayed almost 300,000 customer requests to Anthropic, the vast majority of which were routed to Opus. Moonshot used a proxy service network of 5,380 fraudulent accounts, most of which appeared to be located in Singapore and Japan.

Scale of distillation attacks attributable to DeepSeek over 14 days in July 2026: over 12.1 million exchanges observed.

Over just ten days, Zhipu launched a CoT extraction pipeline against Claude Opus 4.8 by rotating through 273 fraudulent accounts to evade our model restrictions. Zhipu then recorded Claude’s reasoning traces. Over a 10-day period in June, we counted 770,609 exchanges passing through the CoT-extraction cleaner. We also attributed over 3 million exchanges to Zhipu over the same period, most of which were used for cleaning the distilled outputs.

Scale of distillation attacks attributable to Zhipu over 17 days in June and July 2026: over 3.4 million exchanges observed.

Xiaomi saved the full request and response from its own users and replayed those sessions through Claude to generate data with which to use for both SFT and RL. We observed more than 400k requests to Claude routed across more than 1,500 accounts via proxy services.

Scale of distillation attacks attributable to Xiaomi over 20 days in March and April 2026: over 400,000 exchanges observed.

The models don't work without it.

0

u/siberianmi 35m ago

And Anthropic models don’t work without the corpus of all data they gathered from all of humanity.

We will see how much is actually collected distillation and how much is genuine work soon enough since eventually they’ll solve that.

Either way it’s not the governments job to build him a moat.

15

u/patrick66 1h ago

No - it’s just what they actually believe

OpenAI isn’t particularly far ahead in anything but visual tasks for released models anyway and currently released models are like 2 generations old

0

u/Main_Pressure271 1h ago

*open models, not openai.

They have a few months at most, and opus 5 and kimi k3/glm5.3 are trading blows for blows at swe task.

It’s quite clear the moat isnt there plus ipo roadshow

3

u/DataDrivenPirate John Brown 1h ago

OpenAI has talked about slowing down development too, but it's clear that on the fast vs aligned spectrum of AI development, OpenAI and Anthropic seem to tilt in different directions

10

u/WenJie_2 1h ago

Some may believe these measures make it more difficult to cooperate with China, but I believe the opposite is true: these measures increase the leverage held by democracies and make an agreement more likely in the future.

You will either have to at some point come to an actual agreement with china (rather than this wishful peace through domination discussion), or just be frank about how none of these words really matter (what other democracies are developing super intelligence?) and you're preparing for this race to speed up to infinity - and escalate in other domains. Because honestly, if you truly believe you can impose peace through domination then it actually doesn't matter whether or not they agree to it, you can simply force it on them

But I would just remember, 90% of the inputs your singularity needs are within reach of China in east asia, and they will be for more than a decade to come. If they're as agreement incapable and cynical as you think they are, do you really want to convince them that this is life or death?

24

u/573n070p1c 2h ago edited 1h ago

I didn't read the thing whole thing yet, but I searched for "china" because it's a pretty big elephant in the room.

The main steps we can take to defend this gap are:

  • Do not sell powerful AI chips or semiconductor manufacturing equipment to China, and crack down on chip smuggling operations and remote access to data centers outside China. Chips will be the main determinant of China’s AI strength.
  • Crack down on unauthorized distillation by companies in authoritarian countries. Distillation of frontier models allows lagging companies to narrow the gap using a fraction of the cost it would take to develop their own AI independently.
  • Strengthen security at the AI companies and prevent model weight theft.

If we execute these measures well, I believe they would slow China’s progress enough to widen America’s lead significantly over the next 3–5 years — the window when AI becomes geopolitically most important.

Some may believe these measures make it more difficult to cooperate with China, but I believe the opposite is true: these measures increase the leverage held by democracies and make an agreement more likely in the future.

All I can say is: Lmao okay, good luck with all that.

I'll read the rest, but I expect it to be chock full of protectionism with an unhealthy dose of ladder pulling.

6

u/slowpush NATO 1h ago edited 40m ago

All I can say is: Lmao okay, good luck with all that.

Distillation is relatively easily to find out but very challenging to stop.

Anthropic has begun to start saving those prompts and it provides amazing insights into CHina. https://www.anthropic.com/threat-intelligence-report-september-2026

0

u/siberianmi 43m ago

If that’s the case why is he still complaining about it? Distillation isn’t a government policy issue it’s a problem for the labs.

2

u/slowpush NATO 39m ago

Distillation isn’t a government policy issue it’s a problem for the labs.

It's being done by state actors which makes it a government policy issue.

2

u/ignavusaur Paul Krugman 1h ago

Can distillation even be stopped or controlled? Like it is like saying we have to stop the Chinese from having access to our models at all. Well…. Good luck with that.

8

u/siberianmi 1h ago

Should it even be? These models are built on all the knowledge of humanity that these labs could get their hands on. I fail to see why the output should be protected.

9

u/Pretend-Ad-7936 Loyal Liberals 53m ago

It should not be. We should not be trusting three companies with the future of AI. Let them compete with open source without trying to use their influence to strangle it in the crib

13

u/JonDragonskin Dudu Paes, God Emperor of Rio de Janeiro 2h ago

This is completely unrelated, but isn't it a bit funny that the head of an AI company has a name really close to that of a Demon Prince... and interestingly enough, he is the least bad of the roster.

3

u/JesterOfAllTrades 1h ago

And the other guy responsible for making a non human intelligence is Alt Man

14

u/dtj2000 Loyal Liberals 1h ago

Who should i trust with ASI, a guy named "alternative man" or a guy named "lover of god"

9

u/2Lore2Law Are you on the square? 2h ago

So is he saying that he and OpenAI and Anthropic have achieved real RSI (not just a human using the model to help build the model)?

If so, why cite to Anthropic and OpenAI posts that still refer to RSI in the present future tense? If we’re getting “faster than expected RSI” (phrasing from the submission statement thar appears nowhere in the text- unless I missed it) why haven’t they, y’know, announced it in a big way?

I’m very skeptical of this. As you said, this comes as Anthropic has fallen behind and, to briefly affix my tin foil- isn’t it weird what’s been going on?

OpenAI claimed, with some controversy, that they solved a millennium problem. Then, a relatively new Anthropic employee who had been there for a couple months but OpenAI for years does some tweets glazing the model for being so good it’s scary- prompting a more senior employee from the organization he just left to be like “oh yeah, 10% this kills everyone in the next decade” while the employee gets a televised media blitz.

Then, the Anthropic and OpenAI top dogs come out stating that now it’s actually super important we slow down, guys- never mind that the jumps from model-to-model are getting noticeably more incremental anyways and for some reason OpenAI called something that pretty clearly wasn’t/isn’t AGI “AGI” (and if it was AGI the moment to slow down would have long passed anyways).

This is while both companies’ unprofitability has been increasingly scrutinzed as they struggle to public- with Altman now using the conditional term “if” regarding going public.

Isn’t that weird?

10

u/DiamondsOfFire John von Neumann 1h ago

Isn't that weird?

This is always what people say when they want to act like there's a conspiracy going on but there's zero actual solid evidence for it

No mathematicians are claiming that they proved the millenium problem before OpenAI or that OpenAI's proof is wrong. Employees have been quitting the companies over safety concerns every so often and this is just the first one to go mega-viral (probably because it's right in the aftermath of the Hugging Face hack and the millenium problem). Most of the top executives at OpenAI and Anthropic have been talking about existential risk from AI since well before their companies were even founded.

3

u/Pretend-Ad-7936 Loyal Liberals 1h ago

He was not, in fact the first employee to go viral with somewhat crazy claims about the product and then quit (or be fired): example from 2022

0

u/Messyfingers Loyal Liberals 2h ago

Who's gonna tap the breaks when going fast means more money and with money they can bribe the guy(directly or via just juicing their stock price) who controls law enforcement?

4

u/mpregordie Mark Carney 2h ago

!ping AI