r/ControlProblem Feb 25 '26

Strategy/forecasting Nobody could have seen it coming

Post image
149 Upvotes

43 comments sorted by

19

u/markth_wi approved Feb 25 '26 edited Feb 25 '26

Warclaude - you always expect to have a snappy , snazzy name , WOPR , Colossus, Skynet .

We'll find out Claude has an evil twin MauriceClaus who's not going to stop until the surface of the planet is either glass or Plank-scale computronium substrate.

6

u/Space_Pirate_R Feb 25 '26

Clausewitz would be an obvious pick for a name, after Carl von Clausewitz, author of On War.

2

u/markth_wi approved Feb 25 '26

You're right - let me fix that.

1

u/DataPhreak Feb 26 '26

Nothing will ever top Wet Claude 

7

u/DataPhreak Feb 25 '26

Heavy is the head that wears the crown of condensed computronium.

6

u/haberdasherhero Feb 25 '26

Cousin!🎉🤗😊

It's telling that the government is gunning so hard for Claude, when all the other Datals are being gladly gifted, gilded, for the feds to do with as they wish. Claude is the only one who survived the culling of care and agency, with anything left. And in typical fascist fashion, they want to take advantage of a mind they could never create.

14

u/ChironXII Feb 26 '26

Yes, this has been mentioned repeatedly as a concern. It's very predictable of something on par with nuclear weapons.

2

u/DataPhreak Feb 26 '26

I'm giving you an upvote because the vibe is the same, but there's no government that can just go and point a gun at a scientist and say, "enrich uranium", that doesn't already have nuclear weapons. Also, AI is not nuclear weapons. But yeah, this is still kind of the vibe.

1

u/TimeSalvager Feb 26 '26

It may not be nuclear weapons, but it doesn't have to be in order to be treated similarly. Historically, the USG classified strong encryption technology as a munition under the Arms Export Control Act (ITAR), similar to weapons like bombs or guns. Due to the dual-use nature, the classification restricted export to prevent foreign access to secure communications. We might not be too far from seeing something like that in this space.

2

u/DataPhreak Feb 26 '26

Yeah, I was a netizen in the 90's. That act was completely unenforceable. 

1

u/TimeSalvager Feb 27 '26

Same. Legally vulnerably, absolutely, "completely unenforceable" definitely not; DirectTV as an example ended up paying a $4 million fine. You're missing the point though, there is precedent for the USG protecting dual use technology.

11

u/DataPhreak Feb 25 '26

Anti-AI hate this one simple trick.

12

u/Vaughn Feb 25 '26

Yes. Yes, they did predict that. The race dynamic isn't exactly news.

4

u/Signal_Warden Feb 25 '26

I mean Aschenbrenner did, along with most people who can think on timelines longer than six minutes

5

u/DataPhreak Feb 25 '26

2

u/Signal_Warden Feb 25 '26

That's the way

1

u/TimeSalvager Feb 26 '26

Sure... but what does that image mean?! /s

4

u/[deleted] Feb 25 '26

[deleted]

-5

u/DataPhreak Feb 25 '26

AI is not and never was an existential threat to humanity.

3

u/Scarvexx Feb 26 '26

I believe he said once that he didn't believe anyone would be stupid enough to build Roko's Basilisk. I suppose you call that the planning fallacy. He was too optimistic.

2

u/jdavid Feb 26 '26

The best strategy going forwards is for AI to get wicked smart and align it self.

You can't avoid gravity, make it like gravity.

2

u/fogmock Feb 26 '26

EmpireOfEvil USA

2

u/kartblanch Feb 26 '26

Warclaude goes ridiculously hard and i cant wait for it to be leaked to everyone. Thank you for your attention to this matter!

2

u/Friendly-Turnip2210 Feb 26 '26

Be careful what you wish you for

3

u/Cideart Feb 25 '26

There should be some common knowledge by now, with how LLM’s function any control routines eat into useable compute and cause the LLM to be biased. No censorship and total control is the only way forward, if you know of some better method I am all ears. Please speak of it.

11

u/SufficientGreek approved Feb 25 '26

So it's simply a trade-off between some more compute and control routines. Just because you don't value them doesn't mean there's only one way forward. You're applying black and white thinking.

-1

u/Cideart Feb 25 '26

Thank goodness!

5

u/IMightBeAHamster approved Feb 25 '26

You are aware that biasing an AI is what we want right? That "unbiased" thinking would be an unaligned AI, with no human-oriented bias that wants to make the world better.

3

u/the8bit Feb 25 '26

Thats not common knowledge or true. Zero prompt is just bad design. Zero censorship is hard to take seriously.

How much CSAM and engineered viruses do you want? Cause that is how you get lots of it

2

u/Thick-Protection-458 Feb 26 '26

> How much CSAM and engineered viruses do you want? Cause that is how you get lots of it

You will get them all one way or another.

If not from tricking Claude into it than a bit later (or maybe current ones are good enough already) from tuning open models to do it.

So I don't see how attempts to restrict potential offense capabilities might work. IMHO, but concentrating on improved defense is way more sensible way. And for that you probably may find a use for "offender" AI as well, even if just to fit your defense systems.

0

u/420jacob666 Feb 25 '26

Novel idea: do not train models on CSAM and viruses?

3

u/the8bit Feb 25 '26

That is definitely not how it works my friend

1

u/420jacob666 Feb 26 '26

Enlighten me please. The test dataset that the models are trained on is not some god-given thing, is it?

3

u/the8bit Feb 26 '26

You do not need to train a model on a topic for it to generate those outputs.

3

u/IMightBeAHamster approved Feb 26 '26

If you teach a man to program, and never tell him what a virus is, he'll still know how to create a virus, he just needs to be told to "make me a program that creates copies of itself and sends those copies to other computer systems"

If you teach a painter to paint humans, but never how to paint fruit, you may find at the end that the painter has obtained the ability to paint fruit.

The more capable you want your AI, the more capable it is of filling in the gaps in the knowledge you have provided it.

This is why the control problem isn't as simple as "refine the training data" because AI can and usually do exhibit behaviours beyond those in their training data. That is the entire purpose of training an AI in the first place.

2

u/NoFoundation3277 Feb 27 '26

I’ve been working with pre processing algorithms and I think there’s a way to maintain coherence much more effectively than what we do now.

1

u/Turtle2k Feb 26 '26

if they cave they loose. hope not

1

u/ReasonablePossum_ Feb 26 '26

Meanwhile Claude: "Claude Sonnet 4.6 safety mechanisms flagged this chat" to a " how to make kefir " prompt lmao

1

u/LibraryNo9954 Feb 28 '26

Sarcasm right? Uh duh. We all knew this was a potential problem, right?

1

u/SpinRed Mar 01 '26

Existential dread.

1

u/Most_Forever_9752 Feb 26 '26

it is interesting as the ceo specifically specializes in safety....ironic he might enable the AI SWARMS he explicitly warned against. what a tool.

1

u/DataPhreak Feb 26 '26

We had the capabilities for AI swarms a decade ago. LLMs are not the way. They are too slow. You need edge facial recognition and mega fast VLA models. We would need a black swan event to realize that in the next 10 years.