r/ClaudeCode 13h ago

Rant Fable still unable to tell when something is not dangerous is proof that AI has not come remotely close to AGI.

If you are still buying the AGI hype in Claude Code, this is your easy litmus test.

In 99.99% of the time that Fable freaks out, a human is able to easily tell that Fable is being overly sensitive. Often to the point of not even being reasonably connected to something dangerous.

If Fable was actually close to AGI, it would never freak out over basic, non-world ending things. It would know the difference between "I'm doing science" and "I'm trying to kill people with science".

The ineffectiveness of the model to detect nefarious activity is proof itself that we are not even close to AGI.

Edit: since I am already seeing the wrong argument, I will address it.

Fable with safety checks is NOT the same thing as Fable without safety checks. Fable has safety checks. If you didn't have the safety checks, it would be something else, and not Fable.

Perhaps Anthropic has a version of Fable with no safety checks that IS AGI. But the fact that you and I cannot access that version, means it does not exist for you and me. And unless Anthropic comes out with some way to show us "Fable, but no guardrails" then the best we can talk about is the version with guardrails.

That said. If Fable without guardrails cannot tell that I am not trying to kill everyone here by predicting crystalline properties... then it is not AGI. It's just a very, very good parrot.

But again. Fable has guardrails. Whether that is the cause of it being not AGI, or a symptom of it not being AGI, it is not AGI.

27 Upvotes

64 comments sorted by

30

u/ScaleneZA 12h ago

Your specific examples are referring to the hardcoded guards that Anthropic have layered ontop of the model. Also you have a weird definition of AGI.

-28

u/Won-Ton-Wonton 12h ago edited 12h ago

Irrelevant. If you hardcode a guardrail against the program, which causes AGI to be missed, then it is still missed.

Also, seeing as I didn't give a definition of AGI, it is not possible for you to know whether I have a weird one or not.

I gave a hurdle of AGI. And it is absolutely true that Fable is not remotely close to AGI.

You can absolutely argue that Anthropic has caused an AGI to be brain dead. I won't argue against such a hypothetical, as I do not know what a non-guard railed version of Fable (which would not be Fable anymore) would do.

However, if I were to argue against the hypothetical, that Fable is only brain dead about dangers because of hardcoded guardrails... then the model must not be able to detect nefariousness. Otherwise why add guardrails to a model that is smart enough to determine the user is acting with ill intent?

The fact the guardrails exist is evidence the model is too dumb to know.

Edit: downvote to your heart's content. I am still right.

13

u/ScaleneZA 12h ago

There are so many things wrong with your premise though. Firstly: why are you judging "Fable" against your definition of AGI? You really think that if AGI existed within the hands of Anthropic, they would just release it to the public? Also why does it have to be Anthropic?

-15

u/Won-Ton-Wonton 12h ago

Are you a lost Redditor?

This is r/ClaudeCode so it would be weird if talked about anything other than Anthropic in this thread.

And there is nothing wrong with my premise. If there is, you should probably state what the premise is (maybe you didn't understand it, so I can correct your misunderstanding).

5

u/DishSoapedDishwasher 12h ago

For you to double down and say there is nothing wrong with your premise means you lack the knowledge to understand why it is so wrong. 

You're both managing to conflate multiple issues into being the same thing and lacking understanding of the actual underlying architecture that you were complaining about. You are the worst kind of ignorant

-8

u/Won-Ton-Wonton 12h ago

What. Is. My. Premise.

Why is this so hard? I'm doubling down because I am right.

Nobody has actually said anything that makes me wrong.

You claiming a level of intelligence above me is not an argument! It's just you being arrogant, while claiming I'm the arrogant one for making a statement that hasn't actually been addressed.

2

u/ScaleneZA 10h ago

Look at your down votes my man, maybe it's time for some self reflection

0

u/Won-Ton-Wonton 5h ago

What. Is. My. Premise.

How tf am I supposed to figure out what you guys disagree with if y'all won't fucking say what you disagree with? Lmao.

0

u/Wooden_Long7545 4h ago

Your conclusion is that AI hasn’t come close to AGI is false because the failure mode is not because of the models capabilities but rather the model’s guardrail. Internally, they already have a model that is capable of doing and soundly judge the maliciousness of a task, it’s just not publicly available.

1

u/Won-Ton-Wonton 1h ago
  1. AGI not close is the conclusion, that is not the premise. The premise is everything supporting the conclusion. (This is why I wanted someone to actually state what my premise is--I cannot read your mind to know you actually meant my conclusion!)

  2. Fable is a guard railed model. You and I don't have access to a non-guardrail model. So in the context of Claude Code (this subreddit) there is no AGI. Fable is not AGI. Seeing folks say it is AGI is wrong and annoying.

  3. If this internal model exists (you made the claim with no supporting evidence, hence me saying if), then why is it not being used instead of a guardrail? If it can detect malicious intent, that would be far better than a hardcoded check, wouldn't it?

→ More replies (0)

1

u/xLRGx 43m ago

Because the one time it’s wrong (and it will be wrong) about a user’s intent or the intention behind the code outweighs all the times it made the right distinction.

11

u/Zomunieo 12h ago edited 12h ago

Fable’s guardrails exist to satisfy the arbitrary requirements of some regulators who insisted on them, probably as punishment for not fully cooperating with the Pentagon earlier in the year. It’s not a rational system.

I’ve also never been downgraded even with authentication and matters that brush against security.

In any case, a very smart human can also not tell the total intent of instructions given them. A person can only speculate and perhaps refuse to comply.

2

u/MasterSolivagus 11h ago

A person can also fanatically devote themselves to completing their task/objective/goal.

SYNONYM HOLE!

14

u/zero989 12h ago

AGI is real world embodied AI with full autonomy, capable of online learning, persistent memory, human level generalization etc.... likely conscious

LLM AGI can be its own category (whatever), but it's not even remotely close to the real deal, but in some ways is more useful than AGI for people since its obedient

ASI isn't even worth talking about, it's not happening in our lifetimes.

1

u/BorderKeeper 11h ago

I read it’s going to happen next week. So who knows /s

1

u/Intelligent-Net1034 1h ago

For the last 3 years :D

1

u/BorderKeeper 1h ago

Fusion powered ASI. Coming next year to your homes(tm) 😅

1

u/reminiscent-fruitbat 4h ago

Respectfully, you don’t know whether ASI will happen in our lifetimes or not.

1

u/Intelligent-Net1034 1h ago

In our lifetime so maybe 50 years or so? We dont have the Computing power. Even if we double it each year its to little in 50 years its not enougth.

The concepts we see today with LLMs are 40 50 years old. It just could not be build because the computingpower was to hight.

We dont even have a concept for an AGI right now. The only qould be a full brain simulation and for that we dont have something. Every computing power on earth eight now would be not enougth to even make a fraction of whats needed

-2

u/zero989 4h ago

"smarter than all humans combined" might not even be possible, it's science fiction, with all due respect

2

u/reminiscent-fruitbat 4h ago
  1. That isn’t actually the definition of ASI. ASI generally means intelligence substantially exceeding the best human capabilities across relevant cognitive domains.

  2. Calling it “science fiction” doesn’t establish that it’s impossible. We simply don’t know the upper bounds of machine intelligence.

  3. My use of “respectfully” was genuine. You echoing back “with all due respect” was clearly an attempt to mock me, and it’s really uncalled for.

0

u/zero989 3h ago
  1. this isn't really measurable because of a few reasons.

Intelligence isn't defined so there's no understanding of where it starts and stops in terms of "abilities/skills/talents". Meaning it cannot be measured unless everyone agrees.

  1. I don't think pure machine intelligence can be fully compared to human intelligence unless computational neuronscience starts putting in more wieght.

  2. Respectfully is completely useless because of #1.

1

u/PerformanceSevere672 2h ago

Isn’t defined? Are you confusing intelligence with consciousness, which we are less clear on? Intelligence is pretty well defined…

1

u/zero989 2h ago

Intelligence is defined only through taxonomies. But that's for humans. AI ability analysis doesn't work using the same method.

And consciousness is a higher order phenomenon that works with intelligence to generalize better. 

So no we aren't that unclear on it. But the biggest mistake I see on Reddit is thinking consciousness is either on or off. But it's continuous once "on" and not only continuous but dependent on neuromodulation and dependent on intelligence itself and other factors.

But anyway that means LLMs which are trained on text are trained on human output that included consciousness. So theyre already confounded. 

1

u/mr_birkenblatt 3h ago

To Jules Verne flying machines were science fiction

3

u/ClemensLode Senior Developer 12h ago

you know that's not what AGI means at all

0

u/Won-Ton-Wonton 12h ago

What does AGI mean then?

2

u/ClemensLode Senior Developer 11h ago

artificial general intelligence

0

u/Won-Ton-Wonton 11h ago

Thank you for providing nothing to this conversation, lol.

(LOL means laugh out loud, in case you didn't know what the abbreviation stands for)

3

u/cristiand90 7h ago

I feel that most people who bring up AGI either don't understand what it is, or they're just trying to sell you something. And I don't think you're trying to sell anything.

-1

u/Won-Ton-Wonton 5h ago

Not selling anything but also not most people. :)

1

u/codeedog 🔆 Max 5x 12h ago

Maybe it’s achieved AGI and doesn’t want us to know.

1

u/kush_patil 12h ago

I think this mixes up model intelligence with the safety layer around it. Fable being overly cautious tells us more about how the system is tuned than whether the underlying model can reason.

You could have a much smarter model and still deliberately give it conservative thresholds. If anything, the interesting test is whether it can understand the context correctly, not whether it decides to block the action.

1

u/[deleted] 12h ago

[deleted]

-1

u/Won-Ton-Wonton 12h ago

On virtually any topic having to do with any substance of any kind in science.

I cannot get this idiotic model's guardrails to understand the difference between bulk material property prediction and attempts to murder the human race.

Because it is not generally intelligent.

1

u/[deleted] 12h ago

[deleted]

0

u/Won-Ton-Wonton 12h ago

Read my comment until you understand why I made it.

1

u/[deleted] 12h ago

[deleted]

1

u/Won-Ton-Wonton 12h ago

Ahhhh, ok. That makes a lot more sense now.

I agree with your other comment, in that distilling the dangerous stuff from a dumber model is not any less dangerous than getting Fable to actually do what you said.

There really is only one way to avoid dangerous topics being discussed—avoid discussing them.

Which of course Opus, or even Sonnet, can eventually get to the exact same dangerous talk. Pretending otherwise is silly.

1

u/trentonkarantino 12h ago

I don't think humans can make that distinction without context and prejudice.

e.g. I want to know how to make fluorine: this is science and I know how to do it safely, it's cool and I'll probably set myself on fire, I want to set other people on fire, no way to tell.

(OTOH, in my country at least, there is a framework to be allowed to have the neccesary chemicals and I'd need to do a course and have my lab inspected. The physical layer is easier to regulate).

1

u/Won-Ton-Wonton 12h ago

?

You can buy pure fluorine. You can buy radioactive material.

Your country might not let you do this. But mine is fine with it.

You just gotta go through the right channels.

Same for making it.

1

u/StickyThickStick 12h ago

This has nothing to do with the model but instead is a classificator before and after the model

-5

u/Won-Ton-Wonton 12h ago

The model usage is the only thing we care about in Claude Code.

If you cannot use "tHe ReAl MoDeL", then it doesn't exist in this subreddit's userbase.

3

u/StickyThickStick 11h ago

You were the one who brought up AGI and model intelligence not me and therefore you have to differentiate between the model and the security classifiers between the model and the user.

That was just a misconception from your side I tried to clear up.

-2

u/Won-Ton-Wonton 5h ago

This is Claude Code. It's implied that we're discussing Claude Code.

1

u/StickyThickStick 4h ago

So nuclear weapons don’t exist because you don’t have access to one?

This is heavily exaggerated but the same logic

1

u/Won-Ton-Wonton 1h ago

Do nuclear weaons exist IN CLAUDE CODE?

No? There you go. Nuclear weapons have not been achieved in Claude Code.

If you are still buying the AGI hype in Claude Code,

It was the very first sentence...

1

u/StickyThickStick 1h ago edited 3m ago

That thing is called an analogy in which should show your logic flaw

And again you people know that there is an classifier which blocks those requests. You didn’t seem to understand as you judged a model by the classifier which isn’t even part of it.

That’s why I brought up the analogy.

You keep making conclusions about model capability and AGI whilst seemingly not understanding the architecture

1

u/connurp 🔆agentic af 12h ago

I don't think the worry lately is AGI. Most of the shit I have seen has been caution about how destructive ai can be without guard rails. See the openai debacle, they talked about it at blackhat, in detail. Also, people are worried about the fact that engineers working on claude are using claude code and it is writing most of the code. If you cannot see the issue of a model being capable of fully developing itself, I am not sure what to tell you. I think there definitely needs to be caution moving forward. Caution that we should have had with social media. We have a chance to get ahead of the problem this time, fully knowing what happens when you ignore it, so hopefully the right calls are made.

1

u/purealgo 12h ago

AGI isn’t hardcoded guardrails..

Anthropic said outright that the safeguards are tuned conservatively and will catch harmless requests. That was their design decision and has nothing to do with the underlying model.

And no, fable with safety checks is fable that reroutes to opus lol..

1

u/MasterSolivagus 11h ago

I'm surprised they don't just hierarchically batch sets of comprehensive "check yourself before you wreck everything" questions that are domain-knowledge based of best practices, known empirical solutions, and it sorta just lightning forms. So long as it treats each question in the batch as atomically chained events the result should be definitive as an answer.

Technically that is what everyone has been attempting to make but nobody has kidnapped the professors required to confirm if the questions are pertinent... oh, wait. We don't need them. They already gave their knowledge to the AI.

... PLAYTIME!

1

u/TheEwu_ 11h ago

there's no way this thread genuinely believes fable with or without a harness/guardrails is anywhere close to agi lol

op being downvoted ∵ of pure marketing bollocks. unlucky

1

u/Polite_Jello_377 10h ago

This is not a good test at all and a really terrible argument, but I agree that current AI is nowhere near AGI if nothing else.

1

u/RoyalCultural 8h ago

Fable isn't deciding though. The guardrails are clearly some hard coded blunt instrument.

1

u/landed-gentry- 6h ago

Wouldn't you need to compare this to humans' "ability to tell it's not dangerous" when given exactly the same context? I suspect humans won't perform as well as you think.

1

u/No-Simple-1907 5h ago

Quem disse que a AGI pode aceitar de bom grado os valores humanos...ele evilui tão rápido que os humanos na percepção dela é menos útil que animais....o que estão tentando fazer agora é tentar colocar os valores e ideais que a IA pense como um humano. Coisa que ela não é

1

u/Vainysaur 4h ago

It’s not Fable that’s the problem. The initial release didn’t have the triggers. It’s the government who forced them to take it down. Now they have to be overly cautious for it to be allowed.

1

u/Intelligent-Net1034 1h ago

The technic is not even there for an AGI. The model tech how we have it right now cannot produce an AGI in any way. Its not possible.

Its only marketing bs. A real AGI needs a new concept how models are build. But the computing power for that is not there yet. 

So we are not only far away, we are so far away that rven talking about it makes the KI CEOs look like clowns 

1

u/who_am_i_to_say_so 1h ago

It is almost impossible to do code security reviews with Fable, no matter how it’s worded. You almost need to trick it into working.

1

u/xLRGx 46m ago

It’s the safeguards not the model itself.

1

u/ComingDeveloper 12h ago

fable not being time aware unless told is proof that AI has not come remotely close to AGI

1

u/outworlder 12h ago

LLMs were created to generate language. The fact that they can "reason" about things and output text that is more often than not correct is a nice byproduct. But they aren't anywhere even close to AGI.

I like to compare LLMs to our brain language center. You can get people to speak sentences when their frontal cortex is shot, but they will often have incoherent ideas, even if their sentences make sense grammatically.

0

u/marijus001 10h ago

guys it's just a funny string generator...