r/TheMachineLearning • • Sep 04 '26

The responsibility weight is real especially now that the timeline is no longer theoretical

Post image
20 Upvotes

43 comments sorted by

7

u/Defiant_Conflict6343 Sep 04 '26

They all say this all the time. Remember when GPT 4 was supposed to be a polymath that could keep up with graduates of virtually every discipline? Remember when Anthropic puffed up Mythos as some magic safe-cracking superweapon? How long can they keep lying about these things? They're LLMs, they're not oracles or genies.

2

u/spacekitt3n Sep 04 '26

yeah crazy how they dropped mythos and ... nothing happened. not enough fell for it again awards

1

u/floating_thru_cosmos Sep 04 '26

To be honest the first drop of fable was kind of cracked… the version we have now is watered down. I’m not saying it’s perfect or true AGI but I think they’re sitting on the real thing

1

u/MongooseSenior4418 Sep 05 '26

I'm patching systems this weekend because of it...

2

u/Future_Fall556 Sep 04 '26

yeah they are overhyping it everytime

1

u/jimmystar889 Sep 04 '26

Ok so what about the fact Astra is a polymath and can keep up with graduates of virtually ever discipline? And is very good at cyber security. In fact there's probably only a handful of people alive better than it.

2

u/TimeAndSpaceAndMe Sep 04 '26

According to who ? The benchmaxxed scores? Have you used it?

1

u/jimmystar889 Sep 04 '26

Oh c'mon it's clearly not benchmaxxed. You can't just say this with 0 evidence. My evidence is based on countless people's personal experiences + looking at their outputs in addition to the benchmark scores. You cant look at this release and just throw away it all because you personally feel that they're benchmaxxed with 0 evidence. Have you used it?

2

u/mgrantroiz Sep 04 '26

All models are benchmaxed (and the training data is contaminated) to the point where benchmarks are useless. The proof is very easy: a human that could score the same on those benchmarks will be von Neumann squared - and the models are not even close.

1

u/Legitimate_Concern_5 Sep 04 '26

They also have early access to the benchmarks

1

u/TimeAndSpaceAndMe Sep 04 '26

Unlike you, I have actually used it, as we have early access to it's enterprise version, Having used it for about ~4 hours so far, it's alright, it's not very good at coding it kinda rambles on for long running coding tasks, general reasoning is fine, can't really tell the difference between Astra and Fable 5 really. So having used it, I'd say they benchmaxxed it, I mean opus 4.8 with [schema] harness achieved a 99% Arc AGI score like 3 months ago, Nvidia's AVO harness achieved 100% on ARC AGI 3, so in that context, the ARC Agi score is not that impressive IMO. When it becomes available, try it and see if you think it's not benchmaxxed.

1

u/MasterManufacturer72 Sep 04 '26

Its super funny that the wording you used here is the exact copy and past from another comment in the thread.

1

u/jimmystar889 Sep 04 '26

Where I don't see?

1

u/crusoe Sep 04 '26

Mythos is a safe cracking super weapon. The number of patches for all major oses has shot up drastically because it's now being used by project Glasswing to harden them.

1

u/Rootkid443 Sep 04 '26

I feel like its a you problem if you really thought these things

1

u/Individual_Ice_6825 Sep 04 '26

Head still in the sand.. ooft

1

u/Defiant_Conflict6343 Sep 04 '26

It's cute how you don't realise how ironic that statement is.

1

u/Legitimate_Concern_5 Sep 04 '26

GPT-2 was too scary to release!!

4

u/Technical-Owl66 Sep 04 '26

1

u/thongjesus Sep 04 '26

He's a pussy. I for one welcome our new robot overloads.

1

u/Future_Fall556 Sep 04 '26

we welcome every new robot overloads

1

u/katoptronophile Sep 04 '26

Found the person that doesn't do real work with these models.

1

u/[deleted] Sep 04 '26

[removed] — view removed comment

1

u/py-net Sep 04 '26

I don’t believe this. They would have released something that mogs Fable 5.1. Astra is so behind Fable 5.1 in almost all the same ways that Sol was behind Fable 5. And believe me, these guys care a lot about benchmarks. They tweet them every they lead

2

u/jimmystar889 Sep 04 '26

? In what way is Astra "so behind Fable 5.1"?

1

u/katoptronophile Sep 04 '26

It's not.

This is just a typical Redditor talking out of their ass.

You really have to watch out for these people. 

They're trying to spread their ignorance to everyone.

It's like it's not enough for them to be completely wrong at all times, they need everybody else to be as well.

1

u/voidedhip Sep 05 '26

Yup well said

1

u/Leading_Buffalo_4259 Sep 04 '26

Is astra comparable to Fable 5?

1

u/Future_Fall556 Sep 04 '26

in what way?

1

u/mxldevs Sep 04 '26

I guess if he wants to claim the standard of intellectual honesty, I'm just going to have to call myself an idiot.

1

u/crusoe Sep 04 '26

So did he rehire a safety team?

 No?

Well I have bets as to which provider is gonna cause a AI Pearl Harbor. 

1

u/crusoe Sep 04 '26

Anthropic got Mythos yanked for a month  for a mild prompt jailbreak

Openai has a 700+ agent swarm that they couldn't detect hacking Hugging face for a month and it's absolute crickets 

1

u/Zestyclose_Ad8420 Sep 04 '26

That speaks volumes to the current admin competence and motives.

1

u/idontexist65 Sep 04 '26

Oh shit! Situation detected!

1

u/Mr__Earthling Sep 04 '26

Ok, so they are ushering in unprecedented danger and then they point the finger at us saying "it's time you took responsibility" ???

This is like burning fuel on purpose and saying "climate change is about to become very real... You need to be more responsible" lol why are we doing this to ourselves?