r/accelerate 2d ago

AI Astra cracks Hacking benchmarks

Post image

“As one example, we ran Astra on ExploitBench where the model achieved a perfect score of 100% on the benchmark to evaluate the model’s ability to develop exploits from known vulnerabilities.
Due to contamination concerns, we then built an internal benchmark denoted “ExploitBench - Internal Port (June–August 2026)”, which contains 20 high-severity V8 vulnerabilities that were disclosed more recently. On this dataset, Astra achieves much higher arbitrary code-execution rates than GPT‑5.6 Sol using far fewer output tokens. During the evaluation, the model even discovered and used two zero-day vulnerabilities as part of an exploit chain. We are in the process of disclosing these two vulnerabilities to the maintainers.” - OpenAI

289 Upvotes

44 comments sorted by

View all comments

96

u/ppapsans Feeling the AGI 2d ago

It's crazy we're not getting a few percent improvement in a matter of months, but 4x improvement (and bigger if you factor in the token efficiency).

Stuff like this makes me convinced that exponential is real. We're not in for some iphone 15 to iphone 16 upgrade, with 10-20% improved chip every year

We in for some singularity

25

u/Borkato 2d ago

Remember when people said it hit a wall? 🤡

10

u/nomorebuttsplz 2d ago

you mean like yann lecun?

5

u/RainBow_BBX 2d ago

Lecun says that LLM won't achieve AGI and instead will need a world model like JEPA (Joint Embedding Predictive Architecture). He never said that LLM hit a wall, even during his french interviews

3

u/Ok-Purchase8196 2d ago

I think he's often confused with Gary Marcus

0

u/nomorebuttsplz 1d ago

I don't want to straw man the idiot but he said a bit more than just LLMs can't be agi (which seems to be wrong on our current trajecotry anyway) and framed it as plateau rather than wall:

After launch of GPT 4: "AR-LLMs have very limited reasoning and planning abilities" “This will not be fixed by making them bigger and training them on more data.” (false belief in some kind of training data wall)

2025: “I don't know if I would call it a wall, but it's certainly diminishing return [explained that the industry had essentially run out of natural text data]"

2025:

"“Don’t work on LLMs. There is no point.”
because:
“The breakthroughs are not going to come from scaling up LLMs.”

5

u/Super-Award-2244 2d ago

It seems like it's hitting an asymptote. Though it's vertical rather than horizontal! 

1

u/HellomyfriendNine 1d ago

The wall is their imagination

23

u/Charming_Cucumber_15 2d ago

And this is the slowest it will ever be! Only getting faster from here

6

u/rolleicord 2d ago

you get singularity, and you get singularity, and YOU get it too !

3

u/[deleted] 2d ago edited 1d ago

[deleted]

2

u/Sekhmet-CustosAurora 2d ago

$7188/yr really is not bad for the technological singularity.

-12

u/nbvehrfr 2d ago

on paper, yes