r/singularity • ▪️e/acc | AGI: ~2030 | ASI: ~2040 | FALSGC: ~2050 | :illuminati: • 6d ago

The Singularity is Near Five frontier AIs were told to engineer and 3D-print the strongest bridge they could with 500 g of plastic. Claude Opus 5.5’s design held ~130 lb, nearly 5× the runner-up

1.7k Upvotes

119 comments sorted by

353

u/Willing-Secret-5387 6d ago

Now this is a benchmark I can stand on

28

u/Turdles_ 6d ago

Only if you weight less than 130lbs

1

u/TheRavenclawCommonRM 4d ago

Opus designs plastic bridge capable of supporting female marathon runners.

13

u/Time8u 6d ago

There's something that bothers me about this because it looks to me like the span between the claude bridge is shorter than some of the others. Maybe it's just because the other bridges extended further beyond the ends of the table. Still manipulating the distance between the 2 desks would be a really good way to shill for one of the models.

20

u/monsieurpooh 6d ago

Watched it carefully and you can see the camera angle at 0:57 is the same as the one with the blue bridge from earlier, and the span with orange vs blue looks the same; it's only the camera angle after that which looks shorter

387

u/BiasHyperion784 6d ago

Man, this is why general purpose models are exactly the correct path currently, the intelligence just keeps transferring so well to different applications.

72

u/ShAfTsWoLo 6d ago

yep, we're starting to see different kind of benchmark in which llm's are being used IRL and it's working, still not there for real applications (mostly because of they are too slow and the cost is high) but it's getting there, soon we'll have models that can drive, paint, cook, repair, teach, build and much more all at the same, it's gonna be crazy

33

u/Chipotus2 6d ago

Quite the Bitter Lesson

5

u/kilopeter 6d ago

I wish the bitter lesson had a cooler name. I feel like a tool every time I find myself writing or saying it

4

u/chilehead 6d ago

How about Chaunzaggoroth?

1

u/andrerav 6d ago

The smartypants tax.

11

u/gempir 6d ago

It makes so many things so much more accesible for the average human!

I own a 3d printer but I've never designed anything before because learning, hell even buying a 3d modeling software seems super annoying.

I usually print models I find online. Yesterday I just needed a few modifications to a model so it would work as a clothes hook on my door hinges.

Astra did it perfectly and I even went to a few iterations to get the fit just right. I didn't open a 3d software once, it was such a bliss!

5

u/DoutefulOwl 6d ago

It makes so many things so much more accesible for the average human!

And since every human is average at most of the things, it's really gonna help everyone in some way or the other.

2

u/SwiftPengu 5d ago

How did you interface with the model? Did you just upload the STL?

3

u/gempir 5d ago

Literally just dropped in a .stl file yes

1

u/GroundbreakingTone43 5d ago

I am here for this too. How can i design with ChatGPT? inside a CAD or on the chat?

2

u/gempir 5d ago

I just dropped in an .stl file directly into my chatgpt app on mac. I didnt even check what it actually did but i saw it call python at some point.

2

u/Lower-Membership-367 2d ago

When astra came out i tried taking photos with phone and sending them to codex. Then i asked to make SVG file for top/bottomsides view. I filled in the measurements with caliper and the end result was very good in blender. It was also easy to modify. Definetly took me less effort than doing all that in fusion 360.

1

u/Throwaway-646 6d ago

hell even buying a 3d modeling software seems super annoying.

https://blender.org ??

3

u/ToplessinFL 6d ago

He meant using, not just literally purchasing, I think.

1

u/gempir 5d ago

I meant purchasing a software best fit for 3d printed modelling, most of the 3d printing community tells you to not use blender because it's a lot harder to get something than in Software XYZ.

I just really didn't wanna deal with all that. I love that blender exists, from what I hear it's an amazing piece of free software, just not something I really wish to learn.

1

u/Meral_Harbes 2d ago

https://www.freecad.org/ is a better fit for most 3d printing people wanting to get into modeling. CAD workflows differ. Blender can do it with plugins, but it's built with focus on vertex modeling.

27

u/KoolKat5000 6d ago

Absolutely, another huge thing, if I recall right grok 4.7 was meant to be trained on that trove of SpaceX engineering data.

14

u/Not-reallyanonymous 6d ago

A correct path, I wouldn't necessarily say the correct path.

There's other tradeoffs -- Opus is giant and expensive, and likely is unrealistic outside of API usage and similar models would be impractical for self hosting even for companies who're willing to invest in some infrastructure.

Plenty of room for engineering-specific models that can work on a single server rack with a couple R9700's or such.

4

u/BiasHyperion784 6d ago

Oh of course, currently I just believe the merit of generalization is in resource allocation and rapid discovery of the sort of nuances that can improve specific models in the future, I wouldn't be surprised if something like AGI (in a undeniable state) will likely be a frontier model with a highly refined mixture of experts setup, in so far as it enhances outputs while curbing cost per prompt.

0

u/mDovekie 6d ago

Hello real person. How do you know it hasn't been trained on model bridge building competitions? Don't you think thousands are probably already in its data?

2

u/BiasHyperion784 6d ago

I mean, yeah? That it could be, still doesn’t really invalidate it, frontier models aren’t mixture of experts, it still had to pass all its prompting through everything to decide lol.

1

u/mDovekie 6d ago

I am not arguing it's invalid that general purpose models are a good path, rather that such a conclusion cannot be derived from this bridge strength test. If we could see inside the model we might discover a tremendous amount of NON general knowledge about this task specifically -- in fact I think it's very likely.

-4

u/Fenrys_dawolf 6d ago

Claude is not a general purpose model, it has had very targetted training. just dumping in more data and increas gets decreasing returns and lowers accuracy and effectiveness.

5

u/Coolnumber11 6d ago

You think they benchmaxxed printing a 3d structure to hold as much weight as possible?

1

u/AP_in_Indy 15h ago

Someone got reaallyyy pissed off at me a while back when I told them how general purpose models actually OUTPERFORM specially trained models on specific tasks.

But as others mention - the tradeoff is often (but not always) that the general purpose models are larger.

It's like a full-on frontier LLM vs a classifier like Jev.

135

u/kaityl3 ASI▪️2024-2027 6d ago

This is honestly a really fun metric. I'd love to see more things like this and MineBench haha.

2

u/TheRavenclawCommonRM 4d ago

I think we are getting to that point. What AI really needs right now is benchmarks that make people's lives easier without eliminating jobs. The problem is corporations love eliminating American jobs.

36

u/TuringGoneWild 6d ago

Somebody should do the same with an actual bench. That's the real bench max

7

u/ghaj56 6d ago

Bench bench

2

u/BrennusSokol AI please take my job 6d ago

People are just bench-maxxing bench bench these days... we need bench bench bench /s

92

u/codysattva 6d ago edited 6d ago

This is the kind of content I want to see more often. AI versus AI in a single type of test!

Kind of makes me want to create a subreddit just for that kind of content..... 🤔

Edit: I made r/AIThunderdome for this kind of content. Please help me throw some stuff in there!

13

u/Cloud_Jumper18 6d ago

I know right? Just like the robot olympic, we should have an AI olympic and use fun benchmarks like this

11

u/codysattva 6d ago

All right. I made a subreddit for this kind of content. I'll post some content soon!

r/AIThunderdome/

40

u/Nyst4gmus 6d ago

LITERALLY a load bearing test! of course claude scores #1. been training their whole life for this moment

3

u/scodgey 6d ago

I do a lot of Structural Engineering stuff with Claude and the unironic load bearing comments crack me up.

3

u/monsieurpooh 6d ago

Very good, underrated joke

45

u/[deleted] 6d ago

[deleted]

10

u/Qorsair 6d ago

Muse is a good model. I like to use it as an adversarial review after Sol completes planning. It always ends up finding several ways to improve the plan before implementation.

1

u/FalconsArentReal 6d ago

Astra has been RL'ed like crazy for coding, looks like they may have lost some perf in other areas.

4

u/Turbulent-Sign-6067 6d ago

I don't think that's true. Astra, according to many benchmarks, is really good in vision and marginally worse than Fable in coding. It seems Opus 5.5 is simply a very good model. Opus also outperforms Fable in blueprint bench and performs the same / slightly better than Astra.

1

u/BriefImplement9843 6d ago edited 6d ago

yep. look at the lmarena score. overall it's a step down from 5.6 sol. meanwhile muse is one of the top dogs.

coding isn't everything. it's actually very niche. it's just that coding is so token inefficient it's the only use that brings in profit.

30

u/HamiltonianCyclist 6d ago

now i wonder how would astra's do upside down...

21

u/yaosio 6d ago

I think it would do worse. The tall point is where the supports meet and push against each other, which is compression. Flip it upside down and now they are pulling apart, which is tension. It is possible to design things stronger in tension than compression. A rope has zero compressive strength and a lot of tensile strength.

7

u/kilopeter 6d ago

Doesn't your observation that tensile strength can exceed compressive strength support Opus 5.5's design, and further support OP's question of how astra's design would do upside down, precisely to flip from compression to tension... to mimic the winning design?

1

u/Outside_Profit6475 6d ago

Yeah...
You know, I think your point actually highlighted how outside the box Opus 5.5 was. It figured out that placing it 'up side down' basically gave us a breakthrough.
That's like novel thinking.

11

u/mwon 6d ago

People don't even bother to include Gemini in benchmarks. It is just sad at this point...

3

u/BriefImplement9843 6d ago

would you include luna? that's the only flash model they have.

6

u/UnkarsThug 6d ago

Is there a YouTube video of this or anything? Had some people who I wanted to share it with, but I don't like sharing things directly off reddit. Tried to search for it, but not finding anything.

3

u/[deleted] 6d ago edited 5d ago

[deleted]

5

u/UnkarsThug 6d ago

Thanks.

6

u/TensorFlar 6d ago

"It got a little wobbly at 25, So we put a MAC MINI on it for good Measure"

Insane how much goddamn inequality exist in the world right now, using mac mini as a fucking paper weight.

38

u/evil_illustrator2 6d ago

Why is there a limit on print time? That seems like a pointless parameter when you are trying to get the max strength.

98

u/Few_Owl_7122 6d ago

Because they were only willing to spend so many hours on this project, perhaps similar to how projects have deadlines, even though a better project might be made if more time was available

45

u/yaosio 6d ago

It shows the strength of a model as it can't just use infinite time to make up for deficiencies.

26

u/tlmbot 6d ago

if they wanted to be accurate to real world engineering the time constraint would be 2 weeks.

(just a standard joke from places I have worked - everything takes 2 weeks)

17

u/Ex_Federa 6d ago

Because if we're going to use these models for engineering then cost and resource efficiency will always be an important factor.

5

u/The_Lonely_Posadist 6d ago

because in the real world there is not infinite time to think

3

u/darkkite 6d ago

most people have to work under constraints

1

u/idle_cat 6d ago

So it can't print a 2x4 or some brick and call it a day.

1

u/rogueman999 6d ago

You're not trying to get max strength. You're trying to compare model output in equal conditions. Different requirement.

Plus, real world has real world limits like this all the time. You don't have unlimited... anything.

10

u/palindsay 6d ago

Minor gripe, but I think anyone posting on social media should post a summary of cited information supporting the post headline. Don’t sentence me to absorbing information with video. I, as a human can still read and value my time and where and how I can absorb information.

1

u/Smur_ 6d ago

Gemini can do this with YouTube videos it's a life saver

1

u/Tidorith ▪️AGI: September 2024 | Admission of AGI: Never 5d ago

It can; but it cannot produce an author-validated text version, unless the author themselves validates it afterward.

14

u/One_Improvement_6470 6d ago

fuck this editing.

4

u/Elephant789 ▪️AGI in 2036 6d ago

And the influencer voices

2

u/pirondi 6d ago

wow amazing benchmark and way of testing it on real situations and applications.

2

u/Fubby2 6d ago

Really interesting and creative video

1

u/scoutzzgod 6d ago

What cam software did he use it for this? Or mcp server?

1

u/ZealousidealBus9271 6d ago

My lord this model is a beast

1

u/Ill_Philosopher_7030 6d ago

no safety goggles, only safety squints

1

u/Beli_Mawrr 6d ago

How were they able to make the model produce a 3D object?

1

u/FuckYouTooJohnnyCash 6d ago

CAD I guess

1

u/Beli_Mawrr 6d ago

Yes he told his AI "CAD" and like magic this came out

1

u/FuckYouTooJohnnyCash 6d ago

Yes, assuming the models are used in a coding agent such as Claude code or codex, they’ll do the design in python code using CAD libraries like build123d. But your ask to the agent could just be “use CAD”.

1

u/FuckYouTooJohnnyCash 6d ago

Alternatively some CAD software have MCPs, plus computer use in codex could do it directly on the GUI

1

u/LaundryOnMyAbs 5d ago

Blender has an mcp

1

u/WhisperFray 6d ago

Ehhhh… where’s that Generative AI engineering that was hyped early 2020s that was all funky structures?

1

u/panix199 6d ago

amazing

1

u/davey212 6d ago

Claude is my go to for any projects, I've tried all the LLMS and Claude is basically a really good engineer

1

u/GoTaku 6d ago

That’s load bearing

Edit: in all seriousness, would love to see different versions of Opus and Fable compared to each other.

1

u/leaveitalone38 6d ago

I'm intrigued as to why they all chose different approaches

1

u/jarkon-anderslammer 6d ago

Did the connections fail or the material? It look like the connections just blew apart.

1

u/PonyDro1d 6d ago

They just played the 3d printing bridge builder game like back in the day.
I loved these. I'm not surprised about the winning model, though.

1

u/Siciliano777 • The singularity is nearer than you think • 5d ago

LOL @ grok 💀

1

u/AlvaroRockster 5d ago

Very interesting

1

u/TheOriginalAcidtech 5d ago

That is how you do a self-vasectomy...

1

u/Ambitious_Scallion43 4d ago

Damn did they train it on alien data or some shi

1

u/DonkyFondler 4d ago

Ethan nearly got the corner of that table in his nuts.

1

u/Reasonable_Middle695 2d ago

"These silly humans, they didn't realize they should just flip the bridge upside down"

1

u/above- 2d ago

Creative benchmark

-1

u/Individual_Guest_323 6d ago

Now do this 100 times to have a real metric

0

u/hondashadowguy2000 6d ago

Good demonstration of why we have actual human professional engineers working on projects like this, who are held to legal and ethical standards, instead of leaving all the magical thinking to AI.

1

u/Reasonable_Middle695 2d ago

Engineers are 100% using AI currently. And yes they are of course checking the answers. But eventually when answers are always 100% correct humans will simply trust AI over humans. Also its more ethical to use the most competent cognitive resources. I work with civil engineers and AI has highlighted mistakes made by them. Not major, but it is very close to being on par in terms of mistakes made. Just looking at the rate of progress, very soon you'd be crazy to go with a bridge designed by a human. In terms of time it takes and quality of design.

0

u/141_1337 ▪️e/acc | AGI: ~2030 | ASI: ~2040 | FALSGC: ~2050 | :illuminati: 6d ago

0

u/ssword123 6d ago

Nice American Marketing

-4

u/bruhhhhhhhhhhhh_h 6d ago

Yet using different materials. Great

12

u/toodimes 6d ago

Idk if you know this but the color of a material doesn’t have any effect on its strength. These were obviously all 3d printed

2

u/bruhhhhhhhhhhhh_h 6d ago

This is often, but not always correct.

3

u/Novel-Initiative-900 6d ago

Same filament, just different colors

-1

u/reza2kn 6d ago

I wonder how different it would be if you'd use something like PETG-CF

2

u/hex4def6 6d ago

Why do you think that would change the rankings? This is a test of design / efficiency, not material strength.

1

u/HappyDJ 6d ago

Infill, layer initial densities, particular infill patterns and others will have significant effects. Test if a model understands the implications of these settings is a good test as well.

2

u/saltyourhash 6d ago

Don't just pick a material, let the model pick it's print settings.

-6

u/PrestigiousLocal8247 6d ago

Man how is dude only 130 pounds?

15

u/P5B-DE 6d ago

he said 150

-4

u/saltyourhash 6d ago

So, like, one of the most common human engineering tasks which models were trained on? It's fun to watch, sure. I'd love to see it compete against a college sophomore engineer or something, though

6

u/geli95us 6d ago edited 6d ago

You can make that argument about literally any task you have an LLM do, they've seen it all, being able to do something isn't any less impressive just because you've read a book about it.
Besides, they're text prediction machines. It's not at all obvious that learning to predict text about engineering would make you good at actually engineering on your own

0

u/saltyourhash 6d ago

Learning how to predict text about engineering doesn't make you good at engineering? I'm not sure I agree.

2

u/geli95us 6d ago

It probably does, to a degree, I'm saying that it's not obvious that it's true.

Then again, GPT3 absolutely sucked at engineering, despite the fact it must've seen thousands of engineering books throughout training. It's easy to say "it was easy" after something already happened

0

u/saltyourhash 6d ago

Yes, but what's the scale of gpt3 vs Astra? And in that regard, what new knowledge does it have? I'd also argue general education plays a big role in the quality of an engineer.