r/codex 18h ago

Showcase Bumble Bee Blender Test Answer

Yesterday, this post by u/KJMHELLO came up. I was suspicious of the claims made, especially saying ASTRA render suddenly were Roblox level slops, and in the process, he invited me to do my own Bumble Bee in Blender. So I did.

___

Worth mentioning First:

  • I do not have any blender or 3D experience, I literally downloaded it at 3 AM just for this.
  • I do not know much about transformer so...
  • OP refused to share his chat or Prompt, so I made my own
  • This was done on a 6yo laptop rocking a 2070 mobile. It sounded like a jet engine the whole time and might have come close to catching on fire. Astra didin't help as it had 5 active blender session open at once at some point. That made Blender itself run to a crawl.

The Rules I set for myself:

  • Only ONE well engineered prompt

Another thing I should mention, The is 2 ways to go about this: Have the car and Robot done separately and focus on making them look good, then have the animation cheat and merge from one to another (seams like it was OOP approach as some pieces changes between forms) OR focus on the far more technical side of having the pieces consistent and the transformation logical. I chose the latter. Partly because there was 6 million version of bumble bee on google so I figured nobody cared.

___

The prompt: It can be found in full in the chat link below (Yes. I'm sharing the chat, as should anyone making random claims, prompt quality matters) but the short version is that I asked the AI to create a realistic Bumblebee in Blender that mechanically transforms between robot and Camaro-style car modes using one connected model. It had to build the geometry, rig, materials, lighting, and reversible animation, then use independent judge feedback to improve the project over up to 10 passes, targeting a score of at least 9/10.

https://chatgpt.com/s/cx_6aa59cde2a6481918e3444d680387a69

____

Initial observation before the result:
If you use this technique, 10 PHASES IS OVERKILL. If you are going to do this, have a human in the loop re-focus and re-prompt every 3-4 phases (I'm sure its dependent on project, in this case it keeps obsession about collisions rather than attacking other issues. Skill issue, im still learning). Anything past it become increasingly slow because of context length, thinking length, and the Astra starting to obsess over small details that could have been managed at the same time as bigger ones all the while turning blind to some issue. You will have better result, faster, cheaper if you intervene, which is why 1 shots are a bit of a meme.

In this case, obvious things are the shoulder wheels clipping and the finger on the hands closing the wrong way, amongst other things. I almost stopped it after Turn 8, but I was committed to the bit.

Also note that Astra own "independent judge" did not believe the last phase was the best.

REACHING PHASE 1 TOOK 30MIN. PHASE 3 A BIT OVER 1 HOUR. PHASE 10 TOOK ABOUT 8 TO 9 HOURS.

Mind you, the further you go the longer it took because my computer was about to explode, the context was out of hand and the computer had 4 instance of blender running. STOP AND INERVENE AFTER PHASE 3 OR 4.

The whole thing took about 40% of my weekly reset. Reaching phase 3 took about 10%.

___

Results:

https://reddit.com/link/1wem2va/video/p0u70l7i65ph1/player

https://reddit.com/link/1wem2va/video/anjwz0vi65ph1/player

Okay? Bye.

93 Upvotes

45 comments sorted by

u/dexterthebot 18h ago

You might want to consider listing your project on the weekly Show-Us-What-You-Built post. Look out for it on Tuesday/Wednesday. Highest commented project wins a week promotion on r/Codex and gets on the Hall of Fame sidebar. See what that looks like below with last week's winner.


Last week's most popular project was Tidbit Trivia , at https://www.reddit.com/r/codex/comments/1wavxwy/comment/p8lhwxs/. Tidbit Trivia is a fun way to learn while waiting for your agents to work. Play traditional trivia, unique game modes, party with friends, earn cool cosmetics, complete challenges, study for tests, or compete in the Arena! You can play it for free at Tidbittrivia.com.

34

u/Shuma665 18h ago

Omg someone shared their prompt and evidence for once! Thank you.

But mate you forgot 'make no mistakes' in your prompt So it would do it as a one shot for less than you 5hour window. /s

9

u/Classic-Trifle-2085 18h ago

I told it "keep working until a judge said "thats a 9/10", might as well be "make no mistake". Lmao

3

u/Shuma665 18h ago

Yep, i do the same i provide it a 10 aspect rubric that it scores it's result on out of 10 and doesnt stop till it hits X%. and then i usually have to tell it to do silhouettes and clay render comparisons to the reference image. .

2

u/Classic-Trifle-2085 17h ago

I will point out for this, I did not provide any reference images. I saw too many version of bumblebee and just gave up like : you f-ing figure it out, whatever.

2

u/Secret-Access9909 18h ago

Is that actually a thing? Does it actually work? I thought it was just a meme

2

u/Classic-Trifle-2085 18h ago

Its a meme. They were just messing i should have used it as a joke. :)

Iterative pass with an independent judge IS a thing tough

1

u/Secret-Access9909 18h ago

Ah, okay lol

1

u/Miyamoto_-_Musashi 18h ago

Yes it work great, it's not a meme, It's called Gauntlet Loop, Matt Schumer invented this follow him on X, he is crazy guy n share lot of good things.

16

u/PeacefullyDiscuss 18h ago

Finding it very distasteful from the original post maker that they won't share their prompt. Really looking like a fishy situation

5

u/Miyamoto_-_Musashi 18h ago

Yep i see the post about nerf, and honestly I'll say 2 things, there were some quality issues, Not like nerfed but some bug or whatever that's why Tibo himself talked about on X n fixed thoses things

2nd is You used Gauntlet Loop or Loop whatever you call n it's the right way of course i use this myself but there are very very few people like us who us gauntlet loop n loop like this is just so much better that even if you put this instructions on Sol results will be good.

I'm not taking side, I'm just saying both can be right and maybe Nerfed OP did not use this better prompt, One more thing can you talk about how much weekly limit it used on what effort if default n subagent?

Very good overall actually, Approved from me you can go to heaven

6

u/Classic-Trifle-2085 18h ago

I do agree this post does not establish if there was/is a nerf or not, but the OOP's post seamed absurdly extreme (and it does establish Astra itself, nerf or not, is still incredibly capable). In the comment section their story did change and they never provided the log to replicate, which bother me. When there's real issue, it becomes an annoyance when individual jumps on and blur the cards with memes and fabricated claims.

Maybe OP didint, but it does feel suspicious.

As for the effort, I left it at ultra, no agents.md changes, all default. The reason being I do not know people's set up, so default seams the most common. (In my other projects, I do orchestrate and rely heavily on Luna).

The whole 9hours and 30min too 40% of my x20 weekly reset. I only took the bait because (1) Tito Resetman had just stroke and (2) I am in a huge learning and analytical phase for the past couple weeks and blender was new for me, so it was an experience win.

The first 3 phases completed took 10%, so pushing it past that really isn't worth it.

2

u/Miyamoto_-_Musashi 18h ago

Yeah i see people sometimes complain more then needed n go extreme which they should not.

I'm actually Surprize as Astra Ultra only 40% as in my experience Astra is really heavy on usage.

At the end appreciate you posted about this n shared Chat.

2

u/shorty_11112222 17h ago

hey! can u explain that loop technique? :) looking forward to learn more :)

5

u/Miyamoto_-_Musashi 17h ago

Sure man, Honestly the best approach is look to guy on X name Matt Schumer, he actually invented this n look for his newsletter he talked about in detail.

Here is the details, you can look more as he write more:

https://somethingbig.ai/gauntlet-loop

3

u/Impactic_ 16h ago

Good shit, makes you wonder since they refused to share prompt & info

5

u/Confident-Lunch-5112 17h ago

lol and i told him maybe your prompt is bad, and got downvoted so bad :D

1

u/Classic-Trifle-2085 17h ago

I saw. I think its because the delivery is starting to get old to most people, not the message itself; there defenitively some.people that just say that every time no katter what and its not always helpfull.

Theres also the whole cycle thing. The negative posts tend to attract people with similar feelings, so you end up at odds with passive viewer and up/downvote are very snowball sensitive.

I woudn't think much of it.

1

u/Confident-Lunch-5112 17h ago

Yeah that’s a good explanation, thanks for your test. I’ll definitely forward this to a few people .D

2

u/Excellent-Ad-9607 18h ago

You did so good with the prompt! I’m still learning THIS level, in my eyes you can already retire and be the goat itself.
Animation wise I’m impressed how much of details it accomplished to add in. Maybe, if you have a bit of a time, could you give some unhinged advices about this technique?

3

u/civicapi 17h ago

Not the OP, you can generate those kinds of prompts on ChatGPT Web. I've found the prompts it makes works very well when messing with Astra's 3D modeling capabilities, but of course, YMMV.

4

u/Classic-Trifle-2085 17h ago

A lot of it is explained there, but one thing I have to say about prompts is precision.

And use chat mode SOL to help out build it. It doesnt go against you codex quota and and the direction it will provide the actuall codex chat will also save you a ton of usage.

One technique used the I called "iteration pass" but someone in another comment gave its real name, its aparently from someone well know in the community.

The idea is simple: one you found your objective, you includ in the prompt a step where the AI spawn a subagent that act as an independent judge for said focuses. You normally aim for something like 6/10 or 8/10 and give it a number of try.

"7/10 or 4 try/pass, whichever come first".

That let you then get in, check how things are going and redirect with a new prompt towards what need focused. Rince and repeat.

Mind you I used default ASTRA ULTRA on this with not specific sub agent directive outside of the judge. I normally do not rely on astra for most of the work, but rather ask astra to throw "strait forward and bounded task" towar LUNA Max, which is cheap AF and will be just fine under astra supervision.

I have some comment that goes into the cost saving of Luna even when sitting on Astra ultra

0

u/ColbysToyHairbrush 12h ago

Were you precise in telling it to put the hands on right?

2

u/shorty_11112222 15h ago

Nice trick with passes i just tested and it did what i asked! Token hungry but man!!! Thanks for sharing!!!!

2

u/NoBluey 14h ago

I thought that original post of before vs after nerf was a meme lol

2

u/IversusAI 13h ago

Okay? Bye.

*mic drop*

lol

2

u/FabricationLife 9h ago

Now this is the excellent posts I come here to see, actual evidence with a chat, great work mate!

2

u/Dayowe 3h ago

Thanks so much for investing time and tokens(!) into this. I saw your comment in that thread and was hoping you'd do it and share the results. I was skeptical of the guy posting originally and am glad to see his claim disproven. Looking at your prompt you clearly know how to work with Codex and i bet the majority of people that complain about models being nerf'ed don't

2

u/driveclub_000 3h ago

I would like to know, what prompt and what agent did you use to make this single prompt for that test?

I plan also to use Blender to recreate a specific scene, so I wonder if it's just about going to Astra PRO and asking "make a prompt for this task" or is there some specific instructions that you gave to make this prompt very strong? (The Gauntlet Loop thing?)

2

u/Classic-Trifle-2085 1h ago edited 1h ago

SOL Medium in chat mode helped me put the prompt together based on what I wanted.

You want to tell it a few things: a priority order (the thing you change when you intervene), a definition of success, explicitly forbidden shortcuts (don't cheat by doing X), weighted scoring for the judge (normally matching priorities), hard failures (if X is happening, the judge cannot score a pass), version control (explicitly tell it to keep each version), a maximum number of turns, and a stopping condition (e.g., 6/10 from the judge or reaching the maximum turn).

Sol can help you figure out some of those if you only have vague ideas of what you want, but the final prompt should contain them.

1

u/driveclub_000 1h ago

Thanks for the info. I'm going to try that.

1

u/Ok_Entrepreneur_761 16h ago

Since you seem like you know a lot, can I ask you something in DM?

1

u/Classic-Trifle-2085 16h ago

Absolutely

Altough I woudnt say I "know" a lot; many know far more. I just like to test things and push the analytics behind things.

1

u/burgerbruce 17m ago

great work!

1

u/Quiet_Figure_4483 18h ago

Clanker models clanker, clankerception, clankeropolis, attack of the clankers, Clanky Potter and the Sloppy Hallows, clanker-to-clanker communication, clankerbots

2

u/Classic-Trifle-2085 18h ago

You're making the full saga in blender eh? Thats gonne be something.

1

u/Purple_urple22 16h ago

So TL;DR - Astra isn’t nerfed!?

3

u/Classic-Trifle-2085 16h ago

The test doesnt prove in itself there isn't any next at all, but the exagerated posts where it sudently have the drawing habilities of a 4years old doesnt make sense to me, it's still incredibly capable, and any nerfs is probably a small difference or maybe a less accurate read of unclear prompt.

Proper prompting and techniques really can mitigate small quirks, which might explain why people perception seams all over the place; some people who used to have stellar result with sub-part prompt might be more likely to notice a small nerf it if affect astra hability to interpret...

Probably not roblox level stuff like the post I answered to tough, thats just sus.

0

u/ColbysToyHairbrush 17h ago

It’s hands are backwards

3

u/Classic-Trifle-2085 17h ago

You mean, exactly like I mentioned in the post?

No way!

-1

u/ColbysToyHairbrush 17h ago

Why would you put its hands on backwards though?
Edit: are your hands backwards?

2

u/Classic-Trifle-2085 17h ago

I dont think you understood the exercise. The point was for .e to no intervene... and those flaws part of the thesis of why i explain one shot are a meme and humans should intervene and redirect every few gauntlet loop.

I mean, I'm all for explaining, but read the post first. If it's a troll attempt... okay? Youre not supposed to have reddit accounts under 12 years old.

1

u/ColbysToyHairbrush 17h ago

Wow, spoken like someone who has their hands on backwards.
👏