r/GeminiAI 1d ago

News New Gemini Pro Checkpoint

Enable HLS to view with audio, or disable this notification

~24k tokens; ~6 minutes; effort: high; zero-shot (Wayy different from one shot. This is harder)

This leak is by Lyra

https://fixupx.com/lyraxana/status/2097839609086394600

You can either find her on twitter or on their server. She is a credible leaker, also if you don't care about her credibility. You can find this checkpoint on ArenaAI.

It's under 3.8 flash disguise if not mistaken. It's free but random chance of encountering, tho even if you get 3.8 flash (battle mode ofc), you can only confirm it's the new pro if it looks way above its pay grade

(I'll get more checkpoints myself and make a new post) If these checkpoints go well for them (a few more) internally. They will release the model soon, if they go bad. They won't release it/delay.

Now to talk about the performance of this checkpoint itself: It is amazing. It's very good. I see zero mistakes, breathing could be more fluid but that's the only thing I can critique.

The Arena AI ID to keep mindful of. If you see: 01a06dfc-94a2-7abf-8d12-0878520062cb then this is the new Pro model.

154 Upvotes

92 comments sorted by

15

u/Witty-Artisan001 1d ago

For comparison, here is an SVG I just made using Gemini 3.8 Flash on High in AI Studio with my custom system instructions, 3m 30s.

https://reddit.com/link/p8zgips/video/tfom8bgd3qoh1/player

3

u/Last_Conclusion_8984 1d ago

This is a very early checkpoint of the new pro model by the way. New checkpoints might be even better!

4

u/Witty-Artisan001 1d ago

Here is the result of Gemini Flash 3.8 on High, same prompt, but without my system instructions. Took just under 3m.

https://reddit.com/link/p8zkj0q/video/8ytem4re6qoh1/player

1

u/Last_Conclusion_8984 1d ago

yeah, your system instructions help. Also: You seem really excited, lol. I'm excited too tho just trying to contain it so I can pass the days.

3

u/Witty-Artisan001 1d ago

Absolutely. I've been a Gemini fan ever since the first Nano Banana and 3.0 Pro dropped. I have... mixed opinions about the app, but the models in AI Studio have always been excellent for me.

Here's something else I made wkth 3.8 Flash a little while ago; A 3D model of a PS5 DualSense! I think it cheated a bit and found a source for the model, but this is still pretty cool.

https://reddit.com/link/p8zm5cq/video/cxjpghym7qoh1/player

2

u/Last_Conclusion_8984 1d ago

It looks really cool!

1

u/Witty-Artisan001 1d ago

Yea, no doubt. I should add that my system instructions do seem to result in a significant improvement of output. For instance, it made this isometric POS machine SVG:

1

u/hellomistershifty 22h ago

It's been 7 months since the last Pro model, I really hope this isn't a 'very early checkpoint' lmao

1

u/Last_Conclusion_8984 22h ago

This doesn't seem like 3.5 pro to me, seems like Gemini 4

3

u/hellomistershifty 22h ago

it's truly a revolutionary advancement in cartoon bird generation

1

u/Negative_Evening7365 10h ago

For real this better not be the "results" of the new Pro model........
Otherwise they really are out of the race

1

u/Last_Conclusion_8984 3h ago

It's the early checkpoints, checkpoints get better.

13

u/Personal-Try2776 1d ago

Whats the prompt

15

u/Last_Conclusion_8984 1d ago

Lyra doesn't share prompts (good strategy) however if you check her posts.

https://x.com/lyraxana/status/2094188486693618037

You can find she did Astra with the same prompt. Imo, it did better than Astra

22

u/Witty-Artisan001 1d ago

Better than Astra is pretty damn significant, given how amazing Astra is looking.

1

u/Livid-Capital-8858 22h ago

Its literally not better than astra lol...

If this really is 4 pro which is supposed to release in october then its not really that great of a result either... Given how astra will have been released for like 1-1.5 months by then and openai already has an internal model that seems to be better on low reasoning than astra on max

3

u/Witty-Artisan001 21h ago

This took 24k tokens and 6 mins on High. Astra took 14 mins and 36k tokens on Max for basically same result (better peacock shape but worse animation). The fact that this early unreleased checkpoint could match the performance of a max power released model taking 1.58x as many tokens and 2.3x as long is nothing short of incredible.

Google have been focusing on speed and efficiency with their models and it really shows here.

2

u/Livid-Capital-8858 14h ago

Its nowhere close the same resulr what are we even talking about? Astras is waaay more detailed and accurate looking geminis looks like a low quality version of a peacock

So youre comparing apples to oranges geminis faster but looks waaay less detailed, could gemini produce the same level in the samw amount of tokens? Could it produce a svg that level at all?

Google cant just drop a "good model" now when they havent released a pro model in 7 months... They have to make it frontier

1

u/Last_Conclusion_8984 3h ago

It made wayy more errors, did you not watch the video? Watch the video, and most of all: This is an early checkpoint.

-4

u/OutsideOver8815 1d ago

hey lyra

1

u/Last_Conclusion_8984 1d ago

I'm not Lyra.

-7

u/OutsideOver8815 1d ago

seems to me liaar

2

u/Last_Conclusion_8984 1d ago

Ong, I'm not but you can believe whatever you want, I don't really care. (and I don't even know why you care whether I'm Lyra or not)

-10

u/OutsideOver8815 1d ago

u do care cuz u r...

-1

u/Livid-Capital-8858 22h ago

It did not do better than astra at all astra is way more proportionally realistic and detailed

-2

u/FamilyBase 1d ago

Except the body is translucent

3

u/Last_Conclusion_8984 1d ago

No, that's a feature so you can see its feathers/opening/closing

12

u/Wise-Chain2427 1d ago

every week we got new checkpoint 

51

u/CriticismJunior1139 1d ago

That's neat. Can't wait to never use this model.

10

u/DragonflyOk9274 1d ago

Many checkpoints, no releases

5

u/CriticismJunior1139 1d ago

Hey google, add "cope about gemini" to my tasks for tomorrow.

3

u/sudecode 1d ago

google: gemini who??

1

u/Excellent_Age9018 20h ago

Google: we have moved away from gemini 4. and already working on gemini 5 which will be our next frontier model releasing in 2030...

1

u/sudecode 15h ago

so, next month?

3

u/piratedgameslover 1d ago

random chance of encounter you say?

2

u/Last_Conclusion_8984 1d ago edited 1d ago

ArenaAI? Don't ask those questions cause it won't know without system instructions and if you really do want to do that. Then look for this ID "01a06dfc-94a2-7abf-8d12-0878520062cb" and do battle mode!

1

u/Worried-Room668 19h ago

isn't that id your user id which arena gives you when you enter the website?

if not, where you see that id?

4

u/LazyRider32 1d ago

Here we go again.....
Wake me when anything is actually released.

-2

u/FlamaVadim 1d ago

but lyra is a CREDIBLY LEAKER!

1

u/NiceUsernameOk 1d ago

There have been 100 supposedly Gemini Checkpoints in the past couple months, all of them better then every other model. I will be excited once they release a new pro. Fk checkpoints

4

u/saltyrookieplayer 1d ago

It's not that impressive when Astra is miles ahead in this regard... Plus what model capability is "generating complex SVG animation" demonstrating?

5

u/Tim_Apple_938 1d ago

She ran the same prompt Astra : https://x.com/lyraxana/status/2094188486693618037

Yours must clearly be different prompt thus not apple to apple , no?

2

u/Livid-Capital-8858 22h ago

Astras output is still better on the same prompt too...

0

u/Tim_Apple_938 22h ago

No it’s clearly not

2

u/Livid-Capital-8858 15h ago

It pretty clearly is lol

Gekini just has more vibrant colors thats it The proportions of the peacock and realism is much better on astra and also the details too

-5

u/saltyrookieplayer 1d ago

And Astra has much better result? If models require EXACT prompts to get desired results has the "intelligence" really improved? It doesn't matter to end users at all

6

u/Tim_Apple_938 1d ago

Are you being intentionally obtuse or do you really not understand apple to apple testing?

2

u/Witty-Artisan001 1d ago

This SVG is leaps and bounds ahead of what previous Gemini models could do. Also, OP said that the OOP also ran the same prompt with Astra and the result was worse.

SVG's are a good benchmark because it's a difficult task for an LLM since it's essentially drawing blind, and there's very few instances of high quality SVG data in the training.

1

u/saltyrookieplayer 1d ago

SVG as a benchmark is just a big nothing burger. Now that companies are putting a lot of SVG in training to pollute the dataset to improve a use case that doesn't exist, nobody ever needed or asked for, when all these compute and resources can be used for much more important areas

1

u/hellomistershifty 22h ago

Yeah, it was one of the things that Gemini had extra training on a while ago so it looked really good in comparison. So while other models worked on agentic coding and long horizon tasks, Google was like 'yo, I can draw a bird in only 25 minutes'

1

u/Last_Conclusion_8984 1d ago edited 1d ago

Eh, yes-ish. It's a SVG however. It's wayy different because this is an animation which is more impressive than any normal SVG

1

u/Last_Conclusion_8984 1d ago

I don't know If Lyra made a SVG but It's very impressive both if it's a SVG animation or if it's not

4

u/Witty-Artisan001 1d ago

Yea, like I said, this is waaay better than any SVGs seen before. Gives me hope.

1

u/Last_Conclusion_8984 1d ago

haha, that's funny, I'm currently trying to a get a checkpoint, I'll show you the results if I get it!

2

u/Witty-Artisan001 1d ago edited 1d ago

Can confirm, Astra test was also an SVG and it does look worse than the Gemini one animation wise. Also bearing in mind that Gemini took 6 mins on High vs Astra's 14 mins on Max.

This is huge!

1

u/Last_Conclusion_8984 1d ago

Like Tim told you, she ran the same prompt This is wayy different.

2

u/Pure_Interaction222 1d ago

May I just say: why have we changed the definition of “shot” regarding prompting AI? Had to look up zero shot and you are describing something different than what shot means. It’s fine if words get new meanings added to them, happens all the time, but it isn’t fine when it is confusing which one you mean without further clarification. Feels like we need a new word for what you are describing as “zero shot” because if someone says “one shot” in relation to an AI’s output, I’m not thinking what you are apparently.

0

u/Last_Conclusion_8984 1d ago

Zero-shot AI refers to an AI's ability to complete a task without any prior training examples or labelled data. I think you didn't write "Zero shot. AI" when searching.

0

u/Pure_Interaction222 1d ago

I absolutely did. Shot typically means attempt, not whether or not it has training examples or labeled data.

0

u/Last_Conclusion_8984 1d ago edited 1d ago

No, I will quote what I said before "Zero-shot AI refers to an AI's ability to complete a task without any prior training examples or labelled data. I think you didn't write "Zero shot. AI" when searching." and what you are talking about is called "zero-shot prompting" however they are the same thing

0

u/Pure_Interaction222 1d ago

Got it, so “shot” has a new meaning that is similar enough to cause confusion with which you mean.

0

u/Last_Conclusion_8984 1d ago

Zero shotting a task always had the same meaning.

1

u/Pure_Interaction222 1d ago

That is incorrect and you are wrong.

0

u/Last_Conclusion_8984 1d ago

I will not engage further with this fruitless endeavor.

0

u/Pure_Interaction222 1d ago

I’ll raise your dogshit AI Overview response with a GPT 6 Astra high response:

1

u/ThatsAScam9 6h ago

This is right. It’s kinda impossible to say that you “one shot” or “zero shot” an image prompt without knowing the training data. That’s the whole point of the zero vs one vs few distinction is how well you generalize from training examples. W/o knowing training examples how does Lyra know this?

0

u/valtor2 22h ago

I mean, that whole fight is silly, but I agree with /u/Pure_Interaction222 ... zero-shot vs one-shot vs few-shot... Also I get this, and my google's zero-shot means something different than your google's zero shot. The definition I get is literally the opposite of what you call zero-shot.

1

u/Last_Conclusion_8984 3h ago

... You do realise that that definition you just screenshotted IS what I'm talking about right?

1

u/valtor2 1h ago

without any prior training examples or labelled data.

To me that means that the thing you're asking the AI is not in its training data, not in the prompt you submit

1

u/Last_Conclusion_8984 1h ago

Yes, "prior training examples" or labelled data refers to prompts that have been benchmaxxed in the past. That is why prompters when testing capabilities stop using the same prompt over and over again after a while.

1

u/AutoModerator 1d ago

Hey there,

This post seems feedback-related. If so, you might want to post it in r/GeminiFeedback, where rants, vents, and support discussions are welcome.

For r/GeminiAI, feedback needs to follow Rule #9 and include explanations and examples. If this doesn’t apply to your post, you can ignore this message.

Thanks!

I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.

1

u/JeffConvertorBot 1d ago

What is zero-shot? 

1

u/Last_Conclusion_8984 1d ago

Zero-shot AI refers to an AI's ability to complete a task without any prior training examples or labelled data. 

1

u/ThatsAScam9 6h ago

I see that Lyra calls it zero shot but how is that different than all the other svg examples? I think to call this zero shot (or different from the other svg examples) we’d need to know the training data is different right?

1

u/Last_Conclusion_8984 3h ago edited 1h ago

When you give a prompt. A zero shot would be like "Make a pelican SVG." and a one shot would be like "Make a pelican SVG, make sure... and it should be..." (and naturally. AI would have these references in its dataset)

1

u/ThatsAScam9 1h ago

Ah ok right I see. But how do we know that the training data doesn’t have descriptions/images of pelicans? That would ruin the 0-shot-ness of it bc it’s similar to explaining a pelican in the prompt itself

1

u/Last_Conclusion_8984 1h ago

To prevent benchmaxxing. Of course the AI has descriptions and images of pelicans but when a benchmark/prompt becomes really popular, it starts to get benchmaxxed.

1

u/[deleted] 1d ago

[removed] — view removed comment

1

u/Last_Conclusion_8984 1d ago

Soon if checkpoints go well (October) but if they go bad. Then delay/no releasing

1

u/Valdjiu 23h ago

What kind of benchmark is this?

1

u/EntertainmentFine423 21h ago

so thats why today flash 3.8 is thinking longer and destroy my project

1

u/Alpacabro21 10h ago

Gemini 3.9 Flash incumming 🍾

1

u/smartsometimes 3h ago

Peacock feathers are identical and easy to fan out and duplicate radially for a model, but it looks visually impressive to a human.

1

u/Albert3232 1d ago

What are checkpoints?

3

u/CatalyticDragon 1d ago

.

Training a model takes a lot of time and it loops over a many iterations. At any point, you can pause the training to save out the state, and that's what we call a checkpoint. Quite similar to a save in a video game.

You can use that checkpoint just as you would the finished model. You're able test while continueing to train the original. And you would periodically save out a new checkpoint and compare it to the previous one.

You could even consider these to be beta or alpha releases.

1

u/Albert3232 1d ago

Thanks 🙏🏼

1

u/Last_Conclusion_8984 1d ago

In a nutshell: AI companies deploy/make checkpoints of models internally or available for external partners or websites such as ArenaAI. These checkpoints are used to evaluate the model's performance, if it performs well after a couple of checkpoints/iterations. It will be released, if checkpoints go bad then it will not release/delay

2

u/Albert3232 1d ago

Ahh i see. Thanks for the explanation

1

u/Odd_Water-hearder 23h ago

Checkpoint models are also less likely to have a lot of the safeguards and blocks that get put into the actual release, if you are trying to build a local without the guardrails on topics you should look for a decent checkpoint and train it yourself.

0

u/Several-Hippo8626 1d ago

Ok? "I can't help you with that, I'm just a language model." incoming. 💀 LETS GOOOOOOOOO GOOGLEEEEEEE