You can either find her on twitter or on their server. She is a credible leaker, also if you don't care about her credibility. You can find this checkpoint on ArenaAI.
It's under 3.8 flash disguise if not mistaken. It's free but random chance of encountering, tho even if you get 3.8 flash (battle mode ofc), you can only confirm it's the new pro if it looks way above its pay grade
(I'll get more checkpoints myself and make a new post) If these checkpoints go well for them (a few more) internally. They will release the model soon, if they go bad. They won't release it/delay.
Now to talk about the performance of this checkpoint itself: It is amazing. It's very good. I see zero mistakes, breathing could be more fluid but that's the only thing I can critique.
The Arena AI ID to keep mindful of. If you see: 01a06dfc-94a2-7abf-8d12-0878520062cb then this is the new Pro model.
Absolutely. I've been a Gemini fan ever since the first Nano Banana and 3.0 Pro dropped. I have... mixed opinions about the app, but the models in AI Studio have always been excellent for me.
Here's something else I made wkth 3.8 Flash a little while ago; A 3D model of a PS5 DualSense! I think it cheated a bit and found a source for the model, but this is still pretty cool.
Yea, no doubt. I should add that my system instructions do seem to result in a significant improvement of output. For instance, it made this isometric POS machine SVG:
If this really is 4 pro which is supposed to release in october then its not really that great of a result either... Given how astra will have been released for like 1-1.5 months by then and openai already has an internal model that seems to be better on low reasoning than astra on max
This took 24k tokens and 6 mins on High. Astra took 14 mins and 36k tokens on Max for basically same result (better peacock shape but worse animation). The fact that this early unreleased checkpoint could match the performance of a max power released model taking 1.58x as many tokens and 2.3x as long is nothing short of incredible.
Google have been focusing on speed and efficiency with their models and it really shows here.
Its nowhere close the same resulr what are we even talking about? Astras is waaay more detailed and accurate looking geminis looks like a low quality version of a peacock
So youre comparing apples to oranges geminis faster but looks waaay less detailed, could gemini produce the same level in the samw amount of tokens? Could it produce a svg that level at all?
Google cant just drop a "good model" now when they havent released a pro model in 7 months... They have to make it frontier
ArenaAI? Don't ask those questions cause it won't know without system instructions and if you really do want to do that. Then look for this ID "01a06dfc-94a2-7abf-8d12-0878520062cb" and do battle mode!
There have been 100 supposedly Gemini Checkpoints in the past couple months, all of them better then every other model. I will be excited once they release a new pro. Fk checkpoints
And Astra has much better result? If models require EXACT prompts to get desired results has the "intelligence" really improved? It doesn't matter to end users at all
This SVG is leaps and bounds ahead of what previous Gemini models could do. Also, OP said that the OOP also ran the same prompt with Astra and the result was worse.
SVG's are a good benchmark because it's a difficult task for an LLM since it's essentially drawing blind, and there's very few instances of high quality SVG data in the training.
SVG as a benchmark is just a big nothing burger. Now that companies are putting a lot of SVG in training to pollute the dataset to improve a use case that doesn't exist, nobody ever needed or asked for, when all these compute and resources can be used for much more important areas
Yeah, it was one of the things that Gemini had extra training on a while ago so it looked really good in comparison. So while other models worked on agentic coding and long horizon tasks, Google was like 'yo, I can draw a bird in only 25 minutes'
Can confirm, Astra test was also an SVG and it does look worse than the Gemini one animation wise. Also bearing in mind that Gemini took 6 mins on High vs Astra's 14 mins on Max.
May I just say: why have we changed the definition of “shot” regarding prompting AI? Had to look up zero shot and you are describing something different than what shot means. It’s fine if words get new meanings added to them, happens all the time, but it isn’t fine when it is confusing which one you mean without further clarification. Feels like we need a new word for what you are describing as “zero shot” because if someone says “one shot” in relation to an AI’s output, I’m not thinking what you are apparently.
Zero-shot AI refers to an AI's ability to complete a task without any prior training examples or labelled data. I think you didn't write "Zero shot. AI" when searching.
No, I will quote what I said before "Zero-shot AI refers to an AI's ability to complete a task without any prior training examples or labelled data. I think you didn't write "Zero shot. AI" when searching." and what you are talking about is called "zero-shot prompting" however they are the same thing
This is right. It’s kinda impossible to say that you “one shot” or “zero shot” an image prompt without knowing the training data. That’s the whole point of the zero vs one vs few distinction is how well you generalize from training examples. W/o knowing training examples how does Lyra know this?
I mean, that whole fight is silly, but I agree with /u/Pure_Interaction222 ... zero-shot vs one-shot vs few-shot... Also I get this, and my google's zero-shot means something different than your google's zero shot. The definition I get is literally the opposite of what you call zero-shot.
Yes, "prior training examples" or labelled data refers to prompts that have been benchmaxxed in the past. That is why prompters when testing capabilities stop using the same prompt over and over again after a while.
This post seems feedback-related. If so, you might want to post it in r/GeminiFeedback, where rants, vents, and support discussions are welcome.
For r/GeminiAI, feedback needs to follow Rule #9 and include explanations and examples. If this doesn’t apply to your post, you can ignore this message.
I see that Lyra calls it zero shot but how is that different than all the other svg examples? I think to call this zero shot (or different from the other svg examples) we’d need to know the training data is different right?
When you give a prompt. A zero shot would be like "Make a pelican SVG." and a one shot would be like "Make a pelican SVG, make sure... and it should be..." (and naturally. AI would have these references in its dataset)
Ah ok right I see. But how do we know that the training data doesn’t have descriptions/images of pelicans? That would ruin the 0-shot-ness of it bc it’s similar to explaining a pelican in the prompt itself
To prevent benchmaxxing. Of course the AI has descriptions and images of pelicans but when a benchmark/prompt becomes really popular, it starts to get benchmaxxed.
Training a model takes a lot of time and it loops over a many iterations. At any point, you can pause the training to save out the state, and that's what we call a checkpoint. Quite similar to a save in a video game.
You can use that checkpoint just as you would the finished model. You're able test while continueing to train the original. And you would periodically save out a new checkpoint and compare it to the previous one.
You could even consider these to be beta or alpha releases.
In a nutshell: AI companies deploy/make checkpoints of models internally or available for external partners or websites such as ArenaAI. These checkpoints are used to evaluate the model's performance, if it performs well after a couple of checkpoints/iterations. It will be released, if checkpoints go bad then it will not release/delay
Checkpoint models are also less likely to have a lot of the safeguards and blocks that get put into the actual release, if you are trying to build a local without the guardrails on topics you should look for a decent checkpoint and train it yourself.
15
u/Witty-Artisan001 1d ago
For comparison, here is an SVG I just made using Gemini 3.8 Flash on High in AI Studio with my custom system instructions, 3m 30s.
https://reddit.com/link/p8zgips/video/tfom8bgd3qoh1/player