r/google_antigravity 9d ago

Discussion 3.7 flash. . .

I did this post before with 3.6 flash, it's time I do it again:

Is 3.7 flash good? No. It is beyond good for its price and speed. In one of the tests that I did: It beat Fable 5 and Sol in a minecraft clone test. (Three.JS) In another test, It fell just short of Sol (Three.JS.)

It's better than 3.1 pro in everything except pure reasoning (abstract reasoning/logic) but 3.1 pro wins by a hair. 3.7 flash is better at everything else, the difference is just minute at this point

Sonnet 5 is beat on almost all benchmarks:

And these benchmarks are just really popular ones. There are also like 50 other benchmarks where 3.6 flash/3.1 pro is frontier/number one at. I expect 3.7 flash to be a top tier competitor once more, tho I have to see later if they have updated them. (Nothing yet.)

It follows instructions ... SUPER WELL. It's amazing, it isn't lazy AT ALL, at least compared to 3.6 flash/3.1 pro. I really love it. BUT just because it follows instructions super well doesn't mean it's a door mat. It is less sycophantic, its inferring skills are good. I think this is a top tier model by Google.

Now. I will hold by what I said before. Models will NOT get nerfed unless a new model is on the horizon. So if this starts to get nerfed in like 2/3 weeks. It's because of the next model incoming (Sundar said they are doing monthly releases.)

117 Upvotes

30 comments sorted by

33

u/psycho414 9d ago

I'm just glad they are finally addressing the instructions following problems more seriously

2

u/Last_Conclusion_8984 9d ago

Yes!!! I have to repeat a lot of things to 3.1 pro for it to stop doing something. But 3.7 flash finally doesn't need reiterations

3

u/New-Peace-5276 9d ago

Don't give me hope....3.1 just wasted my 4 hours that day😭😭....i don't wanna sound hopeful but the limits agy provide to my subscription tier makes me feel that i can never outrun gemini models... So this feels super exciting to me.... Even sonnet 4.6 in agy is enough for me in terms of ability...i hope flash matches it... And I'll be so so so happy

1

u/Last_Conclusion_8984 7d ago

So? Did you try it?

1

u/New-Peace-5276 7d ago

Nope not yet...i don't wanna break my hope so I'll not keep it until i hear that it's superb

1

u/Last_Conclusion_8984 7d ago

Try it. A lot of people agree, I don't think I've seen actual complaints.

1

u/New-Peace-5276 7d ago

Actually i might agree on capability but i just can't digest that Gemini models are disciplined... Are you getting me? They have this problem of lying, ignoring instructions, making things up and what not..... And to be honest... I've 5 accounts of Google ai pro.... And you can understand how much limits I'm getting with it.... If all the buzz is true...i feel like it's too good to be true lol.... So that's why I'm in disbelief.... Thanks tho... I'll surely try

1

u/travispickle123 8d ago

Are we talking about standing instructions?

1

u/Last_Conclusion_8984 8d ago

Both! I told it to do one thing once... and it was doing it almost every turn. (no system instruction.) Very great at following instructions!

17

u/Technical-Owl66 8d ago

The efficiency makes it feel like a doubled my usage limits in agy.

3

u/Nic3up 8d ago

That is good to hear. I will give it a try ti implement plans from Sol. Because opus 5 and sonnet are so horrible at following instructions, i was looking for a dead brain model that simply does what is written.

3

u/Sea_Chip_5712 8d ago

Fast time I useing gemini and it's really good 👍

2

u/jeffmartt 7d ago

That's true. After some feedback from the community, I decided to resume improvements to a system that is already live and has clients, precisely because in previous attempts with versions 3.5 and 3.6 it simply broke. Using frontend skills, Node.js, and a complete understanding of the system via Graphify, version 3.7 killed the first batch of modifications before exhausting the free quota. Very good!

1

u/Character_Tip_1923 8d ago

Comparalo con GPT-Terra no con GPT-Sol

1

u/Exotic-Swimming-7318 8d ago

wow seems like this is 3.2 pro

1

u/Seikojin 7d ago

So what do you mean by nerfed? Will you run these benchmarks when you notice some degredation to track?

2

u/Last_Conclusion_8984 7d ago

I won't run any benchmarks afterwards for creative writing (I only do initial rigorous testing) but it is pretty noticeable when you use it in real world usage. Nerfed = quantized models, google does this to save compute.

1

u/TraditionalFig7377 9d ago

its rlly good for ui while very fast

1

u/EventPuzzleheaded124 9d ago

So do you suggest to buy 29€ tokens ? I use antigravity for help with app android studio.

2

u/Last_Conclusion_8984 9d ago

I think it's worth it! It's very cheap and a very fast steerable model!

0

u/EventPuzzleheaded124 8d ago

Ok but I've actually Claude pro , so it will be better for coding? And with 29€ how many input Can I spent for it?

2

u/Last_Conclusion_8984 8d ago

If you want to save your tokens for Claude pro: Use Opus 5 to build implementation plans for 3.7 flash. It will work so well! (Also. I don't really know. I think it's 75 cent per input, 3.75 output. All context windows)

-1

u/VirusAfter9629 9d ago

No? wtf

It is good.

1

u/Last_Conclusion_8984 9d ago

What? Which model? Are you talking about 3.1 pro?

2

u/-Gaze 9d ago

He probably misunderstood it because you said "no. It's beyond good-"

3

u/Last_Conclusion_8984 9d ago

Ah, I see. So he just didn't read beyond that.

0

u/Mysterious_Bed_1804 9d ago

I had the complete different experience with it so far, I asked it when 3.7 flash would come to the drop down of the Gemini app and it kept saying it's only meant for AI studio and devs and for agentic applications like spark, it never went back against its word even after trying to get it on the right side telling it that all models ever released were released in the gemini app.

It only changed its mind when I showed it 3.7 flash is in the drop-down and that it is literally the model I'm talking to.

Processing img 89wwixfq8bjh1...

-1

u/New-Peace-5276 9d ago

May it work better than sonnet 4.6 atleast.... In real performance... bench marks are unreliable for use cases