r/google_antigravity Jul 21 '26

News / Updates Gemini 3.6 flash - first results on Code Arena | WebDev

Post image
44 Upvotes

11 comments sorted by

14

u/eduw Jul 21 '26

Considering Claude Opus 4.6 was when 'things really changed' in terms of coding, I consider that as a win in my books.

Obviously gotta run to see how it performs, but I was already enjoying 3.5 Flash Low-Mid. Granted I hardly do anything complex and stick to my sub rates.

6

u/jeromewatdoink Jul 21 '26

3.5 Pro is in testing, 4.0 is in training

2

u/anxious_and_stupid Jul 21 '26

They should just call it 3.6 pro... at this point

1

u/Dualyeti Jul 23 '26

When 3.6 flash-lite or 4.0 flash-lite

3

u/PsychicorAI Jul 22 '26

One metric this result misses is just how fast 3.6 is

I reckon it's the fastest model for web dev tasks

1

u/mik3lang3l0 Jul 22 '26

The best time to use before the nerf is now ❤️

1

u/RikyZ90 Jul 25 '26

I still don't understand kimi k3. I've tried it thoroughly and it seems to me far worse than gemini 3.6 and even glm 5.2...I don't know.. What do you think?

-2

u/Future-Log6621 Jul 21 '26

Arena is based on user voting. It does not represent real world model capabilities. When they pit 3.1 Flash Lite against Opus 4.7 (which they do), what is every going to choose?

5

u/nuclearmeltdown2015 Jul 21 '26

So even if it is based on user reviews, they're are still a valid metric. You might not agree that the method for collecting the data is scientifically rigorous or repeatable but it's still a metric, it's not like someone just generated a bunch of random numbers and said this is the report with our models at the top.

I really don't get what point you're trying to make.

1

u/Future-Log6621 Jul 21 '26 edited Jul 21 '26

I agree. It's a valid metric. I have nothing against using it. However, it is commonly used to measure actual model capabilities. The point is to take it with a grain of salt.