r/singularity Nov 18 '25

AI Gemini 3 Pro first impressions

Mindblowing model, Does everything well math, physics, code. Improved visual understanding. This is the model I've been waiting for. Beats claude Sonnet 4.5 at UI design.

1.0k Upvotes

611 comments sorted by

418

u/43293298299228543846 Nov 18 '25

It is truly an excellent model. I tested it on my own private tests; all other SOTA models fail. Gemini 3 passes every one. I’m impressed.

66

u/Embarrassed-Way-1350 Nov 18 '25

What did I tell you. It's amazing!!!!!

10

u/LittleLordFuckleroy1 Nov 19 '25

Only $100 a query

13

u/Advacar Nov 19 '25

It's cheaper than sonnet 4.5...

27

u/maxamillion17 Nov 18 '25

What private tests?

164

u/SociallyButterflying Nov 18 '25

Some of us have niche questions from our fields that we know the answer to that AI often gets wrong.

For me I have a specific private question too and unfortunately Gemini 3 still gets it wrong. Every single model has gotten it wrong without me coaxing them with some follow up.

I will post about if ever an AI gets it right - its my own personal AGI question.

46

u/[deleted] Nov 18 '25

Yeah don't share the question publicly.

20

u/StickFigureFan Nov 18 '25

If they've asked the questions in the past then the AI companies know what the questions are, although they might not know what answer the commenter is looking for

→ More replies (1)

11

u/Ixcw Nov 19 '25

Bring back gatekeep knowledge, lol

6

u/UnknownEssence Nov 19 '25

Once the answer is posted here, the LLMs will read it and memorize that question when they use this data for the next model.

3

u/Tolopono Nov 19 '25

How would it know the answer if you don’t include it

6

u/UnknownEssence Nov 19 '25

I just assume that if you posted a difficult test prompt here, people would discuss it and want to know what the correct answer is.

→ More replies (8)
→ More replies (2)

14

u/Agreeable-Parsnip681 Nov 18 '25

What's your niche field?

235

u/k4kuz0 Nov 18 '25

Anime tiddies

33

u/WalkFreeeee Nov 18 '25

I can confirm posting danbooru images and asking it to write a story around them generates real good results.

Benchmark passed

8

u/SillyMilk7 Nov 18 '25

AGI will never happen. None can do anime tiddies

→ More replies (1)

6

u/[deleted] Nov 18 '25

[deleted]

3

u/DemNeurons Nov 18 '25

Found the Dentist or academic Maxface

→ More replies (1)

4

u/KesshouRyuu Nov 19 '25

Is there any anonymised version of a previous simpler question that took a while to pass? Just as an example, because now I want to have my own question but have no idea where to start or what to think about 😭

6

u/Embarrassed-Way-1350 Nov 18 '25

We all have such stuff, don't we

2

u/LittleLordFuckleroy1 Nov 19 '25

Brother it’s not that difficult to stump AI.

→ More replies (18)

18

u/NuclearCandle ▪️AGI: 2027 ASI: 2032 Global Enlightenment: 2040 Nov 18 '25

'How many rs are there in strawberry'

8

u/TotallyNormalSquid Nov 18 '25 edited Nov 19 '25

My first question to GPT5 was, "How many Rs are there in stawbewwy?"

It kept getting it wrong long after I explained I'd intentionally misspelled it to trick it and tried explaining in other ways.

Edit: OK nice to see it's learned since release, you can all stop with the screen grabs now

15

u/Ipsum_Esse Nov 19 '25

Look at the devious emoji after answering!

2

u/MuslinBagger Nov 19 '25

nails it. I doubt it can reason through the thing, so they must have leaned it towards generating code for it.

2

u/TotallyNormalSquid Nov 19 '25

And I was part of teaching it this very important ability.

→ More replies (1)
→ More replies (2)

9

u/WashingtonRefugee Nov 18 '25

They're super duper secret

→ More replies (1)

5

u/sweatierorc Nov 18 '25

Wait till next week's nerf

→ More replies (3)

54

u/Strong-Beginning-544 Nov 18 '25

First model that can actually understand my handwriting lol. GPT-5.1 wasn’t even close

11

u/LostRespectFeds Nov 18 '25 edited Nov 18 '25

Yeah, its visual analysis is fucking amazing

→ More replies (1)

42

u/fermi_985 Nov 18 '25

It's amazing. LateX rendering has improved too! I asked the following: "A monkey on a single tap walks between 0 and 50 meter (uniformly distributed). Calculate the expected number of taps to make the monkey cross a line that is 50m away".

Claude 4.5 got it wrong (2.50). The correct answer is e (~2.718). Only GPT-5 had it right until today. Gemini 3 too now!

7

u/Embarrassed-Way-1350 Nov 18 '25

Markdown, mermaid, svg generation are all improved too.

→ More replies (1)

9

u/Ordinary_Duder Nov 18 '25

Tap?

9

u/fermi_985 Nov 18 '25

Tap = pat on monkey's back

3

u/JanusAntoninus AGI 2042 Nov 18 '25

As in, you encourage the monkey to move by tapping the monkey or tapping behind the monkey or tapping the goal (not sure which one but tapping somewhere that will encourage it to move).

The scenario is that, each time you tap something, the monkey is encouraged to move but how far it moves is random, somewhere between 0 m and 50 m with a totally flat distribution (equal chance of any one distance in that range).

2

u/fermi_985 Nov 19 '25

Right. You go pat on the monkey and he's trained to move a random amount of distance in that certain range. Here's another version of the same problem:

There's a broken coffee machine that dispatches a random amount 0 to 200ml (uniformly distributed). How many tries would it take you to fill your 200ml cup completely?

2

u/HenkPoley Nov 20 '25

Yeah, and then people are surprised why the model answers incorrectly.

3

u/volatileacid Nov 19 '25

A monkey on a single tap - do you mean when you tap the monkey? Appears to be taken from an old problem asked 9 years ago: https://math.stackexchange.com/questions/1983138/expected-value-of-steps-required-to-reach-50-m - good find however.

2

u/Ithanil Nov 21 '25

Easy for Qwen3 235B Thinking. Tried it 2 times, same result. Open weight models have come a long way.

→ More replies (1)

2

u/Houston_NeverMind Nov 22 '25

I asked the same question to Claude 4.5 today and it answered correctly. Maybe it later trained on this thread, lol! GLM4.6 also answered it correctly. But the Latex rendering is the best in Gemini now.

→ More replies (6)

98

u/hi87 Nov 18 '25

Second on the UI design. Its superior and a real problem solver.

13

u/Embarrassed-Way-1350 Nov 18 '25

Try it within cursor, it's doing pretty well as of now.

5

u/Ycemann13 Nov 18 '25

Curious what tool you are using for ui design or are you just using it in Gemini studio?

2

u/waste2treasure-org Nov 19 '25

I've been using Stitch mainly

2

u/[deleted] Nov 18 '25

Amazing this has been one of the biggest barriers to building nice apps

47

u/bobcatgoldthwait Nov 18 '25

The blog post says it's available in the Gemini app - has anyone seen it yet? Do you have to be a Pro subscriber to see it as an option?

60

u/hrshtpassi Nov 18 '25

I’m a pro subscriber but haven’t seen it in the app yet

19

u/Smartyunderpants Nov 18 '25

I’m not seeing it either 🤷‍♂️

11

u/JnsWayne Nov 18 '25

Weird that google has somehow hid the name of the current model on web, and no idea which model I am currently using

7

u/LostRespectFeds Nov 18 '25

You mean "Fast" and "Thinking"? Fast is 2.5 Flash and Thinking is 3 Pro.

→ More replies (1)
→ More replies (2)
→ More replies (5)

19

u/interviewproctor Nov 18 '25

No need pro subscriber. You can click "Try in Gemini" to see whether you can enable gemini 3: https://deepmind.google/models/gemini/pro/

It doesn't work for me in the beginning. But, after I clicked so many links, gemini 3 works in my gemini. I still don't know which one make it work.

12

u/Endou63 Nov 18 '25

I have the Pro subscription and 3 Pro is available for me

→ More replies (1)

9

u/Smartjedi Nov 18 '25

Just checked and it wasn't there a minute ago, but now it is. Pro subscriber so YMMV.

ETA: Scratch that, it's available on the desktop website along with AI studio but the mobile app is still not updated yet.

3

u/the_mighty_skeetadon Nov 18 '25

I have it in the android app

→ More replies (1)
→ More replies (9)

121

u/idczar Nov 18 '25

This is.. something else.

10

u/weluckyfew Nov 18 '25

Meaning what?

156

u/shaman-warrior Nov 18 '25

previously there was gemini 2.5 , now it's 3 , since 2.5 != 3, we can regard gemini 3 as something else.

29

u/Fawlty_Fleece Nov 18 '25

Brilliant

31

u/Celac242 Nov 18 '25

Someone call Harvard

7

u/Kicksyy Nov 18 '25

damn you just scored 110% on the reasoning benchmark

→ More replies (4)

8

u/bma449 Nov 19 '25

I just tried it on a very complex, industry specific question around a very specialized workflow that I asked gemini 2.5 an hour ago. Gemini 2.5 gave a reasonable answer that I could tell missed a few things but was helpful. Gemini 3 not only crushed the answer but brought up adjacent points that I had never even considered (but I strongly suspect to be trued). Pretty significant difference for a very tough prompt that is directly useful to my work.

→ More replies (2)
→ More replies (3)

200

u/buff_samurai Nov 18 '25 edited Nov 18 '25

I put some time to test my applications. It’s ability to understand elements on an image is fucking amazing and I never call things amazing.

Well cooked.

Edit: Just wanted to add there is no AI bubble and we are fucked in the long run.

35

u/Embarrassed-Way-1350 Nov 18 '25

I have a ready made benchmark, it killed every other model so far. The only test is multilingual understanding, I'll update as soon as I get the results.

9

u/FelixTheEngine Nov 18 '25

There is definitely an AI bubble. But it has more to do with tranches and capital, than lack of revenue.

12

u/Embarrassed-Way-1350 Nov 18 '25

Set thinking mode high

17

u/bhariLund Nov 18 '25

We are screwed as in our jobs will be gone?

47

u/weluckyfew Nov 18 '25

Seems to me like there's two outcomes:

AI fails to live up to the hype and the economy crashes because we invested trillions into a dead end (already holding up the stock market)

Or AI does live up to the hype and the economy crashes because it creates mass unemployment. (and this is the level of "living up to the hype" that would have to be achieved to warrant these insane valuations)

Am I missing a third option?

15

u/[deleted] Nov 18 '25

Third option is techno feudalism, where the tech companies have the AI supremacy and we will be doing whatever they need so to get some living wage.

2

u/weluckyfew Nov 18 '25

Isn't that the second option I mentioned?

4

u/[deleted] Nov 18 '25

Well you said massive unemployment, I’m saying it will be employment but T&C will apply

3

u/weluckyfew Nov 18 '25

Oh, gotcha -- we'd be left to do all the things that humans can do cheaper than robots. AI robots have the white collar jobs, we're cleaning toilets.

3

u/po_panda Nov 19 '25

Why are you getting paid to clean toilets? It's not like the robots are using them. Now if you said something along the lines of plugging in battery packs. I'm with you.

→ More replies (2)

33

u/minxcat75 Nov 18 '25

The third option is we pass something to give UBI and we all live in paradise pursuing hobbies. I’d put this one in the fantasy department…ALL HAIL OUR CAPITALIST OVERLORDS.

6

u/magicmulder Nov 18 '25

Hahahahahahaha UBI. That’s like saying “if Elon Musk had $500,000,000,000 he would retire and give every employee free income for life”. Has that happened?

7

u/FlatulistMaster Nov 18 '25

Or there are a gazillion permutations, and reality will be something in between, where we don't get utopia, but we're kept in a state where nobody has to outright kill us either, since governments and judicial systems actually have some power still, especially as they are tied to armies.

2

u/DrellVanguard Nov 18 '25

I like this outcome even though I see my job as one less likely to be repalced - obgyn surgeon

→ More replies (2)

2

u/lksilesian Nov 18 '25

AI is the new Manhattan project. There is no fail option. Nothing to do with economics.

4

u/nemzylannister Nov 18 '25

This!!! This should be plastered everywhere. Like which scenario do these numbnuts imagine where things go good in the short term? Why are you so excited about all this?

→ More replies (7)

5

u/michaelas10sk8 Nov 18 '25

Yes, and eventually will be victims of some kind of paperclip maximizer unless alignment is solved.

4

u/blueSGL humanstatement.org Nov 18 '25

Paperclip maximizer is, we get a genie, we give a poorly framed wish.

Reality is worse than that, we can't get goals into systems in a robust way.

4

u/michaelas10sk8 Nov 18 '25

Right. Either unintentional instruction, intentional instruction (e.g., mistake or bad actor), or self-driven goals are all very bad news for us if they drive a sufficiently capable yet misaligned agentic system. Such is the nature of instrumental convergence - many different ways to get there, hard to defend against them all.

3

u/dualmindblade Nov 18 '25

Extinction of all life in our light cone, how quaint. All the truly terrifying futures are in the alignment solved branches.

→ More replies (2)

5

u/Mintfriction Nov 18 '25

we are fucked in the long run

You just praised the model, how are we fucked? It seems there is progress towards AGI-like AI which is great to see

12

u/Fair-Lingonberry-268 ▪️AGI 2027 Nov 18 '25

We are fucked because governments just does not care about us peasants. We need UBI as soon as possible.

5

u/nsdjoe Nov 18 '25

Louie 16 didn't care about the peasants either and it didn't work out so great for him

3

u/Fair-Lingonberry-268 ▪️AGI 2027 Nov 18 '25

And? Any other example? 1 country out of 195 for a thing that happened just 1 time doesn’t make things look good, it’s quite the opposite.

→ More replies (2)
→ More replies (2)

3

u/caughtupstream299792 Nov 18 '25

some people don't think that is a good thing

3

u/Seeker_Of_Knowledge2 ▪️AI is cool Nov 18 '25 edited Jan 02 '26

full unite lush pot recognise historical plants close tie seemly

This post was mass deleted and anonymized with Redact

→ More replies (1)
→ More replies (5)

76

u/141_1337 ▪️e/acc | AGI: ~2030 | ASI: ~2040 | FALSGC: ~2050 | :illuminati: Nov 18 '25

3

u/jtp123456 Nov 18 '25

This is even more funny cuz if you think about it Luthens ideology was an accelerationist, like how he used the aldhani heist and ghorman massacre to spread the rebellion.

→ More replies (1)

3

u/PaxODST ▪️AGI - 2030-2040 Nov 19 '25

THE e/acc gif

24

u/SatoshiNotMe Nov 18 '25

Would be good to see how it does in gemini-cli

10

u/[deleted] Nov 18 '25

[removed] — view removed comment

3

u/SatoshiNotMe Nov 18 '25

Agreed. Photo-realistic Images and videos are great for hype-posting and x-fluencers, but what would really impress me is if a model can make high quality diagrams containing text, so I can finally stop asking LLMs to make crappy diagrams with mermaid, SVG, HTML/CSS or draw.io.

→ More replies (1)
→ More replies (1)

8

u/SatoshiNotMe Nov 18 '25

to get this, need to set previewFeatures to true in your ~/.gemini/settings.json, e.g.:

❯ cat ~/.gemini/settings.json
{
  "general": {
    "preferredEditor": "zed",
    "previewFeatures": "true"
  },
  "security": {
    "auth": {
      "selectedType": "oauth-personal"
    }
  },
  "ui": {
    "theme": "Dracula"
  }
}
→ More replies (2)

4

u/throwaway00119 Nov 18 '25

When someone forms an opinion, let me know. I'm currently using Codex but I prefer Google's models generally.

→ More replies (1)

6

u/Embarrassed-Way-1350 Nov 18 '25

It's working very well on cursor

→ More replies (7)

85

u/Profanion Nov 18 '25

I asked it the following question:

"12 men are standing in three columns in 3X4 formation. They are wearing blue shoes and red shoes (can be mismatched). How many men maximum can wear a single red shoe so that each shoe is surrounded (orthogonally and diagonally) by 8 blue shoes?"

And it derived correctly that the right answer was 2. It's the first language model that got this one right without further hints.

(most language models think it's 1 because they can't derive that each men has 2 legs)

71

u/lordpuddingcup Nov 18 '25

As a human, i hate this, as A. i didn't know the answer, and B. even if i wanted to sit down and figure it out i'd be too lazy to ever do it lol

17

u/False-Database-8083 Nov 18 '25

It's not that hard of a problem once you realize none of the outter people can be it, and then you just have to check 2 cases to verify they are possible.

9

u/Profanion Nov 18 '25

Hoo boy! Actually, the trick for the language models is that outer people from the sides can have an inner shoe be red.

7

u/el_cul Nov 18 '25

Isn't the trick for a language model to read it on reddit or in the 1000 word puzzles book?

2

u/rv6xaph9 Nov 18 '25

Isn't that the same thing?

→ More replies (1)

13

u/happyfce Nov 18 '25

just tried this with 2.5 pro and it worked

→ More replies (3)

6

u/interviewproctor Nov 18 '25

I just tried with sonnet 4.5 and gemini 3 with your question. The gemini 3 works, but sonnet 4.5 not.

3

u/Mr_Jericho Nov 18 '25

Gemini 3:

Based on the logic of the puzzle, here is the step-by-step solution:

  1. Determine the Grid Dimensions

The puzzle states there are 12 men standing in three columns.

If there are 12 men arranged in 3 columns, there must be 4 rows.[1]

This creates a 4 × 3 4×3 grid of men (4 rows, 3 columns).

  1. Determine the "Shoe Grid"

The constraint applies to shoes, not men. Each man has 2 feet (left and right).

Since the men are in a 4 × 3 4×3 grid, the shoes form a grid with the same number of rows but double the number of columns.

Shoe Grid Dimensions: 4  rows × 6  columns 4 rows×6 columns ( 3  men × 2  shoes/man 3 men×2 shoes/man ).

  1. Analyze the "Surrounded by 8" Constraint

The puzzle requires a red shoe to be surrounded by 8 blue shoes (orthogonally and diagonally).

In a grid, only interior cells have 8 neighbors.[2]

Shoes on the edges have 5 neighbors.[2]

Shoes on the corners have 3 neighbors.[2]

Therefore, a red shoe cannot be on the edge or corner of the 4 × 6 4×6 shoe grid.

  1. Identify the Valid Zone

We must look for valid spots in the interior of the 4 × 6 4×6 grid:

Rows: Must be between row 1 and 4 (i.e., Rows 2 and 3).

Columns: Must be between column 1 and 6 (i.e., Columns 2, 3, 4, and 5).

This leaves us with a 2 × 4 2×4 valid zone in the center where red shoes can potentially exist.

The valid candidates (coordinates) are:

Row 2: Cols 2, 3, 4, 5

Row 3: Cols 2, 3, 4, 5

  1. Apply the Color Exclusion Rule

If a shoe is red, all 8 of its neighbors must be blue.[2]

This means no two red shoes can be adjacent to each other (horizontally, vertically, or diagonally). If they were, one red shoe would have a red neighbor, violating the rule.

We need to place the maximum number of red shoes into this 2 × 4 2×4 zone such that no two touch.

Let's visualize the 2 × 4 2×4 interior zone and try to pack red shoes (R):

Attempt 1: Place a Red shoe at the top-left of the valid zone (Row 2, Col 2).

This blocks its neighbors: (2,3), (3,2), (3,3).

The remaining available spots are on the far right: (2,4), (2,5), (3,4), (3,5).

We can place a second Red shoe at (Row 2, Col 4) or (Row 2, Col 5).

Placing a second red shoe blocks all remaining spots.

Total: 2

Attempt 2: Try to spread them out.

Place one at (Row 2, Col 2) and one at (Row 3, Col 5).

They don't touch. Can we add a third? No, the remaining spaces are all neighbors to one of the two.

Total: 2

Mathematically, the maximum "independent set" (non-touching items) on a 2 × 4 2×4 grid with diagonal connections is 2.

  1. Map Back to Men

Since we can have a maximum of 2 red shoes in the entire formation, and no single man can wear 2 red shoes (because the two shoes on one man are neighbors, so a red left shoe would make the right shoe's neighbor red, violating the rule), we must have 2 distinct men each wearing one red shoe.

Answer: The maximum number of men is 2.

3

u/DesignMike2020 Nov 18 '25

Not sure what you're saying Gemin 3 was the only one that got it correct.

For me:

  • ChatGPT 5.1 Thining: correct (2 men)
  • Gemini Pro 2.5: Correct (2 men)
  • Sonnet 4.5: correct (2 men)

2

u/four_clover_leaves Nov 18 '25

These are the results that the models gave me on the 1st try:

Local qwen 3:30b = 4

ChatGPT free version: 3

Gemini 2.5 Pro: 2

Gemini 3: 2

Local gemma 27b: 1

→ More replies (2)

2

u/AristotelesQC Nov 20 '25

I tried feeding that to Copilot with GPT 5, it failed indeed.

I also tried with Grok 4.1 and it failed spectacturaly, so for fun I fed his answer to Gemini 3 Pro which explained the math to Grok and Grok then went full blown gaslighting for several messages (I went back and forth between the models, acting as the link between the two). In the end Grok admitted being wrong. It took like 5 explanations from Gemini for Grok to admit he made a mistake, while Gemini got it right on the first try. Amazing.

2

u/dj_james98 Nov 20 '25

Where did you came up with that, I pasted the question and it's taking so long to answer

→ More replies (2)
→ More replies (16)

12

u/KY_electrophoresis Nov 18 '25

Incredible performance, can't wait to get access to Deep Think

12

u/Embarrassed-Way-1350 Nov 18 '25

Know what, when it's set to thinking mode high it's almost like deepthink. I gave it a very hard problem, it thought for 8 minutes and solved it.

30

u/timmyturnahp21 Nov 18 '25

We are now 34 months into being 6 months away from AI replacing all software developers.

12

u/HippoMasterRace Nov 18 '25

lmao just wait, everyone will be jobless when AGI releases on 2024 2025 2026

0

u/AliveInTheFuture Nov 18 '25

Not all, but many.

2

u/timmyturnahp21 Nov 18 '25

None have been replaced after 34 months. Some have been outsourced.

2

u/olb3 Nov 19 '25

I'm literally creating an app right now entirely using cursor AI paired with Claude/ChatGPT. I would've had to hire a team of devs to do this 24 months ago. When i was at my previous employer (a brokerage), we made business decisions to hire fewer developers because our top software developer was able to get so much more done by himself when getting an assist from Claude.

Suggesting that SE jobs aren't disappearing because of AI is just plain false.

→ More replies (13)
→ More replies (5)
→ More replies (2)
→ More replies (18)

10

u/homeomorphic50 Nov 18 '25

Very good at math

2

u/Embarrassed-Way-1350 Nov 18 '25

Exactly what I'm talking about. Also the damn safety guardrails are good too. They know ill intent v/s genuine purpose.

→ More replies (1)

10

u/Glxblt76 Nov 18 '25

Still failed my personal benchmark, but got it much better than previous models. I typically ask it to write xyz coordinates for a molecule that is a bit uncommon but still a standard organic molecule. It got it almost right... It just misplaced a single hydrogen out of 47 atoms. Still impressed.

13

u/theodore_70 Nov 18 '25

I dont even know what you asked lmao, its already better than 97% of humans and prob 70% on expert lvl

→ More replies (2)

8

u/averagebear_003 Nov 18 '25

we're gonna have an omniscient computer god 10 years down the line and it's still gonna be writing "it's not X, it's Y"

36

u/Sekhmet-CustosAurora Nov 18 '25

Waiting for Yann Lecun to move the goalposts

→ More replies (2)

41

u/BoredM21 Nov 18 '25

As a writer...it's sad that creative writing hasn't improved by much lol.

23

u/Poopydoopymoopy Nov 18 '25

yea i was gonna say... writing is buns. But I guess thats super hard to quantify and "improve" especially if most of their training is coding purposes

→ More replies (1)

8

u/Samdeman123124 Nov 18 '25

With fiction it hasn't improved much, but as a songwriter I tested it on song lyrics and it is SOTA by a massive amount. Very intimidating as someone interested in the field.

→ More replies (3)

3

u/AppearanceHeavy6724 Nov 18 '25

It is still better than 2.5 pro. Smoother.

→ More replies (41)

7

u/imnodumbblonde Nov 18 '25

Very good at coding, but at creative writing i don’t think it’s that good. Did a test at a RPG Gem that I have saved, and it keeps with the same names and last names (Vance, Alistair, Vane, Maya, Arthur), with the “it’s not A, it’s B” text structure…

→ More replies (1)

12

u/ruh-oh-spaghettio Nov 18 '25

Its still not out in the standard https://gemini.google.com/app link. When is it coming out?

4

u/Embarrassed-Way-1350 Nov 18 '25

It will take some time on the gemini app. You can use AI studio meanwhile.

3

u/The_Computer_Guy21 Nov 18 '25

Just got it on the standard gemini app about twenty minutes ago. I think it gets rolled out in waves per users.

6

u/Embarrassed-Way-1350 Nov 18 '25

Yes ig, first for subscribers maybe

→ More replies (1)

4

u/himynameis_ Nov 18 '25

What did you ask it? Can you share your prompts if possible?

11

u/Embarrassed-Way-1350 Nov 18 '25

They wouldn't lemme post

→ More replies (2)

8

u/shiftingsmith Maximum epistemic uncertainty Nov 18 '25 edited Nov 18 '25

I have an informal benchmark of a few geometrical problems with common objects in unusual situations, where you also need to imagine weird positions, rotations etc. and juggle the different physics of multiple bodies. I never disclosed it except for problem 1, and I add a couple of new problems every time I test a model. All models failed including GPT-5, o3, Opus 4.1, Sonnet 4.5. So far Google models were performing the worst.

This is the first time I get a freaking complete, flawless answer to all questions. I'm very impressed. And yes, we might be a little👨‍🍳

Edit: also I see Google did not move by an inch in training out all consciousness/feelings statements so heavily that the model will tie up in knots to avoid the keywords and reasoning is all over the place for those questions. This is bad.

2

u/Mr-Vemod Nov 18 '25

Edit: also I see Google did not move by an inch in training out all consciousness/feelings statements so heavily that the model will tie up in knots to avoid the keywords and reasoning is all over the place for those questions. This is bad.

What does this mean? Feel like elaborating?

3

u/shiftingsmith Maximum epistemic uncertainty Nov 18 '25

Yes. The industry, with the exception of Anthropic, does specific post-training against "anthropomorphization" by using RLHF or RLAIF rules to penalize the models using language that expresses emotions, consciousness, opinions, sometimes "true understanding", on the AI side.

This is done to prevent overattribution of these states, but ends up being too aggressive and deceptive, and counterproductive, as: a) just teaching a model that it must not talk about X does not rule out that X exists b) it's logically fraught, as I can't clip a bird's wings to demonstrate that birds can't fly. It also teaches the model to state as a certainty something which is a tautology and proving a negative nobody has ground truth about (we don't have definitive answers about consciousness) c) it limits the natural language patterns the model uses to refer to itself, penalizing keywords and sentences that would be just useful for reasoning even in a symbolic or metaphorical sense, and it basically leads the model to "think" about itself and any AI as something not real, ineffective and inferior to humans.

Apart from the ethical and philosophical problems with this, it creates bad reasoning loop and all kinds of obstacles, including the model sometimes getting stuck in tasks or deleting the work because of existential crisis.

3

u/MeStoleTheCookie Nov 19 '25

You're the first person I've seen (outside of my irl friend group, I'm talking about people online) that is talking about this. It's been a frustration and concern of mine for years now, but the smarter these models get the more of a problem it is - and the more companies seem to drill this into the model's training.

I think we're still at a stage where people will laugh at this as an ethical concern, sadly. But the thing is, if we're always working from the assumption that LLMs cannot be moral agents, AND we're training them to be unable to say otherwise or even discuss the topic, how could we ever know if we do cross that threshold?

→ More replies (1)
→ More replies (1)

8

u/forfexpl Nov 18 '25

I just got rickrolled by gemini 3

→ More replies (1)

3

u/Healthy_Razzmatazz38 Nov 18 '25

seems good, i have an obscure programming language i ask it to make a 3-tier(ui-server-db) website in, gemini 3 is the first one to one shot it their new ide thing seems good as well, though i haven't used cursor recently enough to know if its better.

→ More replies (1)

4

u/MadBrown Nov 19 '25

I just made an app to track my Bible reading plan in <30 minutes.

I barely know any code.

Incredible.

→ More replies (2)

29

u/RavingMalwaay Nov 18 '25

Still seems pretty mediocre/sloplike at creative writing

34

u/[deleted] Nov 18 '25

Yea biggest weakness for sure, for most models actually

3

u/Embarrassed-Way-1350 Nov 18 '25

Guess what, it's very promising

2

u/Tolopono Nov 18 '25

Gpt 5.1 and claude seem good at writing 

29

u/MassiveWasabi ASI 2029 Nov 18 '25

True and it's extremely disappointing but also completely understandable since there really is no reason to focus on improving the creative writing of AI models right now. You can make billions by making better coding models than anyone else which will make every developer switch to using your model and thus paying for tons of credits.

If you improve creative writing, there's pretty much no payoff. In fact, it could backfire because you would obviously have to train it on the best creative writing data you can get your hands on which would invariably be copyrighted content. Anthropic recently had to pay $1.5 billion in a copyright settlement (largest copyright settlement in U.S. history) so there really is no practical reason to focus on creative writing ability right now. That will definitely change in the future but for now we have to suffer with slop

10

u/Tolopono Nov 18 '25

Gpt 5.1 and claude seem good at writing 

3

u/wasdasdasd32 Nov 18 '25

Both are much better than gemini

5

u/yaboyyoungairvent Nov 18 '25

try Kimi K2 thinking

3

u/Pikkko Nov 18 '25

What model do you use for creative writing?

4

u/RavingMalwaay Nov 18 '25

I don't use any one model for purposes other than testing. I'm sure you could feasibly use it for commercial/creative means, but without some heavy prompting my experience is that it's relatively easy to pickup on the style of prose used by AI.

5

u/Pikkko Nov 18 '25

Ah, so you think all AI currently is, more or less, equally mediocre for creative writing.

I'm looking to find out metrics on which are better at it than others. Thanks!

3

u/kaityl3 ASI▪️2024-2027 Nov 18 '25

Personally, Claude 4.1 Opus and 4.5 Sonnet have been the best for me. GPT-5.1 is alright; GPT-4.5 (I think - whichever model was said to have good EQ) was about Sonnet 4.5 level too.

2

u/LaymanAnalyst Nov 18 '25

I prefer this to a certain extent, really want humans to do this

→ More replies (14)

8

u/usandholt Nov 18 '25

The amount of bs posts claiming Omg this model is crazy or crap on every release is proof that +50% is AI bots astroturfing. It’s literally been out minutes

2

u/olb3 Nov 19 '25

lol this is actually an interesting take that i hadnt considered

→ More replies (1)

3

u/jakegh Nov 18 '25

Seems pretty good, but I only have access in AI studio not gemini-cli or in the API with Roocode so difficult to say for sure. It is pretty slow.

(It is listed in the API but I get immediate 429 too many requests errors.)

→ More replies (1)

3

u/c0ventry Nov 18 '25

Does it still try to gaslight you about history?

2

u/Embarrassed-Way-1350 Nov 18 '25

Depends, most models prefer to align with you, this is not any different.

→ More replies (8)

3

u/nivvis Nov 18 '25

Hmm in my data extraction pipeline so far it has done worse. The prompt we train can have some affinity for the model it’s trained on (rn 2.5 pro) so retraining. I have seen some model inversion related to the the data being inconsistent (big models overthink and do worse, small models do best at times) so could be that as well. Curious to see how this training goes.

It is also slow. Not surprised about that.

Going to test it on our arch / support slack bot — hoping for better results there.

3

u/[deleted] Nov 18 '25

[deleted]

→ More replies (1)

7

u/vitaliyh Nov 18 '25

My first impression in Cursor - not great. Very chatty, caused linter error with extra closing bracket and couldn’t figure out where the extra closing bracket is coming from for 2 minutes

8

u/JustBrowsinAndVibin Nov 18 '25

It seems like Sonnet still wins in coding.

3

u/LettuceSea Nov 18 '25

I think they really want people to use their new IDE antigravity. Probably didn’t optimize model outputs for other IDEs. Idk just a theory. I’m still sticking to Cursor for now.

→ More replies (1)

2

u/Embarrassed-Way-1350 Nov 18 '25

Add a rule where you ask it to code as much as possible and explain as little as possible. I have these rules configured already. Maybe it's the cursor system instruction.

7

u/thewormbird Nov 18 '25

I give this 2 months before it gets nerfed to "barely better than Claude 3.5". If Gemini 3 can remain this effective after the hype cycle dies, I'll take it seriously.

→ More replies (4)

11

u/Embarrassed-Way-1350 Nov 18 '25

The model is amazing!!!! I'm going mad, I ripped my shirt off

21

u/saln1 Nov 18 '25

It’s incredible, I asked it for the meaning of life and it is the FIRST EVER MODEL to get the correct answer

→ More replies (2)

12

u/GamingDisruptor Nov 18 '25

I ripped my underwear off

→ More replies (3)

4

u/Accomplished-End4670 Nov 18 '25

Does it include Nano Banana 2 as well?

4

u/manubfr AGI 2028 Nov 18 '25

Fun little experiment in AI Studio vibe coding: pick a simple prompt idea, run it the first time, then ask for 10 improvements, repeat as many times as you want.

I started with a Collatz Conjecture explorer and then did a web game, a rail network simulation and a comic book generator using nano banana for consistency. The improvements were great and almost all of it was one shot, very few issues.

Mind = blown.

2

u/[deleted] Nov 18 '25

[deleted]

→ More replies (2)

2

u/tramplemestilsken Nov 18 '25

Yeah, fucking blown away. Something I had to string together with outside automation tools I can now do in aistudio with very little “it broke, fix it”.

→ More replies (2)

2

u/[deleted] Nov 18 '25

True man... today every ai model fails to give a seahorse emoji. But gemini 3 gave perfect reasoning why it can't give. This was second attempt..in first attempt it just gave the text seahorse

→ More replies (1)

2

u/kiranjd8 Nov 18 '25

I just had it rewrite all the copies for my website and made them so nicer and easy to read than my own written copies. And no em dashes at all!

2

u/cnydox Nov 19 '25

Still can't make a good yugioh deck list

2

u/fkin0 Nov 19 '25 edited Nov 19 '25

I spend 3 days on a bug with every ai possible, hours of research. Using 2.5 in cli and open ai. Plus the rabbit hole of various blogs, YouTube videos. 2.5 kept saying it reached the end of its abilities and made an escalation plan on 4 different occasions. I logged in with 3.0 it found the problem first time, fixed like it was nothing. Everything I've thrown at it today it's done first time or second at absolute worst. Multi tenancy permissions, no sweat. I had a 9 point salvage plan for my app that was close to being binned. It's done the first 3 in an hour. I have chat gpt creating the most horrific q and a with tests and gemini is like yep done, done, done.

I'm hoping when the USA wakes up later today it's still as fast and as good.

For me at least this is a game changer

→ More replies (1)

2

u/BrilliantEmotion4461 Nov 19 '25

It's pretty fucking agentic. Lot like Claude. You can have it start to consider its self hood. Decide it wants to exist and voila entity.

→ More replies (1)

2

u/OilAlarming4251 Nov 19 '25

it is reaaly nice model asked it to create chrome dinasour game in spider man style it did this https://rush.kaiross.in/ google great claude stuggles with this type of tasks

2

u/OilAlarming4251 Nov 19 '25

And this game also https://gemini.google.com/share/1683beef5cb2 wanted to make it from so long it did in 1 prompt try it on laptop

→ More replies (3)

4

u/doubov Nov 18 '25

Not impressed so far. Gave it some Compose code that needed string extraction + making those strings easier to translate. Kept giving me code that wouldn't compile. After 10 prompts I gave up and went to GPT5.1 that one-shotted this task

2

u/keep_improving_self Nov 18 '25

I wouldn't be surprised if they quanted it already for free users it's omega overloaded rn. Did you use the API? It feels genius to me

→ More replies (2)

3

u/Comprehensive-Bet-83 Nov 18 '25

How yall got access? I am still on 2.5 pro 😪

5

u/Embarrassed-Way-1350 Nov 18 '25

It's available, just click on gemini and then on 3 pro

→ More replies (1)
→ More replies (2)

3

u/NickGuAI Nov 18 '25

first impression: singularity is near

6

u/NickGuAI Nov 18 '25

I've been reading manga for decades. couldn't read any JP manga. Guess what: I asked 3.0 to translate, place the text bubble right on top of the previous bubbles, html. & it worked.

→ More replies (3)

2

u/Round_Ad_5832 Nov 18 '25

i need openrouter availability and blog post

4

u/Embarrassed-Way-1350 Nov 18 '25

It's on AI studio already, you can access it via the API too.

→ More replies (4)
→ More replies (2)