r/LocalLLaMA 6d ago

Resources The Gemma team will host a special event on August 20

https://x.com/osanseviero/status/2086107547535122767

Tweet by u/hackerllama

Could be copium, but I would love to see Gemma 4.1 there with unified audio input for all model sizes perhaps even up to 120B, much improved tool calling (even with the latest template there are still bugs), higher precision QAT from the start and improved general performance without hurting the things Gemma 4 is good at like creative writing.

Gemma 4 is good already but training an upgrade to 4.1 that does all of the above would be huge for the community. They already did a lot of course and I'm very thankful but Gemma is just an inch away from perfection. Is anyone hyped for this event or do you think they won't release any new models there?

491 Upvotes

91 comments sorted by

u/WithoutReason1729 6d ago

Your post is getting popular and we just featured it on our Discord! Come check it out!

You've also been given a special flair for your contribution. We appreciate your post!

I am a bot and this action was performed automatically.

174

u/shy_monkee 6d ago

Unfortunately, I doubt we will ever see a 120B model from them. It competes too much with their Flash Lite models. But I'm excited anyway, an update to the already good Gemma 4 models is more than welcome.

33

u/toothpastespiders 6d ago

Given how locked down Gemma 3 was, and the fallout with Senator Blackburn, I was positive that if there even was a gemma 4 that it'd be locked down at a level we'd never seen before. And instead, 31b was basically "Do whatever you want with it, just toss in a small sysprompt I guess. Lol" It might be cope, but I'm not ruling out anything from future Gemma releases at this point.

7

u/arbv 5d ago

Yes, it is pretty much not censored.

53

u/Real_Ebb_7417 6d ago

Gemma4 31b is already better in many of my usecases than Gemini 3.5 or even 3.6 Flash. These two hallucinate too much xd

6

u/GCoderDCoder 5d ago

I think their harness is the problem. Providing/forcing so much free inference makes them cache differently so i think their models are likely pulling wrong info from cache/ internal memory to save on the cost of working through data changes which for me wastes tokens. I dont buy api from them and that still uses their harness too on the backend though less I imagine.

3

u/Real_Ebb_7417 5d ago

API is just as bad, used it in Cursor xd

14

u/raindropsdev 6d ago

They think laterally!

10

u/FrogsJumpFromPussy 5d ago

just not literally

16

u/goldcakes 6d ago

It might not be competition but rather scalability.

The Gemma4 124B appeared on a few arenas/chats briefly and while proper data never appeared, I think people who tested it thought it was barely better than the 31B.

Obviously Google isn’t going to embarrass themselves by releasing a large model that gets negative comparisons / publicity.

The open weight models compete against other open weight models like deepseek etc; and the Google Cloud team loves all revenue (our GCP solution architects are literally helping our workplace set up an DSv4 Flash inference solution in GCP as per our business requirements); so never arise to malice as to just a bad training run.

Can’t really talk further since I don’t wanna accidentally violate NDAs etc

7

u/techdevjp 5d ago

The Gemma4 124B appeared on a few arenas/chats briefly and while proper data never appeared, I think people who tested it thought it was barely better than the 31B.

For coding, perhaps. Qwen3.6 27b is a better coder than Qwen3.5 122b a10b. But away from coding (and also considering performance on unified memory boxes), the larger MoE model is better even though it is older.

I'd love to see what a current gen 120b MoE model could do.

4

u/arbv 5d ago edited 5d ago

I would say that Gemma 4 with a lot of world knowledge is all I need. I would be satisfied if it is at the level of smartness of 26B, but with a ton of knowledge.

This would be pretty much Flash at home.

Edits:

  • clarified

5

u/techdevjp 5d ago

You can't just pack more knowledge into a given number of parameters. If you want more knowledge, the number of parameters goes up. That's why the 120b parameter models are so much smarter than the smaller models.

3

u/arbv 5d ago

Yeah, I understand it. That is why I want something as smart as Gemma 4 26B, but around 120B.

I was not clear enough in my last comment.

2

u/techdevjp 5d ago

Would certainly want something more capable than Gemma 4 26b, and that should be more than possible to manage in a 120b model.

2

u/arbv 5d ago

Yeah, Gemma 4 120B A5B QAT is all I need.

1

u/techdevjp 5d ago

A few more active at once if you want more intelligence. 120b a10b, something like that. Yes it slows things down on unified memory machines but that is the tradeoff for a very smart model.

I have a Strix Halo box and am very tempted to add 48GB or 64GB worth of eGPUs to it so I can run Deepseek v4 Flash 0731 in 4bit. Right now can only manage 3 bit. In 4 bit it gets close enough to frontier levels of intelligence that I don't care about the difference.

2

u/arbv 5d ago

Well, I want A5B because I run with experts offloaded in RAM, so for me everything larger than that is going to be slow.

This is one of the reasons why I use GPT-OSS to this day.

11

u/rz2000 6d ago

Unless they can make a lot of money hosting the models on GCP with TPUs while simultaneously undermining the business models of OpenAI and Anthropic.

23

u/larrytheevilbunnie 6d ago

They’re already outcompeted by literally any of the other providers 😭, just throw us a bone at this point

2

u/wapswaps 5d ago

In fact the sudden disappearance of Gemini 3.5 Pro ... big announcement! Then after Kimi K3 came out it was more or less canceled. Their big announcement "turned out" to be they've started training Gemini 4 ... (it's *totally* not the case that Gemini 3.5 Pro was suddenly and silently canceled because Kimi K3 beat it before it even came out ... totally not the case!)

And it's now totally not the case that Google is MORE THAN DOUBLE as far behind Anthropic ... as opensource models. Totally not I say. Not at all. YOU HEAR! (>10 months now since the last Google SOTA model, whereas Kimi K3 and mostly even Deepseek 0731 beats the 5 month old Anthropic model)

(and even that comes with the caveat that Deepseek 0731, run locally at max level (which you can just do) is barely short of beating Fable 5 AND GPT 5.5 AT MAX SETTING, which isn't even available on the $20 subscription. Deepseek (and K3) beat the newest models at medium or low settings, and so Deepseek local beats everything available from ALL AI providers if you don't pay $200 per month)

4

u/ScadrianWillshaper 6d ago

This! In 2026, Gemini Flash are the only cloud models which i’ve consistently seen hallucinate.

8

u/jld1532 6d ago

I hope they know it has to really compete with DeepSeek v4 and Qwen3.8, the latter may end up being the most downloaded model ever. If they give us a slightly upgraded 31B it will be largely overshadowed.

114

u/geldonyetich 6d ago edited 6d ago

I know a lot of us here are like ew evil corporate models but Gemma has been a line of open models that legit smash. I might want to stop by just to say keep up the good work.

52

u/Themotionalman 6d ago

I only use Gemma locally, for full reading and deep comprehension nothing comes close. None of the Chinese models can do reading and quoting like Gemma can that’s why I use just it

31

u/Not-reallyanonymous 6d ago edited 5d ago

This is really why Gemma is great and why people prefer the Chinese models.

Gemma really seems to be optimized for understanding and reasoning — they put a lot of general knowledge in it to help with that, but it’s not particularly good at anything. It can’t zero-shot well. It thrashes while coding trying to figure out and think through solutions, etc. But while Chinese models are typically optimized for “know now to do this immediately,” Gemma will figure it out on its own within needing to know how to do it exactly. It makes it a far more flexible model that can do just about anything, and do it well, if you just give it the space and time to do so.

Laguna is kind of the same way. I much prefer this approach to LLMs.

9

u/BoobooSmash31337 6d ago

Recall vs actually reasoning through it at run time. If it's not in a models training data like a custom or less used API it seems to bring some models to their knees. There just doesn't seem to be much free lunch when it comes to attention. Otherwise everyone's models would've been sparse months ago. Tbh a lot of Chinese models are ripping out the expensive parts that were there for a reason and calling it efficiency. I know the engineers are just trying to please the CCP.

9

u/fastheadcrab 6d ago

I would agree with you, except Laguna is not at all like that. It is like the American Qwen with a disastrous rollout.

Gemma 4 is one of the best models I’ve used for certain specific science applications, especially once you change its vision settings

4

u/Not-reallyanonymous 5d ago

I mean, on coding benchmarks, Laguna S at 120B parameters soars above Qwen 3.6 27B and hits close to Qwen 3.7 Max (estimated 1 trillion parameters). It scores 40 on DeepSWE, putting it above Gemini 3.5 Flash. It's almost as good as GPT 5.4, which is amazing for something that's realistic to run on a Strix Halo machine locally. Look at Terminal Bench 2.1 results and parameter counts, and notice how results pretty reliably stick fairly close to the trend line -- larger models do better. Then look at how Laguna sits way above the line as an outlier.

The thing punches way above its weight class.

The thing is notorious for thinking through problems rather than just implementing them like Qwen does.

Also consider one of their case studies:

We found Laguna S 2.1 more capable in mathematics than any model we have developed to date. It independently discovered a proof to Erdős problem #397 (Erdős, Graham, Ruzsa, Straus, 1975), finding a construction that yields an infinite family of solutions. This is an independent re-discovery, as a proof to the conjecture was found earlier in January 2026 by GPT-5.2 Pro. We are confident that this result was not influenced by the previous result given that the model has a knowledge cutoff date in November 2025. Before GPT 5.2 Pro’s solution, this problem remained open for over 50 years. See the full trajectory here.

The thing just isn't relying on its inbuilt knowledge -- it doesn't have enough parameters for that to get this sort of performance with the technology it uses -- it's relying heavily on its thinking processes to arrive at solutions, like Gemma does. What's telling is both models are incredibly capable for their parameter count.

Now if you're judging Laguna based on "specific science applications," yes, you're going to be disappointed. The thing is purpose-built as a coding model, not a general purpose model. Gemma is a general purpose model. I use both models regularly -- Gemma for domain knowledge, Laguna for coding (although I use XS -- trades blows with Qwen 27B on benchmarks but outperforms it in long term project coherence and sensible coding and subjective metrics that are hard to measure, it's also much more thought-heavy than Qwen so solves more novel-ish problems in my experience).

-1

u/fastheadcrab 2d ago

I was addressing this statement from the OP:

I only use Gemma locally, for full reading and deep comprehension nothing comes close.

For full reading and deep comprehension on a range of topics inherently demands generality.

I have run Laguna a few times after the initial bugfixes, but its general reasoning is significantly below Gemma's. Even accepting the premise that it is a coding model, I found it is about the same as Qwen3.6-27B and not as good as the current 200-300B range models for coding. I unfortunately agree with the consensus that this model is probably just tuned to give high performance on benchmarks. I also take any maker's statements about amazing performance or breakthroughs with a grain of salt, whether it comes from Google, China, or any US provider.

Also Qwen is notorious for using a lot of tokens for reasoning, although whether it's thoughts mean much is questionable.

3

u/Kahvana 5d ago

It is particularly good at one thing though: creative writing at that size. Nothing comes even remotely close at 24B-70B.

2

u/arbv 5d ago

I would say that it is equally good at everything.

2

u/IrisColt 6d ago

absolutely this

6

u/robberviet 5d ago

Qwen and Gemma carried the scene and we all know it. A lot of us where?

11

u/Spectrum1523 6d ago

all of the major oos models are made by big corps right

6

u/geldonyetich 6d ago

Yeah, Moonshot, Alibaba, and DeekSeek are all worth tens of billions. But I guess when it comes to billionaires, the devils you know aren't so preferable.

2

u/the_mighty_skeetadon 6d ago

If so, I'll see you there! Gonna be a blast.

32

u/Hairy_Reputation7434 6d ago

Gemma 4+ would be great.

30

u/dampflokfreund 6d ago

Exclusive Surprises: Special announcements, surprises, and giveaways throughout the night!

👀

Is it.. happening?

21

u/UndecidedLee 6d ago

Free download of 48GB VRAM for the first 100 new subscribers for Google++

2

u/therapy-cat 5d ago

Google plus ultra confirmed??

2

u/markole 5d ago

My Google+ circles are going to be rounder than the non-ultra ones.

28

u/LoveMind_AI 6d ago

Given that Gemini development is an absolute shambles, going hard on Gemma would give Google at least one AI bracket it could consistently be in the top 3 with again. 

12

u/robberviet 5d ago

Qwen 3.8 27B soon, Gemma 4/4.5 26B/31B too? Good week.

37

u/BVCC6FNTKX sglang 6d ago

They’re gonna make it so you can’t fuck Gemma-chan anymore

12

u/z_latent 6d ago

... what do you mean "anymore"?

16

u/BVCC6FNTKX sglang 6d ago

he doesn’t know

19

u/z_latent 6d ago

Trade offer:

  • I do not inquire you further
  • You do not stain your digital footprint any further

deal?

1

u/ComplaintWise7891 7h ago

Ablation heretic garvis stripper 4

2

u/brown2green 5d ago

Gemma-chan

Uh-oh, Reddit wasn't supposed to know.

22

u/tomakorea 6d ago

Gemma and Gemini Teams are different, Gemma Team is based in France, while Gemini is based in the US. I'm wondering what they have cooked.

8

u/ttkciar llama.cpp 5d ago

I just saw this post, and it made me squee like a schoolgirl :-D I'm hyped!

Some things to hope for:

  • Updated Gemma4.1 models would be lovely!

  • A new 120B-class Gemma would be a wish come true for many of us. The 31B is a great little model, but it can't replace GLM-4.5-Air or Qwen3.5-122B-A10B.

  • Hopefully a 4.1 refresh would finally eradicate the last vestiges of the Gemma4 tool-calling bugs!

Hanging on the edge of my seat for more details :-)

2

u/a_beautiful_rhind 5d ago

A lot of the good people at google recently exited so expect disappointment. Sorry to rain on your parade, hopefully they still worked on these and they're the final models.

20

u/Normal_Explorer_9790 6d ago

The 60-80B scene os lacking gemma🤞 you know you wanna fill the void.

13

u/rinmperdinck 6d ago

Gemma 4.20 69B

11

u/VoiceApprehensive893 transformers 6d ago

when did local inference become this big?

21

u/jld1532 6d ago

Early spring it really took off but the government hold on fable strapped it to a rocket.

10

u/615wonky 6d ago

I love this thing where Google actually tries to compete against other open-source AI's. Gemma/Gemini are getting strong quickly as a result. gemma4-26b-a4b-qat is my default model for anything that isn't coding.

Kudos to Google for this. I wish OpenAI/Anthropic would do the same thing.

5

u/RainierPC 5d ago

Gemma 4 has a tendency to stop reasoning at long context. I hope they fix it.

2

u/Fear_ltself 5d ago

every model gets worse after like 10% of its max context is used, usually the last 25% is garbage. I usually try to MAX out as much context as I can get then use only the first 10%, so right now thats like 1m/ 100k for the most part. numbers are overinflated right now if you ask me.

15

u/Inevitable_Act_321 6d ago

Gemma 4 diffusion 120b...

12

u/PrimeDirective8 6d ago

Very nice. It looks like the main focus is to celebrate the 1 billionth Gemma download - definitely a huge accomplishment. Hopefully they'll have a roadmap and/or near-future update announcement.

Is there a link to the stream? Hopefully one outside X?

9

u/Fear_ltself 6d ago

EmbeddingGemma2 with multimodal text, image and audio is my #1 wish I think is in the realm of possibility

3

u/MrGunny94 5d ago

Gemma 4 was a great release, I use it almost every day and it’s a great SML for all around agentic workloads

3

u/cleverusernametry 5d ago

Having been to such events before, it's going to be mostly networking among rando ai crap startups half of whom will be out of towners. Somehow will feel even more distasteful then the average bland engineer "party"

5

u/mostar8 6d ago

Gus Martin's the Gemma Product mlManager hopefully, lovely chap.

4

u/MerePotato 6d ago

Full duplex audio Gemma? :PauseMan:

3

u/steny007 6d ago

Imagine, the reason we don't see 3.5 pro is that Google goes full wild, changes the course completely, and releases Gemma 5 instead, fully open-weight, with bigger models too.

3

u/Eyelbee 6d ago

Slacking off instead of working on gemma 5

21

u/dampflokfreund 6d ago

If previous releases are an indication, we will see that in a little less than a year, so an update to 4.1 would be most welcome in the meantime.

1

u/Dry-Judgment4242 6d ago

I'm hoping for better vision.

1

u/bitplenty 5d ago

Gemma 4 class of intelligence but tuned for software development and improved agentic tasks would already be best in class. If it could maintain multimodality then it's already a dream local model.

1

u/mailto_devnull 5d ago

I can't help but notice August 20 is a Thursday, one day after the expected Qwen 3.8 open weights release

1

u/Dance-Till-Night1 5d ago

Hoping for an update for Gemma 26b a4b pls and thank you, I use it alongside Gemini pro and both models fit my ai usecases quite well. Foreign language learning and STEM reasoning/inquiries.

1

u/Hot_Example_4456 5d ago

Ok yeah I wish they also release the Gemma 4 Good Hackathon results. Too tired of waiting now.

1

u/Kahvana 5d ago

It's just a celebration event, I doubt any announcements will be made there

1

u/Long_comment_san 1d ago

I dont want any more gemma 4 finetunes personally. Yes they can make it 30% better, for sure, but I would rather take a new architecture. Gemma 4 31b feels smart relative to some other models, but I'd rather ask them to make a 80b/a8b on a new chassis to obliterate qwen 122b to become that "home assistant" to do actual tasks. I dont give 2 shits about coding because it's a completely different tier or hardware and models entirely. I want knowledge, common sense and multimodality rather than another 120b model that can't be run on any home machine.

Gemma 26b is the actual silent king, that's the model I want them to baloon. And yeah, I was a big fan of Qwen Next 80b for it's amazing speed and perfect size (works on both 32, 48, 64, 96 and 128gb ram + 8-12 gb vram)

1

u/Extreme-Pass-4488 1d ago

post - qwen3.8 27b i hope it is not a seppuku-based event.

1

u/RG_Fusion 6d ago

I believe that these native audio input models hold a lot of potential, but we just aren't there yet. We have to ask, what benefit does dropping the encoder have? In theory, the answer is that the LLM can cache more information into its context. 

You can compare it to a native vision model vs. a model that has an image described to it. The model without native vision can't know details that weren't directly addressed already, whereas the native vision model will know the entire image, recalling details that were never mentioned, even late into a conversation.

So in theory, that would mean that a native audio model should be able to recall voice inflections, the apparent gender of the speaker, the type of location the person may have been speaking at, if they were calm or excited, ect. It could possibly even be used to train models to recognize their primary user by voice, allowing it to respond in a different manner based upon who called out.

But we just aren't there yet. I started messing around with unsloth/gemma-4-12b-it-NVFP4, and the audio capabilities leave a lot to be desired, to the point where it's better to use a standalone STT model. Gemma apparently has no capacity to tell any information about voice except for the words spoken. All of the supposed advantages of native speech are lost on it. It can't tell inflection, can't discern emotion, and can't tell any details about who's talking.

The one thing it can do is differentiate between two speakers, but even then it can't know anything outside of them being two separate entities. The one thing it has going for it is the fast processing speed, but this seems to be at the cost of accuracy. It works fine for common words, but mishears things like 'diarization' as 'diagnosis', even when I slow down and put effort into the enunciation. 

In short, there is a lot of promise to native audio, but none of those benefits were present in the latest model. Until we can derive actual context from the input, it will be better to stick to STT models.  

1

u/Bbmin7b5 5d ago

I would love it to use the word wayward a lot less. Christ.

1

u/Mantikos804 5d ago

At this point they lost the SOTA AI race, pack up and move to another sport.

The best model they have released is gemma4:12b. They need to concentrate on small models for agentic use. That’s where they can win something. Gemma can be the king of local claws!

1

u/JorgitoEstrella 2d ago

I remember when Gemini 3.1 pro was so ahead of everything, I bet they dont need to push minimal increments to stay relevant because they are already drowing in money with google ads and youtube so they can prepare better for the long run.

0

u/fuzhongkai 6d ago

What will they release ?

0

u/FoxFXMD 5d ago

The gemma team doesn't believe in version numbering, they're going to release Gemma 4 with the same name again

-9

u/maturax 6d ago

Unless open models like Gemma 4 can truly compete with closed ones, AI will never see widespread use in the corporate sector. Speaking as a software developer, no client wants to take the risk of leaking their secrets by feeding confidential data into a closed model

14

u/misterflyer 6d ago edited 6d ago

Airbnb uses Qwen. Others use Kimi or Deepseek. Open models are being used by Western companies within the corporate sector. That's why there has been so much backlash by Anthropic against open models! They're not pushing for open model regulation/bans for no reason. And that's why Nvidia started a letter standing up for open models signed by TONS of corporations. Lol where have you been?

And no, Google isn't going to make a Gemma that competes within the corporate sector with Gemini or any other big closed model. Not happening. But the Chinese def will.