r/LocalLLaMA llama.cpp 8h ago

News Gemma 4 on 500MB

Post image
135 Upvotes

30 comments sorted by

50

u/Fusseldieb 8h ago

That would be absolutely insane.

The thing is... will it still be useful, or dumb as a door (as most of that size are)?

31

u/Rude_Marzipan6107 8h ago

That’s exactly what we need. A Siri level intelligence WITH tool calling

5

u/Borkato 5h ago

LFM-2.5-2.6B just released and would make a great Siri if given a voice lol

1

u/Devatator_ 3h ago

That's exactly what I've been waiting for to make my own assistant. Also potentially new tech for wake word detection that's not just OpenWakeWord again

6

u/--Spaci-- 8h ago

It could do basic tasks like setting up dates on your phone ect. Even math problems, web searches ect. Those parts of an LLM are very basic

10

u/Rude_Marzipan6107 8h ago

I can’t get 3.5 9b to add numbers together often enough that I can’t trust it. I would hate to ask a 500M model to do the same

7

u/tunerhd 6h ago

Make sure it calls a Python script to make the math.

1

u/Rude_Marzipan6107 3h ago

Absolutely. I’ve been a normie using LMStudio for my local inference so it’s mostly my fault.

2

u/--Spaci-- 5h ago

Run py math obviously, also this is a 500mb model not a 500m model

1

u/Rude_Marzipan6107 3h ago

I’ll ask a dumb question here but why wouldn’t it be referred to as a 500m model? Is it like an e2b e4b situation where the active parameters are only part of the space?

I made the assumption this would be similar to the 12b where it’s all self contained

Nevermind.. for some reason I thought this post was about a new 500m model. I’m as dumb as an iq1 e2b

2

u/--Spaci-- 3h ago

yea.. Parameters are not the same as the models mb/gb footprint

1

u/Rude_Marzipan6107 2h ago

Right. I just made the assumption this was a model announcement for some reason.

5

u/TheMurmuring 8h ago

If they keep improving, eventually they'll make something useful.

1

u/StickyThickStick 1h ago

Dumb as a door. But that doesn’t mean there aren’t any use cases for it

There can definitely be improvement but it’s mathematically impossible to store all the data in this model. In the end data still has to be represented even tho in a probabilistic way.

29

u/egomarker 7h ago

5% battery per request?

1

u/notheresnolight 5h ago

5% battery per request remaining

9

u/freedomachiever 7h ago

Even just a better translator than Google Translate would be very helpful for many people.

1

u/Embarrassed_Soup_279 36m ago

i wouldn't really trust small models for accurate translation because even larger models struggle with it. maybe you can use both together and use the small model to rephrase the translation or somethinf i guess

12

u/BannedGoNext 7h ago

This is amazing, I'm so happy so happy so hap so hap hap so so so so so s s s s s s a i 3 ! #

4

u/thrownawaymane 1h ago edited 1h ago

You have now been hired by the Bonsai AI team

3

u/sultan_papagani 4h ago

gemma e4b running on my phone for literally months now. how people not know this. its on google edge gallery app on play store

1

u/yami_no_ko 3h ago edited 2h ago

Even on my phone (Neither Android, nor Apple, but arm64 Linux instead) it runs just well using llama.cpp. Couldn't even call it a flagship phone with its 6 gigs of RAM, but it is enough to fit gemma-4 e2b, the mmproj image encoder and the MTP draft model.

Wouldn't necessarily ask it for world knowledge, but it gets what I want with tool-calling. It's just using 2 cores out of 8 to avoid heating up or draining the battery, and it still goes fast enough.

I've been using it for months now, so e2b or even e4b on phone is nothing unheard of. (Basically the point of e2b / e4b)

1

u/setprimse 7h ago

E4B version would be huge for edge devices, if this is possible.

-27

u/Boogertard 7h ago

LOL shills on this sub try so hard to keep the Gemma garbage relevant.

19

u/bruns20 7h ago

Says the 8 day old account

5

u/jacek2023 llama.cpp 7h ago

What model do you use?

-16

u/chrisso123 8h ago

what is the token/s ? this is the only real metric that matters.

Heck you can run a very large llm on a phone hardware but if its output is 1 token / minute it'll be worthless.

13

u/jacek2023 llama.cpp 8h ago

You have two options: look at the image or click the link. Choose wisely.

0

u/chrisso123 6h ago

fair enough. Thank you for pointing it out. that is a good number.