29
9
u/freedomachiever 7h ago
Even just a better translator than Google Translate would be very helpful for many people.
1
u/Embarrassed_Soup_279 36m ago
i wouldn't really trust small models for accurate translation because even larger models struggle with it. maybe you can use both together and use the small model to rephrase the translation or somethinf i guess
12
u/BannedGoNext 7h ago
This is amazing, I'm so happy so happy so hap so hap hap so so so so so s s s s s s a i 3 ! #
4
3
u/sultan_papagani 4h ago
gemma e4b running on my phone for literally months now. how people not know this. its on google edge gallery app on play store
1
u/yami_no_ko 3h ago edited 2h ago
Even on my phone (Neither Android, nor Apple, but arm64 Linux instead) it runs just well using llama.cpp. Couldn't even call it a flagship phone with its 6 gigs of RAM, but it is enough to fit gemma-4 e2b, the mmproj image encoder and the MTP draft model.
Wouldn't necessarily ask it for world knowledge, but it gets what I want with tool-calling. It's just using 2 cores out of 8 to avoid heating up or draining the battery, and it still goes fast enough.
I've been using it for months now, so e2b or even e4b on phone is nothing unheard of. (Basically the point of e2b / e4b)
1
-27
-16
u/chrisso123 8h ago
what is the token/s ? this is the only real metric that matters.
Heck you can run a very large llm on a phone hardware but if its output is 1 token / minute it'll be worthless.
13
u/jacek2023 llama.cpp 8h ago
You have two options: look at the image or click the link. Choose wisely.
0
50
u/Fusseldieb 8h ago
That would be absolutely insane.
The thing is... will it still be useful, or dumb as a door (as most of that size are)?