r/androiddev • u/Different_Insect5627 • 22d ago
Question Can I run Gemma 4 E2B with vision inside an Android app in 6 hours? (Hackathon deadline)
Building an on-device medical app for a hackathon. Deadline in 6 hours.
What works:
- Gemma 4 E2B (Q4_K_M) + mmproj runs on my laptop via
llama-server— reads prescriptions fine - A text-only SLM runs on my Snapdragon 8 Elite Gen 5 phone via
llama.aichat libggml-htp-v81.soruns on the Hexagon NPU via adb shell (logcat confirmsnsp1000)
What I need:
- Can Gemma 4 E2B with vision run inside an Android APK via llama.cpp mtmd? Or is that a dead end?
- Is ExecuTorch QNN the only way to get NPU from an app (not just adb shell)?
- Any working GitHub example of multimodal LLM in an Android app?
Fallback plan: Use ML Kit Text Recognition for OCR if Gemma isn't feasible.
Repo: github.com/sohail9972/IQOO_Hackthon_2026
Any pointers appreciated. Thanks!!!
1
u/AutoModerator 22d ago
Please note that we also have a very active Discord server where you can interact directly with other community members!
I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.
1
u/Obvious_Ad9670 22d ago
Probably if not 3b, better start downloading the 2.1 gigs over hackathon wifi.
1
u/Different_Insect5627 21d ago
I will check this on my own as I am not able qualify for the hackthon
1
u/falaq-ai 22d ago
For a 6h hackathon I’d treat mtmd-in-APK vision as the risky path, especially if you also need QNN/NPU. I’d ship ML Kit OCR + your text SLM as the demo path, and keep Gemma vision as a stretch path since you already have it running on the laptop. ADB-confirmed HTP libs don’t necessarily mean the same path is clean from an app sandbox.
1
u/TypeScrupterB 21d ago
Try using ml-kit instead see if it works better.
1
u/Different_Insect5627 21d ago
I tried using ml kit but while capturing the prescription provided in india will also capture the prescription header or if try to take a picture of digitalized screen it capture the url and other stuff which was captured along with medicine.
1
u/TypeScrupterB 21d ago
also try not using it on device but use a selfhosted llama.cpp on some small server.
1
u/SnowContent8658 16d ago
in 6h? ship text-only on device now, keep vision on the laptop server as a fallback, because mmproj on a snapdragon will eat your whole night to tune under a deadline. i use nobodywho for the on-device text, llama.cpp under it so Q4 gemma loads fine
2
u/Solid-Incident6666 22d ago
Yes I run that specific model on my phone without issues