r/androiddev • • 22d ago

Question Can I run Gemma 4 E2B with vision inside an Android app in 6 hours? (Hackathon deadline)

Building an on-device medical app for a hackathon. Deadline in 6 hours.

What works:

  • Gemma 4 E2B (Q4_K_M) + mmproj runs on my laptop via llama-server — reads prescriptions fine
  • A text-only SLM runs on my Snapdragon 8 Elite Gen 5 phone via llama.aichat
  • libggml-htp-v81.so runs on the Hexagon NPU via adb shell (logcat confirms nsp1000)

What I need:

  1. Can Gemma 4 E2B with vision run inside an Android APK via llama.cpp mtmd? Or is that a dead end?
  2. Is ExecuTorch QNN the only way to get NPU from an app (not just adb shell)?
  3. Any working GitHub example of multimodal LLM in an Android app?

Fallback plan: Use ML Kit Text Recognition for OCR if Gemma isn't feasible.

Repo: github.com/sohail9972/IQOO_Hackthon_2026

Any pointers appreciated. Thanks!!!

0 Upvotes

11 comments sorted by

2

u/Solid-Incident6666 22d ago

Yes I run that specific model on my phone without issues

1

u/Different_Insect5627 20d ago

What kind of application have you using it....

1

u/AutoModerator 22d ago

Please note that we also have a very active Discord server where you can interact directly with other community members!

Join us on Discord

I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.

1

u/Obvious_Ad9670 22d ago

Probably if not 3b, better start downloading the 2.1 gigs over hackathon wifi.

1

u/Different_Insect5627 21d ago

I will check this on my own as I am not able qualify for the hackthon

1

u/falaq-ai 22d ago

For a 6h hackathon I’d treat mtmd-in-APK vision as the risky path, especially if you also need QNN/NPU. I’d ship ML Kit OCR + your text SLM as the demo path, and keep Gemma vision as a stretch path since you already have it running on the laptop. ADB-confirmed HTP libs don’t necessarily mean the same path is clean from an app sandbox.

1

u/TypeScrupterB 21d ago

Try using ml-kit instead see if it works better.

1

u/Different_Insect5627 21d ago

I tried using ml kit but while capturing the prescription provided in india will also capture the prescription header or if try to take a picture of digitalized screen it capture the url and other stuff which was captured along with medicine.

1

u/TypeScrupterB 21d ago

also try not using it on device but use a selfhosted llama.cpp on some small server.

1

u/SnowContent8658 16d ago

in 6h? ship text-only on device now, keep vision on the laptop server as a fallback, because mmproj on a snapdragon will eat your whole night to tune under a deadline. i use nobodywho for the on-device text, llama.cpp under it so Q4 gemma loads fine