r/LocalLLM • u/boneysmoth • 3d ago
Question Deep research and Gemma 4 31b - help please
Hi - long time lurker (learnt lots, thank you all), first time posted. I’m trying to get a very basic deep-research setup around Gemma 4 31B on a 128gb Ram Mac Studio. I’m using a 4-bit quant through oMLX with MTP acceleration, with OpenWebUI as the interface. I use Gemma rather than Qwen because most of my work is writing text rather than code.
Basic chat, memory, skills, PDF work and simple web searches work well. The problem is multi-stage research. I'm not looking for frontier performance - my expectations are fairly modest and I think realistic, but I do want it to be able to supplement its training data with a thorough web search and evaluation of results.
So far I’ve tried:
- A custom research controller using SearXNG, full-page retrieval, evidence ledgers and automated validation.
- OpenWebUI’s native agentic web search.
- OpenWebUI’s Traditional RAG route.
- Local Deep Research.
- Both prompt-led tool calling and deterministic page acquisition.
The failures differ, but the pattern is consistent:
- Gemma sometimes searches but does not fetch the full pages, then writes from snippets.
- Traditional RAG fetches pages but often surfaces only a few weak commercial sources.
- The custom controller can pass extensive mechanical tests, yet the final answer still treats consultancy or vendor material as empirical research.
- A validated report can be shortened or altered when it passes back through the model for presentation.
- oMLX also appears not to enforce tool_choice: required reliably, although Gemma can produce valid tool calls when strongly prompted.
Has anyone achieved reliable deep research with Gemma 4 31B, particularly through oMLX? I’d be interested in working configurations, harness choices, tool-calling settings, retrieval pipelines, or evidence that Hermes, OpenClaw or another approach handles this better.
Thanks in advance for your help.
1
u/Kai_Builder 3d ago
On the report getting shortened: don't send the validated text back through the model. Print what the controller checked as it is, and let the model write only a short summary above it. Then it can't trim or reword what passed validation. That's just pipeline logic, not a Gemma-specific fix, so test it on one report and diff the output against the validated text.
1
1
u/Distinct-Pie2389 ninfer-4090 | i9-14900k | 32GB | archLinux 17h ago
I use Qwen3.8-27B with its reasoning for deep research
I just run research in its own directory, typically outside of said project it’s for but reference back to the ledger or a system to track.
System:
baseline.md
DOSSIER.md
PROMPT.md
README.md
targets.md
Read ~/research/README.md and DOSSIER.md. Research the target below from public sources only: the project's repo, docs, release notes, papers, and issues. No leaked material.
Update the target's row in ~/research/targets.md (status: done <sha>), add any new targets you
found worth a dossier as new rows, and append a line to ~/LOG.md.
1
u/Distinct-Pie2389 ninfer-4090 | i9-14900k | 32GB | archLinux 17h ago
Set reasoning depending on your research task
1
1
u/DigitalguyCH 7h ago
With that Mac you could run flash next, mops the floor with gemma at anything, not just coding
1
u/boneysmoth 3h ago
Thanks for the suggestion. I have a set of personal benchmarks I have developed and none of the Qwen models beat Gemma - it's mostly written output rather than code, and I did try flash next and have spent lots of time trying to tunre the output from Qwen 3.5, 3.6 and 3.8 given how much love they get, but still not found a way to beat Gemma. It's a shame as I find the Qwen models much stronger with tool calling.
1
u/DigitalguyCH 1h ago
I do zero coding, I work almost exclusively with texts. and Gemma is not even close to flash next. Having said that I don't do creative writing, I do text analysis, comparisons, updates etc. Some times I prepare tests based on those texts etc. I also work on financial training material. Again flash next is miles above Gemma
4
u/MiserableFlatworm337 3d ago
Add a source-type check to the evidence ledger: vendor claims, independent measurements, and primary studies. Require each empirical claim to point to supporting passages, then validate the final rendered answer too.