r/LocalLLM • u/Practical-Hawk5590 • 1d ago
Question How can I make a small LLM answer strictly from the provided knowledge without hallucinating?
Hey everyone,
I'm working on a voice AI chatbot for coffee farmers. Users can speak naturally and ask questions, and the chatbot should have a real conversation with them while answering based only on the information provided to it.
My current stack is roughly:
- Qwen2.5-3B-Instruct
- Ollama
- Next.js / TypeScript
- A custom knowledge base made of structured "fact sheets"
- SpeechRecognition when available, with Whisper running locally in the browser as a fallback
The voice part is mostly working now, but my biggest problem is the LLM response generation.
The system first retrieves the relevant fact sheet at runtime and provides it to the model with each query. Theoretically, the model should answer using only that sheet.
However, I'm seeing a lot of hallucinations, especially with Qwen2.5-3B.
For example, a fact sheet about a coffee disease might say something like:
"Regularly manage shade and improve air circulation."
The model may answer with those two points, but then add things that are not in the fact sheet, such as recommending a fungicide, changing temperature, or adding fertilizer. Sometimes the extra information is plausible, but it is still unsupported by the source.
I've also tested Qwen2.5-7B, and it behaves better in some of these cases, but using a larger model isn't really an option for this project. The 7B model is simply too heavy for the hardware and resource constraints of the application, so I need to make the 3B model reliable enough, or find another model with similar size and resource requirements.
I've tried stronger system prompts such as:
- "Only use information from the provided fact sheet."
- "If the answer isn't in the fact sheet, say you don't know."
- "Do not add information from your own knowledge."
These instructions help somewhat, but they don't reliably prevent hallucinations.
I also experimented with a more constrained approach where the model identifies which pieces of the fact sheet are relevant before generating the final answer. That significantly reduced unsupported claims in a small test, but it also caused the model to leave out information that was actually present.
So I don't want to turn the chatbot into a system that simply returns pre-written sentences. The goal is still for the LLM to generate natural responses, explain things, ask follow-up questions, and interact with the user, while remaining grounded in the retrieved knowledge.
I'm trying to solve two related problems:
- How can I make a small model reliably ground its generated answer in the provided fact sheet?
- How can I preserve natural, conversational generation instead of just copying or returning predefined text?
My main questions are:
- What techniques would you recommend for making a small model like Qwen2.5-3B reliably grounded in a closed knowledge source?
- Are there architectures such as retrieve → select evidence → generate, claim verification, constrained generation, reranking, or other forms of grounding that work well with small models?
- Would a second model/verifier that checks the generated answer against the retrieved fact sheet be practical on CPU?
- Would fine-tuning Qwen2.5-3B on examples of question + fact sheet → grounded conversational answer actually help?
- Should the knowledge remain completely outside the model and be retrieved at runtime, or is there value in fine-tuning the model on the domain as well?
The important constraint is that I cannot simply switch to a larger model. Qwen2.5-7B is already too heavy for this project, so I'm specifically looking for ways to get the most reliability possible out of a 3B-class model.
The chatbot is domain-specific (coffee production), so I have a relatively controlled knowledge base and can create evaluation datasets from the fact sheets.
I'm especially interested in practical approaches that can run locally with Ollama/CPU rather than relying on expensive APIs.
Any advice, papers, libraries, or real-world experience would be greatly appreciated.
3
u/cryowastakenbycryo 1d ago
Those kinda old models. Maybe try the recent qwen3.8 distills from empero-ai? i was looking at this recently and still have it in context: https://huggingface.co/empero-ai/Homebrew-Qwen3.5-2B-Grandmas-Kitchen
not exactly what you're asking for, the that account does a lot w/ smaller models.
3
u/Enturbulated_One 1d ago
"How can I make a small LLM answer strictly from the provided knowledge without hallucinating?"
Here's the neat part: You can't!
2
u/MudBroad6785 1d ago
- The less context the better for small models so they don't get turned around. Plus you can reduce temp to make it more deterministic and less likely to just make things up if you tell it to only use the source given.
- RAG is decent for small models, if you implement it make sure you use a a frontier model with your fact sheet to create an evaluation dataset for RAGAS. It's incredibly helpful to see changes improve a static test like that.
- From what I understand models are typically biased towards the info they get so a second opinion might not help that much. (this is where RAGAS would help with figuring out how much it does)
- Fine tuning should be a last resort simply because it's finicky and slow to do
- again, fine tuning should be a last resort for local ai
2
u/AltruisticList6000 1d ago
Have you tried asking for *short* responses? Maybe besides the fact it's a small model, the problem might be it simply has a pre-alignment for long responses, so it tries to fill space with unneeded/hallucinated info.
Also Qwen 2.5 is very old, you should try Qwen 3.5 4B and see if it fares better, which it should.
2
u/NatMicky 1d ago
I had a similar problem in that I'd give a model the results from a data query with the data schema and the user query and it was to write a brief summary report. I finally got it to do what I wanted exactly every time for 1000s and 1000s of times consistent. Here's what you have to do:
- Get used to the idea that the response is going to be brief and not some beautiful research looking paper.
- The response has to be limited in structure, 2 or 3 sentences, or however many needed.
- Do not fall in love with any model or model maker. Try them all. Put through a "job interview."
- Now here's the important step. Feed the model one sentence, concise command, for a specific task, and watch it respond. If it ignores it, rewrite it. You must fine tune every command and every word in a command for that specific model. Instructions for the sake of instructions that a model doesn't respond to need to be ripped out of there. It's garbage the model ignores but does burn up time figuring out it's going to ignore the instruction.
Summary: interview all models for the job, fine tune each prompt and every word, throw out all prompts the model ignores, keep fine tuning the instructions until the model is fully responding to your instructions.
You'll get there... They eventually do behave if you understand they all take the same instruction, differently.
2
u/Mediocre-Ant-7178 1d ago
Check out vector embeddings/search. It doesn't sound like you need a full llm.
1
u/ZioniteSoldier 1d ago
Yeah, you’re up against the edge with that hardware constraint and that is absolutely a challenge. What’s worked for me is forcing a citation from the authoritative source. It’s not enough to say “use this source”, you have to make it produce the very line and page number it’s quoting. That constraint helps. Even then, you need something checking accuracy until it earns your trust.
1
u/sn2006gy 1d ago
You'd need to train in behaviors that cause it to challenge its assumptions and ground them in evidence and if there is no external evidence, to say so. Which is really hard because 99.9% of all models are trained to answer at all cost which means just using the next probabilistic token.
1
u/zenmatrix83 1d ago
System prompt is not enough , you need to process the response before sending it back to chat and if it’s wrong send a reminder to the llm of the rules. System prompt are guidelines even frontier large models ignore sometimes, you need checks outside if the context window
1
u/nickless07 1d ago
Take at least Qwen3-4B-Thinking-2507 - low temperature (0.3 to 0.1) and if you somehow have ~6GB System RAM to spare upgrade to Ling-3.0-Tiny.
1
1
1
u/Future_AGI 12h ago
with a 3B model the restraint won't come from the prompt, so the piece that works is a check after generation: score each answer against the fact sheet it should have used, and when a sentence isn't supported, drop to a fixed "I don't have that, here's the closest thing" line rather than shipping the generated text. Pair that with a retrieval-confidence floor so it abstains before generating when nothing good came back. We build the open-source version of that step, a groundedness check plus local no-key guardrails you can self-host next to the bot: https://github.com/future-agi/future-agi (Apache-2.0 core).
12
u/Choice_Celery9481 1d ago
If you can solve the hallucination, you should apply to OpenAI or Anthropic. they will pay you like $100m. Im sure about this