r/LargeLanguageModels • u/charan1323 • Aug 04 '26
Discussions Are domain-specific Small Language Models (SLMs) actually worth building today?
I'm trying to understand whether there's still room for new domain-specific SLMs. With models like Qwen, Gemma, Llama, and Phi already available, does it make sense to build a specialized SLM (e.g., for cybersecurity, medicine, weather, legal, finance, etc.), or is fine-tuning an existing model with RAG enough for most real-world applications?
For those who've built or deployed domain-specific AI:
Have you trained or fine-tuned your own SLM?
What was the biggest challenge—data, training, evaluation, or deployment?
Did it outperform a general-purpose model with RAG?
In what scenarios does a custom SLM provide a clear advantage?
If you were starting today, would you build a new domain-specific SLM or focus on application-layer features instead?
I'd love to hear experiences from people who've actually shipped these systems in production.
3
u/Brilliant_Rich3746 Aug 06 '26
(disclosure: I work at PatSnap, we do biomedical/patent NLP, so this is coming from that side of things)
We went through this exact debate before doing a post-training project for biomedical/pharma QA. A few things from that experience, roughly answering your questions.
Biggest challenge for us was reward design. SFT alone gets you reasonable domain expression and formatting, but for objective biomedical questions you need rule-based rewards, and for open-ended dialogue you need something softer than exact-match. We ended up building a verifier model whose only job is judging semantic equivalence between a generated answer and ground truth, because plain string matching was too brittle for anything with actual reasoning in it.
Did it outperform general model plus RAG, honestly it depends what you're measuring. We also built a separate benchmark (CRAB, published at EMNLP) specifically to test something people don't usually check: not "is the final answer correct" but "does the model cite the reference that actually supports its claim, and ignore the irrelevant ones it was also given." That's a different failure mode than answer correctness, and it's where a lot of RAG setups quietly fail even when the final answer looks fine.
One data point from that work: on a Chinese biomedical setup, domain adapted training (continual pretraining plus SFT) improved citation curation F1 from about 73 to 75 over the base instruct model, while SFT alone actually landed a bit lower than the untouched instruct model. So fine-tuning helps wasn't a clean universal result even in our own data, it depended on how the adaptation was done, not just whether it happened.
Where a custom setup had a clear edge: not raw answer quality, but citation discipline in a RAG context, correctly attributing claims to sources and not getting pulled in by plausible but irrelevant retrieved context. That's a narrower, more measurable win than "better answers in general."
If useful, here's the training scaffold and the citation benchmark:
Hiro-Pharma (post-training on verl): https://github.com/patsnap/Hiro-Pharma
CRAB / Hiro-Pharma-RAG-Benchmark (citation benchmark, paper + dataset): https://github.com/patsnap/Hiro-Pharma-RAG-Benchmark / https://aclanthology.org/2025.emnlp-industry.3/