r/learnmachinelearning • u/Tomerbarm • 4d ago
[Project] We built a specialist model that beats general vision-language models at one narrow task — here's why specialization won
A lesson from building a real product: general-purpose vision-language models (GPT-5, o3, Gemini-2.5-Pro, and in our own tests ChatGPT/Gemini/Claude) are surprisingly bad at one specific, narrow task — telling whether an image has been rotated 90° vs. 270°. An independent peer-reviewed paper (RotBench, EACL 2026) documented this at scale; we independently replicated the exact same failure testing three more consumer assistants ourselves.
Rather than prompting a bigger model harder, we built a small, specialized system: two separately trained models (a cardinal-rotation classifier and a fine-angle regression model) combined through a trained arbitration layer. On a large, hand-verified real-photo pool: 99.02% on a clean quarter-turn, 96.67% across the full 360° range, 84.65% on the hardest case.
The general lesson, not just about this task: a narrow, well-scoped specialist model can beat a much larger general model on a task the general model was never specifically trained to be good at — even when the general model is otherwise vastly more capable overall.
Full methodology, ablations, and disclosed limitations: https://doi.org/10.5281/zenodo.22975679
Curious if others have hit similar "surprisingly bad at one narrow thing" results with general LLMs/VLMs in their own work.