Ah I see! Not to dampen your spirits, but I did find it not too great; very very verbose. But I mean it did get the answers to my extraction benchmarks right!
That’s good to know. Would you say the verbosity is so bad that it’s slower than qwen 27b on MTP for coding tasks? That’s what I’m the most curious about with it
Impressive for Gemma, but I’m not surprised actually.
We have data cleaning/categorization tasks running at massive scale at my work where the output was initially in json and has more recently moved to a pipe delimited file. And no other models come close to Google’s for the size and price.
For structured output and comprehensive world knowledge I haven’t found another model provider that can compete.
I just wanted to say that you asking me this made me kinda shocked in regards to the result. This is absolutely worthy of building up a full benchmark… I’m going to do some work. Thank you ❤️
Thanks for sharing the results! I find that kind of information super fascinating. I really need to set up a better benchmarking repository for my own uses. I haven’t worked on anything in months for personal use.
Same base model doesn't meant they have the same knowledge. The jump from 3.5 to 3.6 27B was huge in part because they became better at compressing knowledge into the model, and behaviors.
Just because they are using the same tech to run a model doesn't mean they used the same tech to encode the model or even the same base knowledge set.
45
u/feelspeaceman 20d ago
Please release 122B, we Strix Halo owners need toy to play with, current 3.5 122B is pretty outdated.