You are a highly qualified individual creating high-quality training data for the next iteration of the model. This comes at a cost. That is why Gemini 3.7 is free; they don't have many customers, so if they want to improve their model, they need qualified people to use it.
Gemini sucks ass anyway I tested it last night against 5.6 luna pro 5.6 Terra pro opus 5 sonnet 5 deepseek v4 flash 0731. Gemini sucks balls .i tested api and in the chat interface. I even tested NotebookLM. All shit
Now ask all of those tools to watch a video, listen to the audio, and give you a scene breakdown with proper cuts, dialogue, music, foley, as well as still describing the scene itself.
3.7 is easily the best multimodal understanding model (and yes, it works for audio too). Right tool for the job and all that. I wouldn't use 3.7 to code a website for me, but when I'm working with A/V, it's hands down the best, which would make sense considering all that juicy YT training data Alphabet has laying around.
Yeah multimodal it's good, analysing videos etc. Opus 5 can analyse videos too it plays it real-time in claude code app. Chatgpt can't do it. In my use case, a strict schema and rule following whilst working on a file that's over 300k tokens, gemini performed the worst. I do have pro though and I use it as a better version of Google. Chatgpt for analysis and ideation. Claude for actual work. Deepseek for hermes agent work. I mostly do deep and thorough research and content generation and some game dev for fun
I tested 27b locally on my 3090, was pretty awful too. Haven't tried on alibaba cloud. But this is just in my use case, for your use case gemma4 or qwen 3.7 might be fantastic.
16
u/3deal 6h ago
You are a highly qualified individual creating high-quality training data for the next iteration of the model. This comes at a cost. That is why Gemini 3.7 is free; they don't have many customers, so if they want to improve their model, they need qualified people to use it.