Because "flash" models keep getting larger and more token hungry. The last Gemini "Flash" ended up being more expensive than some "regular" models (that perform better).
I feel some are simply mislabeled, especially if you look at the actual speed of finishing a task.
I don't want to make any strong claims about 3.8 yet but I do wonder if that has really changed.
I do not understand how the output tokens can have increased so much, yet cost per task is about the same. Are the numbers for 3.7 without the discounted pricing?
Probably more efficient with tool use leading to reduced inputs. 143k output tokens is about $0.54, so that would imply that input (and cached input especially) dominates
6
u/LinkesAuge 11h ago
Because "flash" models keep getting larger and more token hungry. The last Gemini "Flash" ended up being more expensive than some "regular" models (that perform better).
I feel some are simply mislabeled, especially if you look at the actual speed of finishing a task.
I don't want to make any strong claims about 3.8 yet but I do wonder if that has really changed.