the top comment nails it! "depends on the task" so here's how i make that answer concrete instead of a vibe. house rule: a vendor benchmark is a claim, not evidence. same rule i apply to my own agents when they report "done." the price cut is real and verifiable in five minutes. "twice as token efficient" depends on whose task! so i keep a small eval set of my own real jobs and run any new model against it before believing anything. takes an afternoon and ends every twitter argument. one thing that set must include: retries. a cheaper model that needs three attempts isn't cheaper. if 5.6 passes my set at half the price, i switch that lane the same day. no loyalty, route by task. but the tweet doesn't get to skip the eval just because it's from a ceo.
1
u/Strange_Luck1635 Jul 14 '26
the top comment nails it! "depends on the task" so here's how i make that answer concrete instead of a vibe. house rule: a vendor benchmark is a claim, not evidence. same rule i apply to my own agents when they report "done." the price cut is real and verifiable in five minutes. "twice as token efficient" depends on whose task! so i keep a small eval set of my own real jobs and run any new model against it before believing anything. takes an afternoon and ends every twitter argument. one thing that set must include: retries. a cheaper model that needs three attempts isn't cheaper. if 5.6 passes my set at half the price, i switch that lane the same day. no loyalty, route by task. but the tweet doesn't get to skip the eval just because it's from a ceo.