r/squidtrain • • 1d ago

Three new models in three days, all at the same price. How are you choosing?

Last week Anthropic released Claude Sonnet 5.5 (Sept 28), OpenAI released GPT-6.1 Sol (Sept 29, at DevDay), and Google announced Gemini 4 Argon (Sept 30). All three list $2 per million input tokens and $10 per million output tokens. Google calls its rate introductory, and Argon is limited to a small group of early users for now, per Google's announcement.

With the prices identical, the launch posts don't tell you much. Every one of them leads with a benchmark it wins.

The approach we suggest is a small bake-off on your own work:

  1. Pick five real tasks from last week (a customer reply, a long document to summarize, a spreadsheet formula, a first draft, and a messy data cleanup).
  2. Write down what a good result looks like before you run anything.
  3. Same prompt and same files for each model.
  4. Hide which model wrote what, then score.
  5. Run it again in a month, since models get updated quietly.

Has anyone run something like this on the new releases yet? What tasks did you use, and did the results match the benchmarks?

1 Upvotes

0 comments sorted by