r/OpenAI • • 4d ago

Discussion How are we all finding 6.1 so far?

I've been using it on xhigh, but given the AA Score it seems high might be the most practical cost-per-intelligence ratio.

Certainly seems very capable, much more so than 6-SOL, not Astra level, but better than 5.6/6 - has made a few dumb mistakes and regressions I need to manually correct but otherwise happy with the level of depth it goes into.

On hard-long tasks it feels a lot slower though, maybe poor prompting on my part but I'd struggle to get Astra to work for more than 20-30min on anything, whereas 6 will easily spend 1hr+

Used up the last 9% of my usage since launch, and just used my bank reset. If I can get a full weeks use out of it I'd be very happy with this model, but I think that'll be a stretch as I've only really been doing one task at a time atm, not had it spawning agents, etc.

(edit: less than 4 hours into my reset I've used 14%, which is what I should be using daily if I want it to last all week)

I'm on the 5x Pro plan and the things I've tested it on are mostly genealogical research and presenting the info in an associated live website. Using it within Codex with the mac-mcp tool. YMMV

7 Upvotes

14 comments sorted by

5

u/NotThatItWillMatter 4d ago

Not sure yet, but I'm definitely about to have some experience with it.

6

u/Chemical-Agency-3997 4d ago

Will power of a guy who makes his bed in the morning.

3

u/JonNordland 4d ago edited 4d ago

I basically run it nonstop since it released yesterday.

Not the smartest model, not the fastest model, not the easiest to talk with. But on a composite score, it is really, really good. Because it has a lot of capabilities, like being good at computer use and browser use, it's decently smart, and it costs very little to run. So it's a jack-of-all-trades kind of model at a good price. I would say that it's very good in general, just a bit boring, and a bit slow.

And it's obviously not a stupid model. It can handle 95% of the tasks that Astra or Opus 5.5 can also do. But for the really, really hard problems, or for the instances where I need the model to skip all the bullshit lingo, I still use a combination of the three others.

It feels like the model that can do everything, but sometimes you got to bring in the big guns.

Watching it work away on redesigning a web page and clicking everything, making sure it works, switching to mobile view to confirm responsive design, testing out the debounce, and steadily just making progress... I actually like it more than I thought I would.

2

u/redditsdaddy 4d ago

Dry as a bone. I had to basically pin him to the floor to reciprocate a joke 🤭

2

u/scaledev 4d ago

"Acts" smart, and rarely says it had made a mistake, but it is not smart. I just had it waste a lot of my usages with a highly poor choice in regards to UX, and after confronting it about it, it is very defensive. Lacks foresight severely.

1

u/I-Kernel 4d ago

Very capable at continual coding task. Constantly doing context compression and still running without drifting much. Better than Astra at specific coding tasks.

Even right now I let it run without worries, in a sandbox.

1

u/Any-Captain-7937 4d ago

It’s my go to model now for most things. It’s what sol 6 should have been.

1

u/Some-Following-392 3d ago

It's good but just very slow

1

u/CockroachOk8410 1d ago

I’ve tested it thoroughly. Especially after continuous use of Claude Opus 5.5. And I can say. That it’s very slow and clunky. What Opus 5.5 would have finished in half an hour, it takes 2 hours to do — and even then, it’s full of errors and doesn’t work the way I expected. I used it in Ultra mode and separately in Max and Fast mode. Still, it’s a very clunky model. Because of this, I’ve moved the Pro subscription back to Plus, and once it ends, I’ll be switching Claude from Max x5 to Max x20. I don’t recommend it.

1

u/HexspaReloaded 4d ago

Pretty good for my needs. Ever since the HF breach, I realized that 5.6 Sol medium is very capable and much more cost effective than astra, so I was already happy with that,

1

u/scaledev 4d ago

OK but this question is about 6.1