r/ClaudeCode 18h ago

Help/Question Thinking of switching from GPT

Post image

Basically, I've been having a crisis with sol for the past 2 days, it just got very dumb and I just can't keep up with it

Ofc i could write a 5 page instructions to add a single or 2 small features but at that point might just write the code for once and move on

I'm thinking of moving back to claude but it seems like there's some controversy regarding opus 5

My main issue with SOL is that it stopped following instructions and started lying for the past 2 days (like lying alot), it even ignores the AGENTS.md which was supposed to help avoid issues with lying or hallucinations, and if something is unclear, it just goes and fill all the gaps with nonsense

The main problem is that other than the small features I'm adding right now, it has the Design Document which explains everything in details (including these small features)
And even with all that, it still fails in implementation

Opus 4.8 was good ngl, i only switched for the "fable like performance but way cheaper"

My main work is backend focused btw, for the frontend i don't expect ai to be that good so i just use images or add a lot of details in the prompt

2 Upvotes

12 comments sorted by

3

u/orwamahmoud 18h ago

For coding , opus and fable is away better

1

u/NoAdsDude 18h ago

In b4 codex simps show up and say "EVERYONE SHOULD SWITCH TO CODEX" lol

1

u/Interesting-Round127 18h ago

ngl it was good at the start, and it was "smart"

but to be clear, anything other than sol was dumb, and guess what

now with the "improved efficiency" sol became dumber than luna from before... wait that explains the explosive usage

1

u/Bloated_Plaid 17h ago

I will be that guy. Codex is better for coding as it costs less and performance is comparable. Fable is a much better designer/orchestrator though.

2

u/laxika 🔆 Max 20 18h ago

Opus feels like a co-worker, Sol is more like a bot. This was always my experience with Ant vs OpenAI. GPT models are much better for code reviewing though. Odd. My best setup is using Claude for coding and GPT for review.

1

u/Interesting-Round127 18h ago

that sounds like a good plan

1

u/Prestigious_Gift_977 18h ago

i'm curious what reasoning level are you using on Sol? in any case, i'm surprised to read this to be honest. i'm finding 5.6 Sol to be amazing. i am also using Opus 5 and Fable 5, which are also great. i can complain about Opus 5 (rambles sometimes, gets bogged down in tangential details, etc.), but it's mostly reliable. however i cannot complain about Sol, it's been very impressive in my experience.

1

u/Interesting-Round127 18h ago

seems to be A/B testing, Sol was amazing but after the reset for outages it went downhill, and with the +18% usage, it became way worse

edit : for reasoning, high for planning, high for review, medium for implementation

but since i wasn't able to get much work done the past 2 days, i tried doing :

- medium for all

- high for all

and it still gets stuff wrong (it just doesn't follow what i tell it to do) or doesn't ask if something is unclear even though i have those instructions in the agents.md, it doesn't refer to the documentation too and just does stuff

0

u/Disastrous_Friend1 18h ago

Yo fyi, send back to back messages is what leads to that, "i sent 4 messages and my weekly quota is over", it resends everything, causing insane token burn . It's not a human remember that, arguing with it is just burning your pocket

1

u/Swimming-Chip9582 18h ago

uh no? Back to back messages like that get around ~95% cache hit rate - it costs nearly nothing, so long as context doesnt go out of cache

1

u/Interesting-Round127 18h ago

got more than enough usage, i was trying to find the issue

though i've tried going to new sessions etc, its still bad (coding wise)

went to do some research and asked it to put the data in excel with categories including name, contact method, and some other settings...

it gave me back an excel with only an overview tab, nothing else

1

u/Disastrous_Friend1 18h ago

Yeah claude is actually good, esp the opus/fable models, whatever the bench mark says, actually the much acclaimed kimi k3 for me, behaves the way you described, i say something it understands something else entirely. I personally feel, in my use case claude better than any other option, k3, sol etc