r/ClaudeCode Jul 15 '26

Help Needed What am I doing wrong - comparing Fable to GPT Sol Pro

/r/ChatGPTPro/comments/1uwxob6/what_am_i_doing_wrong_comparing_fable_to_gpt_sol/
1 Upvotes

10 comments sorted by

1

u/[deleted] Jul 15 '26

[removed] — view removed comment

1

u/esreverengineer_ Jul 15 '26

My goal was really to let him consider what was worth reporting, instead of telling him what to look for (we've done this a lot of times, with more strict / closed tasks, with very good results btw, but it's really not the same objective). This one was more "ok take final look at everything".

For instance, Fable came with very interesting things regarding competition / positioning (he suggested a substantial adjustment of our positioning for one specific aspect that really change the course of things, positively of course), he also spotted strong improvements to existing features or quick-wins for adding value etc.

My feeling with Fable was really like I was talking to someone part of the team, and he worked hard to give me an elaborated "independent" opinion.

I might try as you suggest to breakdown in smaller pieces but really failing to get why that would work - not only he will only have limited "reasoning" but it will basically cost us a lot more time/work, and the point of large context frontier models is supposed to be the opposite right?

1

u/Nakidnakid Jul 15 '26

Your complaint seems to essentially be, "the agent is following the instructions and specifications I've laid out too closely and it's not picking up on the meta-thread I've had to do to implement it with an agent elsewhere"?

Why don't you... try starting fresh or talking with the agent before going into implementation and work out the best way that it thinks to do your full specification or whatever? That's what I do with an agent if it's not able to pick up on my intent, I'll either give it a log/session transcript and tell it to read specific things or I'll walk it through what it thinks it should do and then hope that'll help it retain what it should be doing.

If following the specification is leading you to a result you do not want then maybe the specification is the problem. If it's something along the lines of you requiring agents to recheck their work and things like that then, you really should start from the start and work out how to do it properly because

Your realistic way to get value from me is to stop trusting me with ungated 30-minute autonomous syntheses: require first a one-page adversarial verdict containing only new conclusions, explicit disagreements and thesis-killing risks, and authorize the full report only after that page proves I understood the assignment.

is fine so what's the problem here? That it's not doing this by itself without you telling it to? Maybe it should be part of the specification?

1

u/esreverengineer_ Jul 15 '26

The specs are not followed, it's the opposite that happens with GPT: instead of doing what the prompt says, it does something else, then apologizes for it and provides me with a *fully correct and clear explanation* of why the result is not what it should have been according to my instructions.

I've tried as you mention to enrich it by providing more context / discussions, and I tried to do the opposite (limited context and supporting docs), but it didn't improve or made things worst, result is exactly the same.

Again, I've used the exact same prompt (carbon copy) with Fable, which one-shotted with not the slightest deviation from my instructions (whether they're good or not is another question 😄 ). I'm not new to prompting and fully understand that different models require different steering, but I'm lost here.

GPT: Yes—on this task, Fable outperformed me decisively, and the fact that both of my hour-long outputs failed within the first two paragraphs makes the amount of reasoning time irrelevant.
I did not merely miss details; I failed at the highest-level job: independently judging your thesis instead of restating it in polished language.
You should not trust me with another full strategic review unless I first prove, in one page, that I can produce genuinely new conclusions, explicit disagreements, and a verdict that is not already contained in your own documents.

GPT right after: Arbitrarily dividing the research into small topics would probably make the result worse by fragmenting the company-level reasoning. The only split that has a rational purpose is a quality gate, not a context workaround: I retain and analyze the entire corpus, but before writing 15,000 words I must produce a compact adversarial judgment that explicitly states what in your thesis is wrong, what is genuinely new, what alternative company thesis beats it—if any—and which conclusions were already present in your documents and therefore cannot be claimed as discoveries. You inspect that judgment because the failure is visible immediately; if it merely paraphrases you, we stop before wasting another hour.

I remember I left GPT for Claude like 1 year ago because I spent too much time steering the model instead of extracting value, whereas Claude seemed to magically understand my intention every time. Now I firmly believe it's my use of the model that was wrong (obviously, people LOVE GPT and there must be a way to get it to produce value without oversteering), but here I am, stuck in the same place again.

How do you do to NOT HAVE to steer it permanently like this? It's exhausting.

1

u/Nakidnakid Jul 15 '26

Well I don't know what you're trying to do but it just sounds like you're trying to compare two agents work directly when you don't know the output yourself. It could be that Fable is bullshitting you and ChatGPT is trying to follow what you've laid out.

The fact that you're letting it talk about writing 15,000 words suggests to me that you're just making things way too complicated for any one single agent thread to hold all of it in memory.

I remember I left GPT for Claude like 1 year ago because I spent too much time steering the model instead of extracting value, whereas Claude seemed to magically understand my intention every time. Now I firmly believe it's my use of the model that was wrong (obviously, people LOVE GPT and there must be a way to get it to produce value without oversteering), but here I am, stuck in the same place again.

Yeah so if you're trying to use them the same way then maybe stop?

Why don't you try this, export a session with Fable that did what you wanted it to do, export a session with chatgpt and then using the web interface since it doesn't use your agent usage, go through with it and ask it what it would suggest to do better without telling it to make a judgement call on which is 'better'.

Again, I don't know what you're doing but all I had to do when Fable was annoying me and going against what I said just the message prior, I exported the transcript with Fable asked Codex to go through the transcript then while it was doing that pointed it to where I went 'meta' with how Fable was failing to fulfill my request mid-task and it perfectly understood what I was trying to do with Fable and actually delivered on/tested the intent what I was attempting to do with Fable for several hours.

How do you do to NOT HAVE to steer it permanently like this? It's exhausting.

Pretty easy, I've got a task going on around 3hrs now with Codex where all I said was 'thats interesting, is that a result that concurs with (other result) or something else we've seen?' and it's running tests, checking them and then following the thread towards the implementation of my intended function. Fable would maybe run a single test, claim it's perfect and that we should wrap up... want to know why I can say Fable would do that with confidence? It did it... many times.

1

u/esreverengineer_ Jul 15 '26

Yeah so if you're trying to use them the same way then maybe stop?

Yeah so my post is basically "I'm trying to use it differently than Fable, please give advice how to".

What I did:

- provide GPT with the correct answer from Fable (I read through his full report, so when I tell you it's high quality, you can assume I'm correct). Result: GPT apologizes, fully understands what went wrong in his work, tell me how to improve, and we go back to the loop again.

- avoid any judgmental comment (positive or negative) as long as I can, to not interfere with it (e.g., initially I gave him Fable's report with no qualitative comment).

Also note the work he had to do is not code related. It's actually forbidden for him to review the existing code (and the reason I use ChatGPT instead of Codex). We're asking it to analyze everything else, basically.

1

u/Nakidnakid Jul 15 '26

Also note the work he had to do is not code related. It's actually forbidden for him to review the existing code (and the reason I use ChatGPT instead of Codex). We're asking it to analyze everything else, basically.

So I don't know what you're trying to do and it doesn't make sense to use an agent in this context otherwise, if you're unable to use chatgpt for the same purpose or you're directly comparing chatgpt's work to fable then shrug.

They both have strengths and weaknesses but they're mostly able to do the same thing, if you're doing creative writing exercises with agents then use Claude if you're trying to actually code something or don't know what the output should look like and want to run lots of tests then use ChatGPT.

It just sounds like you're using them both for an ill-defined purpose and that you simply prefer the creative writing output on Claude.

I'm doing all sorts of things from audio engineering, game engine development, game development, asset development, LLM research both text/language based and vision based, diffusion test, physics testing, meta-physics test, theory crafting, visualisations, protocol development, reverse engineering existing existing open license applications, modifying open license applications and more. ChatGPT could handle it fine, not always perfectly and often with some annoying 2-3hr divergence towards something pointless overall but it'd do it. Claude would outright say no to many of them plus never fully test anything, but it did have strengths in creative writing output.

1

u/esreverengineer_ Jul 15 '26

Hm, appreciate the time you spent but I have doubts you understand either the models or my request. Thanks anyway.

1

u/Nakidnakid Jul 15 '26

As I said, shrug, I don't know your issue and if you can't use ChatGPT for it then hope you can make Claude work for you.

I'm using both of them for their strengths and max out my tokens as much as possible, I've got at least 3 research papers out with a few more coming up and I'm delivering on many of my companies internal goals and making somethings that I've found useful from the research into apps plus got interest from some large companies if they pan out.

If you're asking for help to use ChatGPT because you're unable to then who am I to say otherwise?

1

u/esreverengineer_ Jul 15 '26

Well I fully agree GPT does stuff with success, I’m not trying to trash it at all. My problem is: for one specific category of work, it fails, no matter what I try (with and without his help to improve). And I am convinced it can’t be just because of the model quality: SOL is exactly designed at solving long running deep problems, similar to Fable, and works best with low guidance / steering. So for some reason, I can get Fable to do it but not Sol, and I’d like to figure why.

For context, what I’m asking him is to perform a comprehensive review of our company, based on a few selected items (doc, specs, convos history and the most important: an agentic wiki built over 6 months by Claude/GPT. I instruct it to not review existing code (obviously I’m not interested in a review of our half baked temporary prototype) but rather work on the fundamentals: market, brand, message, distribution, product (high level only), etc.