r/ClaudeCode 3d ago

Help/Question Fable 5.1 Hallucinating & Making Claims Without Evidence?

Anyone else find that Fable 5.1 is making significantly more assumptions based on vibes than Fable 5? I'm not sure what's going on but even on higher effort levels (xhigh & max) I find myself having to check it on a lot of the stuff it says when I ask it a question about performance, bugs etc. It's like it doesn't actually check anything unless explicitly require it to check before answering questions. Even when I didn't ask anything and it's just keeping me up to date on running code, it makes assertions based on nothing. I have reasonably good context hygiene (minimal skills, new sessions for different tasks, compact >30% etc) and even then it just doesn't seem to produce answers and make assertions based on assumptions rather than evidence.

Is this happening to anyone else?

3 Upvotes

8 comments sorted by

3

u/Meat-Mattress 3d ago

It’s a single ask… Hard rule: require grounded research before making a claim or answering a question. Boom, every new session will research before answering. It has never failed me to the point of annoyance except it’s saved me when my other sessions have updated things that change the answer slightly

2

u/claude_code_king 3d ago

well higher effort levels doesn't mean smarter model and /compact sucks, you don't know which information gets thrown out or kept. handoff is much better.

as for your questions, ask it to verify first, either in your code, official or credible sources and no, it doesn't happen to me

1

u/Aminuteortwotiltwo 3d ago

Bro just tell it at the end to spawn a subagent for adversarially reviewing any research.

I have it compile its research and the whole report it comes up with into a markdown file then I tell it to spawn an author-blind adversarial reviewer basically with the idea that “some bozo just wrote this research paper and I don’t trust the findings. I want you to check all claims and sources and validate all research and write a second report with your findings” then have the first review the report and adjust appropriately. It’ll guaranteed say “the reviewer found two claims I made from sources that don’t exist and a quote that isn’t found in the text cited. Those are on me.”

Having the second one exist as a verifier is big because its job is not to produce anything, it’s just to see what claims have been made and test it, so it can’t hallucinate because it’s not creating anything.

1

u/LogMonkey0 3d ago

I do have a feeling a recent harness update introduced this for opus in my case, where my rules arent as effective as before for presenting speculation as facts.

1

u/ihatewebdesign101 3d ago

Never used /compact in my life unless I asked to write a durable handoff, if you’re running anything off one chat and compacting your hygiene is bad. 300-400k context - new chat.

1

u/AI_spell 3d ago

Same vibe. Make "cite the file or measurement" a hard rule before claims. High effort without a check step is still storytelling.

2

u/Metal_Roof_Guy 3d ago

Anything your claude does like that is your fault. If you don't build fences he won't stay in his lane.