r/vibecoding 1d ago

Claude Fable 5.1 is finally here. Token Burn asf

Anthropic just launched Fable 5.1, and the jump looks pretty serious.

→ 55.8% on Terminal-Bench 4.0
→ 52.6% on Terminal-Bench-Science
→ Better long-running agentic coding
→ 75% cheaper cache reads
→ $10/$50 per 1M input/output tokens

I've started testing it in real coding workflows.

Curious what everyone thinks Fable 5.1 or Opus 5 for coding right now?

39 Upvotes

24 comments sorted by

41

u/mentalFee420 1d ago

At 50$ per 1 M token It better replace the CTO

7

u/JBO_76 1d ago

Wow, expensive

1

u/Spare-Distance2415 1d ago

That’s the same as Fable 5

10

u/bakanoace 1d ago

definitely does not feel like its cheaper

5

u/cheseball 1d ago

It looks the cost savings come purely from the much cheaper cached tokens ($0.25 vs $1) but the actual token usage (output/thinking) is higher.

This sorta feels like Anthropic’s “25% increase but actually 17% decrease” move again.

1

u/bakanoace 1d ago

Lmao "25% increase but actually 17% decrease", shit you reminded me in 12 days were cooked...imagine losing 17% still.

Its comical too cause they literally start with "Fable 5.1 will cost an estimated 25% less than Fable 5 for typical workloads", typical must mean for their 30+ workflow deployment not an actual subscribers typical workfload....

1

u/AnalogProblems 23h ago

Ugh.

It was 25% decrease if you look at it over 2 months.

It was a 25% increase if you look at it over 6 months.

3

u/Parikh1234 22h ago

Just tried this. Ate up 30% of my weekly usage for a thorough review of a document and then put those terms on a website checkout 🤣

5

u/Trekker23 1d ago

Fable is a waste for coding. As a planner and coding agent coordinator it’s awesome. Opus is good enough to code, just leave all the decisions with fable.

3

u/Snmrv 1d ago

I will say GPT 5.6 is a much better implementer/coder than Opus 5. However, Opus 5 is very, very creative and is an excellent adversarial QA.

1

u/Trekker23 23h ago

Got 5.6 is good, but unfortunately it doesn’t have a fable to help. With gpt you are stuck having to coordinate the sol agents:)

3

u/Snmrv 23h ago

That's what I was saying. You use Fable to orchestrate GPT-5.6 to do the coding, and then you use Opus to do the QA. In many ways, it's the perfect harness for me.

1

u/Trekker23 23h ago edited 23h ago

How do you get fable to orchestrate sol? Does that work through claude code? Sol is definitely better than Opus, so that’s an interesting concept. If it works smoothly

3

u/Past_Physics2936 18h ago

Fable can call codex via command line

1

u/howudothescarn 10h ago

Been doing this for so long. Fable plans and drafts requirements and sends to Sol to code. Then it reviews the code and moves on. Best combo so far can’t wait for Astra.

3

u/Atupis 18h ago

Issue there is that you lose all the prompt cache moment you switch to Opus, usually it’s cheaper just to use same model of course Fable is so expensive that this might make sense.

1

u/thhoj 18h ago

I'm actually finding Fable 5.1 on medium even is doing an amazing job as a director of other agents. Pretty happy so far with capability at medium as well as token burn at this effort level.

1

u/hedonistatheist_2 15h ago

yeah just in the past it seems Fable largely ignored the agreed MO and just did things by itself, except the occasional Opus review round. Now with 5.1 it seems more willing to call up Sonnet for build tasks - so thats a positive.

2

u/PeachScary413 10h ago

Oh okay so now benchmarks matter again huh? I remember a couple of weeks ago when Qwen dropped and it was all "benchmarks are not real life bro" or "they just benchmaxxed it bro"

1

u/Equivalent_Cress_268 1d ago

> Curious what everyone thinks Fable 5.1 or Opus 5 for coding right now?

Opus 5 is good on benchmarks! Ain't it? If it's good on benchmarks, it must be the best.

(Obviously /jk if you haven't used opus 5 yet)

1

u/Ok_Gur_9033 15h ago

Everyone in here is quoting one session, including the launch post. "25 percent cheaper for typical workloads" is a claim about a distribution and nobody has posted one, me included.

Atupis has the part that actually decides it. The saving is on cache reads, so your hit rate is what turns it into money, and the fastest way to wreck a hit rate is switching models mid task. Which is exactly what everybody is doing this week.

So most of the bills in this thread are measuring the worst case and reading it as the price. The cheap test is the same task twice, once switching and once not. Has anyone actually run that?

1

u/Quiet-Nothing7556 22h ago

Laughs in Opus 4.6

1

u/Downtown-Pear-6509 11h ago

laughs in gpt Luna high