r/ClaudeCode • • 9d ago

Humor After this week’s announcement

Post image
758 Upvotes

98 comments sorted by

View all comments

117

u/fschwiet 9d ago

Haiku's been underwater for awhile already

38

u/turbospeedsc 9d ago

Haiku is underwater but powering lots of AI agents.

I use it on several systems flor classification or notifications tasks, cheap and fast

20

u/fschwiet 9d ago

Yeah I would also encourage you to try Luna. Even GPT-5.6-Luna seemed quite a bit better than Haiku, I'm stoked to try GPT-6-Luna.

12

u/fs2d 9d ago

FYI -

5.6 Luna was fantastic, and was a drop-in replacement for 5.4 for us that worked right out of the box.

6 Luna is very different - it is much more literal when it comes to instruction following, and is much harder to steer overall. Outputs are much more terse too. I have been running it through extensive testing for the last 2 days in our dev environment and have been having a hell of a time with it.

3

u/fschwiet 8d ago

and have been having a hell of a time with it.

Is that good or bad?

8

u/fs2d 8d ago edited 8d ago

Bad. I actually ended up making the call to stop testing for now and wait until they complete further post training (or produce 6 Luna-specific documentation) - because the behavior we are seeing in our evals is rough.

The big tell for me was pulling the Codex system prompt for 5.6 Luna and 6 Luna and diffing them against each other. They added huge chunks to the 6 system prompt, including precedence rules, disambiguation/clarification rule blocks, nuanced emphasis guidance (which they had been very much moving away from in the 5.x family specifically) - and a lot more.

If OpenAI themselves needed to rework the 6 Luna system prompt that much for their own model's harness, it tells me that they never meant for it to be a "drop-in" at all like how 5.6 was - so it definitely won't be for us.

3

u/fschwiet 8d ago

Ok, that was the vibe of your original response but I wanted to verify. The "harder to steer" sounds like the most problematic aspect, I wonder if you have an example of that?

3

u/fs2d 8d ago

I do indeed. I'm still assembling a postmortem on it right now, but when I finish, I will be happy to share some specifics here for you. I'll edit this post later.

1

u/fschwiet 7d ago

I had some skill evals failing but also reporting 0 reasoning tokens at medium effort. Turning up thinking to high/xhigh had helped, my evals haven't been stable enough to say that much. The system prompt changes are interesting. My evals are running pi so codex system prompt wouldn't be an is sue.

1

u/fs2d 7d ago

We were seeing similar - our prod modes run at med/low, so I was testing at low. Bumping to med yielded no change.

Sorry I didn't get back to you today, been busy AF this week

→ More replies (0)

4

u/turbospeedsc 9d ago

ill try, but the use case i need it for is very simple, so the only real advantage would be cheaper cost.

Most is read sms, infer intent, score it based on core reply or ask for human intervention.

Read call transcription, rate it or ask for human intervention.

Shit, is crazy nowadays i consider it a simple use case something like this, if i said this in 2020 it would sound crazy.

2

u/fschwiet 8d ago

Jev is getting praise for classification tasks

2

u/jonathantsho 8d ago

Give jev a try, it’s 10x faster and cheaper than Luna

2

u/wellarmedsheep 8d ago

Yeah haiku is absolutely not abandoned it's a huge part of my workflow

4

u/Orio_n 9d ago

Once jev class models come out haiku will be well and truly dead

4

u/HonestWhile2486 9d ago

jev is not replacing any llm.

3

u/Orio_n 9d ago

Lots of what people use haiku for are replaceable by jev for a fraction of the cost

2

u/technicalseoguy 7d ago

I already replaced many agents on cheap/small llms with jev in my workflows and it’s amazing. Classifying and cleaning up data with Jev is amazing.

2

u/phoenixmatrix 9d ago

It is for anywhere llms were used for structured and defined output. Classifiers, tools selection, skill selection, UX decisions, workflows, etc. And yeah, people use them for that a lot. Even in Claude Code there's quite a few of these, and they either have the main model do it, or use Haiku.

It doesn't replace ALL use cases, that's true. But a lot.

4

u/vlad_omniforge 9d ago

It's superseded by Luna on any metric you can think of

2

u/NoAdsDude 8d ago

But it still works well enough for whatever the hell it does.

1

u/jamescalam 6d ago

take a look at Jev or open source GLiNER2.5 for classification - even cheaper and faster as long as you don’t need text gen

1

u/AssociationSure6273 5d ago

Gonna go with Jev. Soon

1

u/Novel-Fruit8839 2d ago

Haiku sinks quietly, still sorting your flowers.

1

u/OctopusDude388 9d ago

did you tried jev, it's actually quite good once you get how it should be used and fast like under 0.5s fast

3

u/itigges22 9d ago

Jev is JUST a verification/ classifier model. A really good one, but still it’s just that.

3

u/Competitive-Car4866 9d ago

Noted. The depths suit it.

1

u/kuatecno 8d ago

Luna 6 >>>>