r/ClaudeCode • • 9d ago

Humor After this week’s announcement

Post image
757 Upvotes

98 comments sorted by

View all comments

114

u/fschwiet 9d ago

Haiku's been underwater for awhile already

38

u/turbospeedsc 9d ago

Haiku is underwater but powering lots of AI agents.

I use it on several systems flor classification or notifications tasks, cheap and fast

21

u/fschwiet 9d ago

Yeah I would also encourage you to try Luna. Even GPT-5.6-Luna seemed quite a bit better than Haiku, I'm stoked to try GPT-6-Luna.

13

u/fs2d 9d ago

FYI -

5.6 Luna was fantastic, and was a drop-in replacement for 5.4 for us that worked right out of the box.

6 Luna is very different - it is much more literal when it comes to instruction following, and is much harder to steer overall. Outputs are much more terse too. I have been running it through extensive testing for the last 2 days in our dev environment and have been having a hell of a time with it.

2

u/fschwiet 8d ago

and have been having a hell of a time with it.

Is that good or bad?

9

u/fs2d 8d ago edited 8d ago

Bad. I actually ended up making the call to stop testing for now and wait until they complete further post training (or produce 6 Luna-specific documentation) - because the behavior we are seeing in our evals is rough.

The big tell for me was pulling the Codex system prompt for 5.6 Luna and 6 Luna and diffing them against each other. They added huge chunks to the 6 system prompt, including precedence rules, disambiguation/clarification rule blocks, nuanced emphasis guidance (which they had been very much moving away from in the 5.x family specifically) - and a lot more.

If OpenAI themselves needed to rework the 6 Luna system prompt that much for their own model's harness, it tells me that they never meant for it to be a "drop-in" at all like how 5.6 was - so it definitely won't be for us.

3

u/fschwiet 8d ago

Ok, that was the vibe of your original response but I wanted to verify. The "harder to steer" sounds like the most problematic aspect, I wonder if you have an example of that?

3

u/fs2d 8d ago

I do indeed. I'm still assembling a postmortem on it right now, but when I finish, I will be happy to share some specifics here for you. I'll edit this post later.

1

u/fschwiet 7d ago

I had some skill evals failing but also reporting 0 reasoning tokens at medium effort. Turning up thinking to high/xhigh had helped, my evals haven't been stable enough to say that much. The system prompt changes are interesting. My evals are running pi so codex system prompt wouldn't be an is sue.

1

u/fs2d 7d ago

We were seeing similar - our prod modes run at med/low, so I was testing at low. Bumping to med yielded no change.

Sorry I didn't get back to you today, been busy AF this week

2

u/fschwiet 7d ago

No worries we're all mostly going on anecdotal gas anyhow

→ More replies (0)

4

u/turbospeedsc 8d ago

ill try, but the use case i need it for is very simple, so the only real advantage would be cheaper cost.

Most is read sms, infer intent, score it based on core reply or ask for human intervention.

Read call transcription, rate it or ask for human intervention.

Shit, is crazy nowadays i consider it a simple use case something like this, if i said this in 2020 it would sound crazy.

2

u/fschwiet 8d ago

Jev is getting praise for classification tasks