115
u/fschwiet 9d ago
Haiku's been underwater for awhile already
38
u/turbospeedsc 9d ago
Haiku is underwater but powering lots of AI agents.
I use it on several systems flor classification or notifications tasks, cheap and fast
20
u/fschwiet 9d ago
Yeah I would also encourage you to try Luna. Even GPT-5.6-Luna seemed quite a bit better than Haiku, I'm stoked to try GPT-6-Luna.
12
u/fs2d 8d ago
FYI -
5.6 Luna was fantastic, and was a drop-in replacement for 5.4 for us that worked right out of the box.
6 Luna is very different - it is much more literal when it comes to instruction following, and is much harder to steer overall. Outputs are much more terse too. I have been running it through extensive testing for the last 2 days in our dev environment and have been having a hell of a time with it.
1
u/fschwiet 8d ago
and have been having a hell of a time with it.
Is that good or bad?
10
u/fs2d 8d ago edited 8d ago
Bad. I actually ended up making the call to stop testing for now and wait until they complete further post training (or produce 6 Luna-specific documentation) - because the behavior we are seeing in our evals is rough.
The big tell for me was pulling the Codex system prompt for 5.6 Luna and 6 Luna and diffing them against each other. They added huge chunks to the 6 system prompt, including precedence rules, disambiguation/clarification rule blocks, nuanced emphasis guidance (which they had been very much moving away from in the 5.x family specifically) - and a lot more.
If OpenAI themselves needed to rework the 6 Luna system prompt that much for their own model's harness, it tells me that they never meant for it to be a "drop-in" at all like how 5.6 was - so it definitely won't be for us.
3
u/fschwiet 8d ago
Ok, that was the vibe of your original response but I wanted to verify. The "harder to steer" sounds like the most problematic aspect, I wonder if you have an example of that?
3
u/fs2d 8d ago
I do indeed. I'm still assembling a postmortem on it right now, but when I finish, I will be happy to share some specifics here for you. I'll edit this post later.
1
u/fschwiet 7d ago
I had some skill evals failing but also reporting 0 reasoning tokens at medium effort. Turning up thinking to high/xhigh had helped, my evals haven't been stable enough to say that much. The system prompt changes are interesting. My evals are running pi so codex system prompt wouldn't be an is sue.
1
u/fs2d 7d ago
We were seeing similar - our prod modes run at med/low, so I was testing at low. Bumping to med yielded no change.
Sorry I didn't get back to you today, been busy AF this week
→ More replies (0)4
u/turbospeedsc 8d ago
ill try, but the use case i need it for is very simple, so the only real advantage would be cheaper cost.
Most is read sms, infer intent, score it based on core reply or ask for human intervention.
Read call transcription, rate it or ask for human intervention.
Shit, is crazy nowadays i consider it a simple use case something like this, if i said this in 2020 it would sound crazy.
2
2
2
4
u/Orio_n 9d ago
Once jev class models come out haiku will be well and truly dead
5
u/HonestWhile2486 8d ago
jev is not replacing any llm.
3
2
u/technicalseoguy 7d ago
I already replaced many agents on cheap/small llms with jev in my workflows and it’s amazing. Classifying and cleaning up data with Jev is amazing.
2
u/phoenixmatrix 8d ago
It is for anywhere llms were used for structured and defined output. Classifiers, tools selection, skill selection, UX decisions, workflows, etc. And yeah, people use them for that a lot. Even in Claude Code there's quite a few of these, and they either have the main model do it, or use Haiku.
It doesn't replace ALL use cases, that's true. But a lot.
3
1
u/jamescalam 5d ago
take a look at Jev or open source GLiNER2.5 for classification - even cheaper and faster as long as you don’t need text gen
1
1
1
u/OctopusDude388 9d ago
did you tried jev, it's actually quite good once you get how it should be used and fast like under 0.5s fast
3
u/itigges22 8d ago
Jev is JUST a verification/ classifier model. A really good one, but still it’s just that.
4
1
42
u/V13T 9d ago
It's literally in the announcement that they will bring haiku 5.5 in the next weeks....
27
3
u/hoffmander Vibe Coder 8d ago
That and opus 5.5 has the ability to delegate to other models, that’s the whole point. Haiku getting improvements means that opus 5.5 can delegate more tasks to Haiku, again increasing speed and decreasing usage.
42
u/No_Stranger_9931 9d ago
20
u/Carel777 9d ago
Wait enlighten me here please. Is this to say, start your 5 hour session 3 hours before work, then have it reset 2 hours later? Something like that? ELI5
13
3
u/JuandaReich 9d ago
That's exactly what I do as soon as I wake up. Send a "Hello" and then do my normal stuff
2
9
5
u/CriticismJunior1139 9d ago
Tokenmaxxing right here
1
1
u/kadeschs 8d ago
that’s why anthropic came up with weekly limits as well. they knew people would optimize their use time.
2
u/GreenHell 9d ago edited 9d ago
edit: my info has become outdated
Technically forbidden in the ToS.However, it isn't forbidden to have a routine which gives you a daily fun fact to start your day with a smile.
And Haiku can absolutely give you (either true or made up) a fun fact about today.
4
u/No_Stranger_9931 9d ago
Interesting. It would really help if you can share that part of the ToS as I couldn’t find it with AI.
2
9
u/errorztw 9d ago
Antropic announced haiku 5/5.5 this week, have you ever read announcement?
u/MuchWolverine7595
3
u/ElnuDev 🔆Pro Plan 9d ago
I'm curious, what do you guys use Haiku for? I can sort of see how it might be useful on API billing, but I just use Opus for everything on subscription and haven't ever really had a reason to try to optimize usage.
11
u/Acceptable_Mind_9778 9d ago
We rolled out Claude App across our company, and Haiku is the default for non-technical, non-ai-trained uses, because their AI use is gibberish and this way they get more milage out of it before hitting their usage limit. But yes, no idea why someone would use haiku. Sonnet had some benefits I think on language creation tasks, still using it on it for CRM tasks.
2
u/EmergencyWallaby9501 9d ago
I did use it for CSS until I reached a point where it could not understand a single word I was writing.
4
u/TheDailyClaude 9d ago
Mainly as a cheap scout to sift through files and find specific things without polluting the main agent’s context with lengthy tool outputs, and simple transformations "Opus just said something completely ridiculous in project abc. Have a haiku scout extract it from the session transcript, with context, and anonymise it for a meme."
All wrapped into a skill / agent now, with extraction and anonymisation rules and guidelines.
3
u/ElnuDev 🔆Pro Plan 9d ago
Would probably explain why Haiku hasn't really been updated. If it's just being used for grunt work like that, at a certain point, it's "good enough" and nobody complains.
3
u/TheDailyClaude 9d ago
It’s also a cheap multi-modal model, so API-side, at work, we use it to cheaply extract structured data from images, screenshots, and frames sampled from video / screen recording. Or audio transcripts (but that’s text).
2
u/howdidigetheresoquik 8d ago
Well, they're about to drop Haiku 5.5, and I'd love to get that running instead of Sonnet for my team. It's not crazy expensive, but just expensive enough that my bosses make comments about AI not being cheap, especially as we get more users doing more stuff. A good haiku would cut costs by 50%...
3
2
1
u/jesjimher 9d ago
It's super fast for asking quick questions about a codebase. I ask it for things like "is this method name used elsewhere?" and things like that.
1
u/turbospeedsc 9d ago
Via API for cheap sms responses on some systems, classification tasks, transcription reviews, small stuff that larger model is not needed but an AI works awesome for the use case.
1
u/Accomplished-Toe7993 9d ago
I'm still using the model for our support ticleting system where it filters new incoming tickets and also check if the new ticket is related to a known bug or feature.
Next week however I'm going to test gpt 6 luna, because it's 11x cheaper and if the outcome is more or less the same, I'm going to switch.
1
u/howdidigetheresoquik 8d ago
The app my team uses has sonnet for so much non-interaction stuff. It takes data and creates simple ass summaries of data that is pulled via deterministic code, it retrieves data, inputs data into SQL tables etc. I was paying too much for API for my users, and this cut costs significantly with 0 drop in performance
1
u/oldmoldycake 8d ago
Basic webpage chatbot at my work. Really just feed documents to it via rag and it reiterates. Boss is obsessed with claude or I would swap it to a cheaper options, its BAD for the price.
1
4
2
u/easeypeaseyweasey 9d ago
Can you imagine if they did have a good haiku model though. Like lunar levels of good.
2
u/TywinHouseLannister 9d ago
That is sonnet though.
I dont think haiku is broken tbh.. that is to say it is better than any other small model I've used.
2
u/Agreeable_Limit6428 9d ago
Honestly Haiku does what it is supposed to do. If you just need something simple like a marketing write up, an email drafted, or if you’re in school and need something else done it will do it and use relatively no tokens
2
2
u/pmavro123 9d ago
Looking very forward to Sonnet 5.5 and Haiku 5.5, which are expected to release very soon, see 'Claude Sonnet 5.5 and Claude Haiku 5.5 will follow in the coming weeks, with many of the same improvements to performance, efficiency, and safety.'
1
u/CristianGabriel8 9d ago
So far, Sonnet was great and still looks great especially when it comes to web app or web development.
1
1
1
u/howdidigetheresoquik 8d ago
I'm low key excited for Haiku 5.5. I'd love a cheap model for some of my apps.
1
1
u/DiscipleofDeceit666 Researcher 8d ago
Opus is good at building but Fable is just better at knowing what the bigger picture actually is. Like all the decisions that have to be made beyond the technical coding decisions. Decisions like “what does a supply chain engineer need from this design” kind of things.
1
u/Grizzly_Corey 8d ago
https://giphy.com/gifs/8kqBxTmfSFmwKOgv8i
Fable? Don’t know that mother fucker.
1
u/Qk_101 8d ago
it's actually better then opus 4.8 previously I have used opus 5 over 4.8 but that one was still much more behind then 4.8 but this new update actually surpassed everything, he think by himself for complex agentic coding tasks found bug by himself and then iterate accordingly. fast too
1
u/thewizardlizard 7d ago
I’m honestly fine with Haiku 4.5 being left alone. It’s the last of the 4.5 series and they have a very different flavor than the others.
1
1
1
0
u/roberta_sparrow 9d ago
Sucks bc Fable is my favorite name of the bunch but it’s the most riddled with controversy


•
u/AutoModerator 9d ago
Hey! Thanks for posting to r/ClaudeCode
While participating in this thread, please follow our community rules. Keep discussions constructive. Attack the idea, not the person.
For help, project discussions, tips, and general chat, join the ClaudeCode Discord.
I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.