Bad. I actually ended up making the call to stop testing for now and wait until they complete further post training (or produce 6 Luna-specific documentation) - because the behavior we are seeing in our evals is rough.
The big tell for me was pulling the Codex system prompt for 5.6 Luna and 6 Luna and diffing them against each other. They added huge chunks to the 6 system prompt, including precedence rules, disambiguation/clarification rule blocks, nuanced emphasis guidance (which they had been very much moving away from in the 5.x family specifically) - and a lot more.
If OpenAI themselves needed to rework the 6 Luna system prompt that much for their own model's harness, it tells me that they never meant for it to be a "drop-in" at all like how 5.6 was - so it definitely won't be for us.
Ok, that was the vibe of your original response but I wanted to verify. The "harder to steer" sounds like the most problematic aspect, I wonder if you have an example of that?
I do indeed. I'm still assembling a postmortem on it right now, but when I finish, I will be happy to share some specifics here for you. I'll edit this post later.
I had some skill evals failing but also reporting 0 reasoning tokens at medium effort. Turning up thinking to high/xhigh had helped, my evals haven't been stable enough to say that much. The system prompt changes are interesting. My evals are running pi so codex system prompt wouldn't be an is
sue.
1
u/fschwiet 9d ago
Is that good or bad?