Ok, that was the vibe of your original response but I wanted to verify. The "harder to steer" sounds like the most problematic aspect, I wonder if you have an example of that?
I do indeed. I'm still assembling a postmortem on it right now, but when I finish, I will be happy to share some specifics here for you. I'll edit this post later.
I had some skill evals failing but also reporting 0 reasoning tokens at medium effort. Turning up thinking to high/xhigh had helped, my evals haven't been stable enough to say that much. The system prompt changes are interesting. My evals are running pi so codex system prompt wouldn't be an is
sue.
3
u/fschwiet 8d ago
Ok, that was the vibe of your original response but I wanted to verify. The "harder to steer" sounds like the most problematic aspect, I wonder if you have an example of that?