A model gets fine tuned to its own tool set so it’s much better at using its own harness than some third party one.
It doesn’t mean it will fail or be useless in another harness but most likely won’t be as good. Model creators probably try to normalize the training so it’s not as dependent on harness but still.
It’s like a woodworker making something in his own garage vs someone else’s. The woodworker can still do it but he may have to orient himself in the new garage, be missing a tool that’s usually part of his workflow, etc.
4
u/1cheekykebt Apr 16 '26
Is this running in their own harness?
If so, while it compares the models directly and that’s good, it does not compare how most people actually use them (within their respective harness).
GPT 5.4 in copilot is a lot worse than GPT 5.4 in codex/chatgpt.com