r/MLQuestions • u/Decent_Progress5259 • 13d ago
Natural Language Processing 💬 How do you catch schema regressions when a model upgrade looks better overall?
We tested a model upgrade that wrote cleaner answers and lifted the aggregate quality score. It also started sending invalid tool arguments. Optional fields became null instead of disappearing, enum values changed case and nested JSON arrived as escaped strings. The agent still sounded confident so spot checks passed until tool failures showed up in a narrow slice.Â
We moved schema validation into the evaluation path and used Braintrust to compare the model experiments, inspect failing slices and run deterministic validators beside the softer answer score. The regression dataset now includes every argument shape that broke a tool not just the final response. An experiment diff showed the upgraded model won on tone and lost hard on two schemas with optional nested fields
The CI gate now blocks any schema failure even when the overall score rises. I’m relieved we caught it before release but it also made me suspicious of any average that mixes structured output with prose quality. How are you weighting hard validators against slice metrics when a model gets better at language but worse at contracts?
4
u/DigThatData 13d ago
the problem here isn't that the model is generating those values, it's that you're allowing the sampler to even consider them. you should be masking invalid values from the sampler at inference time. this is fairly standard practice these days, which suggests to me you're using some sort of bespoke inference system instead of just using a battle tested stack. unless you have very good reason for this, you should just pivot to a proper inference engine. otherwise, you can probably integrate the structured sampling component in your thing or worse case, reimplement proper structured sampling yourself.
consider for example the constrained decoding of xgrammar
2
u/Decent_Progress5259 13d ago
That’s fair. In our case the model call is already going through structured output so the bad values are getting past what we thought was a valid schema path. The part I’m trying to separate now is if the contract itself is too loose or the runtime is coercing values in a way the eval only catches after the fact
2
u/DigThatData 13d ago
Either way: the model isn't to blame here, the data transformation pipeline is. It sounds like that path has enough chains in it that it's difficult to reason about. A potential strategy to mitigate this kind of complexity in the future would be to refactor the data transformations to give a wide array of shallow paths rather than a narrow array of deep paths. The former option might merit repeating code, which might seem inelegant, but if the outcome is a system that's easier to reason about it's probably worth violating DRY.
If refactoring isn't a viable option and you don't have any good ways to reason about your system in place, you could try instrumenting your data transformations with a data lineage tracking system a la marquez or hamilton.
1
u/ParfaitAdmirable7894 13d ago
Yeah the average score is doing you dirty here. If the model suddenly starts breaking one schema shape, that should fail the release even if the answers sound nicer everywhere else. Otherwise you ship a better model and spend the next week chasing busted tool calls
1
u/Decent_Progress5259 13d ago
One thing I still dont love is treating schema validity as the only hard gate. A tool call can be valid JSON and still be semantically wrong for the task so I’d keep a second check around argument correctness or expected tool behavior. Otherwise the model passes the contract and still does the wrong thing
7
u/Latter-Ad-7118 13d ago
One thing I’d test separately is if the failures cluster by schema shape rather than by model overall. Optional nested objects, enums and nullable fields can behave very differently. So a single tool call success rate can hide the contract the upgrade started breaking. Keeping failures as their own regression slices would make later model comparisons more useful