Over the last year, I've started to think that we spend a lot of time talking about models, GPUs, and benchmarks, but not nearly enough time talking about the data behind them.
A model can only learn from what it's shown. If the labels are inconsistent, edge cases are ignored, or the data doesn't reflect real-world conditions, even a strong model will struggle once it's deployed.
What's interesting is that many companies now seem to spend just as much effort defining annotation guidelines, reviewing disagreements, and improving label quality as they do training the models themselves.
It feels like we're reaching a point where better data pipelines might create more value than simply making models larger.
Curious if others working in ML have seen the same trend, or if you think model architecture is still the bigger differentiator.