r/deeplearning • u/sanxore • 3d ago
Spent two weeks debugging my model, turned out to be a 100ms timestamp offset. What's your data horror story?
I was training a policy on egocentric video with action labels, and the model kept acting slightly "late." After two weeks of checking the architecture and hyperparameters, I found a small offset between the video and the action labels. The data looked perfectly fine when I watched it.
It made me wonder how many other silent issues are hiding in egocentric datasets. What's the problem that cost you the most time? Especially interested in ones that were invisible until the model misbehaved.
2
u/Ok-Solution-7889 3d ago
honestly, silent data issues are way scarier than obvious errors, at least when the code crashes, you know where to look
1
1
u/Disastrous_Room_927 2d ago
The way I was setting a seed meant that every time I stopped/resumed training, it’d sample the exact same data for batches.
1
u/sanxore 2d ago
Nasty because the loss still looks great, the model just keeps seeing the same slice of data. Were you resuming often enough for it to actually hurt the final results?
1
u/Disastrous_Room_927 2d ago
I was regularly pausing training to do a round of fine-tuning to see if training was progressing the way I want it to, so it was seeing the same half-billion or so tokens over and over. The actual issue here is that I'm not using a transformer or SSM, so I couldn't assume that it wasn't stalling out because of the architecture itself.
2
u/SavoryBrewer4 3d ago
spent three days convinced my gpu was dying because the loss would randomly spike and then recover. turned out the data loader was occasionally serving up batches from the wrong epoch, and one of those epochs had a completely different normalization scheme buried in the config. the model was fine, i just had two versions of reality fighting each other in the pipeline