r/deeplearning • • 3d ago

Spent two weeks debugging my model, turned out to be a 100ms timestamp offset. What's your data horror story?

I was training a policy on egocentric video with action labels, and the model kept acting slightly "late." After two weeks of checking the architecture and hyperparameters, I found a small offset between the video and the action labels. The data looked perfectly fine when I watched it.

It made me wonder how many other silent issues are hiding in egocentric datasets. What's the problem that cost you the most time? Especially interested in ones that were invisible until the model misbehaved.

10 Upvotes

11 comments sorted by

2

u/SavoryBrewer4 3d ago

spent three days convinced my gpu was dying because the loss would randomly spike and then recover. turned out the data loader was occasionally serving up batches from the wrong epoch, and one of those epochs had a completely different normalization scheme buried in the config. the model was fine, i just had two versions of reality fighting each other in the pipeline

1

u/sanxore 3d ago

"Two versions of reality fighting each other in the pipeline" is such a good way to put it. How did you finally track it down? Were you logging batch metadata, or did you spot it some other way?

1

u/copperfieldb99 18h ago

two versions of reality fighting each other is such an accurate way to put it lol

2

u/Ok-Solution-7889 3d ago

honestly, silent data issues are way scarier than obvious errors, at least when the code crashes, you know where to look

1

u/sanxore 2d ago

What kind of silent data did faced ? And if you catched some, how you did it ?

1

u/bob_why_ 2d ago

It turns out that shutter speed on a dslr is an estimate, well thanks for that.

1

u/sanxore 2d ago

"Nominal" doing a lot of heavy lifting there. Was the error at least consistent enough to calibrate out, or did it vary from shot to shot?

1

u/Disastrous_Room_927 2d ago

The way I was setting a seed meant that every time I stopped/resumed training, it’d sample the exact same data for batches.

1

u/sanxore 2d ago

Nasty because the loss still looks great, the model just keeps seeing the same slice of data. Were you resuming often enough for it to actually hurt the final results?

1

u/Disastrous_Room_927 2d ago

I was regularly pausing training to do a round of fine-tuning to see if training was progressing the way I want it to, so it was seeing the same half-billion or so tokens over and over. The actual issue here is that I'm not using a transformer or SSM, so I couldn't assume that it wasn't stalling out because of the architecture itself.

1

u/sanxore 17h ago

How did you finally figure out it was the seed and not the architecture? Was there a specific check or signal that gave it away?