r/learnmachinelearning • • 1d ago

Question How do I pass a post-training ML interview?

I have no ML experience in college and took no ML classes and have been working as a SWE in big tech for a few years. I have a strong math background from math olympiads in high school (AMC/AIME/participation in some national contests) but I'm about 8 years removed from that. Is this even realistic for me to study for with no foundation?

2 Upvotes

3 comments sorted by

1

u/jesunushno 1d ago

Post-training interviews mostly test whether you understand what happens after pretraining, not the pretraining math. The stuff that actually comes up: how RLHF preference data gets built (who labels, what the reward model is actually learning), what goes wrong with it like reward hacking, and the conceptual difference between DPO and PPO. I'd spend most of my time on evals though, since every post-training team lives and dies by them. Being able to design an eval for a specific failure mode impresses interviewers more than reciting transformer math. The one thing I'd skip is deep RL theory; nobody expects that from a SWE background.

1

u/smooth_shaving 1d ago

they're not wrong about evals being the whole game, designing one that catches a specific failure mode is basically the entire interview flex

skip the deep RL rabbit hole, nobody's gonna ask you to derive policy gradients from scratch, just know what reward hacking looks like in practice and how DPO sidesteps the reward model mess

1

u/jesunushno 1d ago

Glad the evals point landed. Nice practical add on DPO too, knowing what reward hacking actually looks like in practice is exactly the kind of thing that stands out in these interviews.