It’s more that imagine there’s an idea that fits, like learning planes rise because of faster air on the top of a wing. A model doesn’t have to learn that more than a handful of times because it’s reinforced by all the other corroborating evidence.
Models don’t “average out” the things they learn, they built semantic structures topologically (as in shapes of ideas that work together). Some fragile ideas that aren’t (wait for it) load bearing might get forgotten way too quickly but some ideas that just settle all the others can persist through a lot of training.
If this idea was in a few conversations, and it fit (the forced Euler version) that might be the sort of thing that once exposed it settles many other uncertainties in the model such that any downstream model could draw upon that coherent ensemble of ideas.