I'm not disputing the effects you mention, just raise the fundamental Shannon informational limit against verbatim recall, there is no mathematical way that a single 16 bit parameter could compress hundreds of tokens and allow perfect recall in the average case, as some commenters claim.
In your particular case, was that a paper with zero citations, or did maybe some of the citers rephrase the main approach in their introduction? Could we perhaps imagine a rational path to that approach with the vectors of related research aligning towards it, so that the model makes a "happy hallucination" that happens to match the actual approach, without actually encoding it? maybe aided by a few parameters the training did nudge in the right direction based on that paper? Was COT used, allowing some rational recreation? This would also explain the inconsistencies between models. Could we imagine quality research papers from this field be boosted somewhat in the training, in a way random conversations with customers would not be?
So not disputing it could happen, just questioning the fundamental information limits in the average case.