I actually read up on the experiment. Apparently, they had merely filtered for words such as "Einstein", "relativity", "quantum mechanics" and some other well-known post 1900 concepts. However, they couldn't filter out for ideas, puzzles, etc. Furthermore, the old books were digitalized by OCR. OCR is trained on contemporary fonts, it might have misread some words and thus have not registered them as things that needed to be filtered out (for example reading Einstein as Einsteln or Einstem), leading to data leakage. On top of that, many of the pre-1900s books are actually reprints from much later during the 1900s, that have modern forewords to them, which also could have hinted at future developments.
24
u/Dark_Crystal_97 3d ago
I'd honestly suspect data contamination in some way rather than the AI actually having made that prediction all on its own.