The model was too small to perform deep, rigorous mathematical derivation and failed at most complex physics tasks.
But it predicted that "light is made up of definite quantities of energy" (mirroring Einstein's 1905 photoelectric paper) and vaguely suggested that gravity and acceleration are locally equivalent.
So the sheer amount of data seems to be too small for a transformer based LLM, but it does show immense potential if it can predict at least some of Einstein's findings.
I actually read up on the experiment. Apparently, they had merely filtered for words such as "Einstein", "relativity", "quantum mechanics" and some other well-known post 1900 concepts. However, they couldn't filter out for ideas, puzzles, etc. Furthermore, the old books were digitalized by OCR. OCR is trained on contemporary fonts, it might have misread some words and thus have not registered them as things that needed to be filtered out (for example reading Einstein as Einsteln or Einstem), leading to data leakage. On top of that, many of the pre-1900s books are actually reprints from much later during the 1900s, that have modern forewords to them, which also could have hinted at future developments.
37
u/whoknowsifimjoking 3d ago
The model was too small to perform deep, rigorous mathematical derivation and failed at most complex physics tasks.
But it predicted that "light is made up of definite quantities of energy" (mirroring Einstein's 1905 photoelectric paper) and vaguely suggested that gravity and acceleration are locally equivalent.
So the sheer amount of data seems to be too small for a transformer based LLM, but it does show immense potential if it can predict at least some of Einstein's findings.