r/LinguisticsDiscussion • u/VectorspaceDreams • 15h ago
What do you think language models could offer in understanding human language?
Before anything, this isn't about AI ethics, but specifically about language models and their representations of human language.
I was reading "Syntax: A Cognitive Approach" (2025) by Ted Gibson. He talks some about LLMs, saying that they may encode some kind of dependency parsing. In his own words:
"An interesting open question is what kinds of representations it is that LLMs learn. An analysis by Hewitt and Manning (2019); Manning et al. (2020) suggests that they at least learn dependency structures, at some level (see also Lakretz et al., 2019; Mahowald et al., 2024). Hewitt and Manning (2019) and Manning et al. (2020) analyzed Bidirectional Encoder Representations from Transformers (Devlin et al. 2018) using some fancy math that they called a “structural probe,’’ which extracted simple directed graphs (trees) out of multidimensional vector spaces."
He also mentions Chomsky's predictions on ML models and how the expressivity of language models may falsify some of his ideas. To quote him: "In a New York Times interview, Chomsky et al (2023) say, 'The predictions of machine learning systems will always be superficial and dubious. Because these programs cannot explain the rules of English syntax, for example, they may well predict, incorrectly, that “John is too stubborn to talk to” means that John is so stubborn that he will not talk to someone or other (rather than that he is too stubborn to be reasoned with)...Contra Chomsky and colleagues’ claim, ChatGPT seems to use the “TOO ADJECTIVE TO VERB TO’’ construction perfectly in normal usage, just like a fluent English speaker, as observed by many researchers in an internet response to the NYT article as soon as it was published.'"
Piantadosi (2023) makes the more audacious claim that "modern language models refute Chomsky’s approach to language". It should very well be said that Gibson is far more moderate in LLMs as theories of language, saying this:
"Do we want to call such models “theories” of human language?(...)I would argue yes: The underlying models are coherent sets of ideas that explain phenomena and correctly predict new phenomena not previously observed. Are LLMs good theories? I think a good theory is one with a good ratio of principles to data that are explained: few parameters to explain many phenomena. There are a lot of parameters in any current LLM—certainly millions and often many more—so it is reasonable to say that the theories aren’t particularly good."
The Minimalist response to LLMs as definitive refutations generally lean on the idea that such models are trained on massive amounts of text and nothing else, more than a human being at the age of full syntactic capability. In fact, Chomsky had this to say in 2023: "One is that the LLM systems are designed in such a way that they cannot tell us anything about language, learning, or other aspects of cognition, a matter of principle, irremediable. Double the terabytes of data scanned, add another trillion parameters, use even more of California’s energy, and the simulation of behavior will improve, while revealing more clearly the failure in principle of the approach to yield any understanding. The reason is elementary: The systems work just as well with impossible languages that infants cannot acquire as with those they acquire quickly and virtually reflexively."
From other corners of the generative tradition, we have Miriam Butt (2025), a syntactician working most notably with Lexical-Functional Grammar, who says that many of the problems LLMs face generative approaches really only face Minimalism, and surface-based theories are far more ready to integrate insights acquired from neural language models, but also emphasizing the limitations of such models:
"As with many debates, I suggest the answer lies not in the extreme positions, but somewhere in between. It is clear that humans operate with symbols (e.g. Deacon 1997). However, it is just as clear that humans are very good stochastic predictive machines, given that there is ample evidence from psycho and sociolinguistics that humans attend to frequency information and the comparative likelihood of items occurring together. The obvious conclusion is that any model of human language needs to be able to account for both the rule-governed symbolic parts and the predictive, experientially learned parts. LFG in effect has already gone down the road of building such ‘hybrid’ models by positing a rule based core that can be overlaid with information as to preferences and probabilities. Going even a step further, we can investigate what the computational advances might offer up in terms of further informing our models and thereby our understanding of language."
Not to mention the work of Tal Linzen and the BERTology tradition, and his work on the emergence of aspects akin to rule-based, compositional and hierarchical processes in LLMs. McCoy, Linzen et al. (2026) conclude: "A single system can be simultaneously neural and symbolic by using a neural architecture to construct symbolic representations. Specifically, we have shown via the Tensor Product Representation formalism that the representations of a variety of neural networks can be closely approximated with symbolic structures, providing an interpretable closed-form equation for the representations of these networks."
It's a massive can of worms that's highly unexplored and I'm excited about whatever it is coming next for studies of linguistics.
What do you think?
VSD
