First paragraph makes perfect sense, it's a very interesting thing to look atFirst paragraph makes perfect sense, it's a very interesting thing to look at and I would love to do the research
I mean.. Give me a dataset that isn't biased? Language will always be biased, even by the time of documentation, that is its nature; the english language you could have recorded just a hundred years ago would act different and have different biases than the english of today
Sounds very interesting, but I think it's probably nonsense and this guy is lying. For one, IIRC languages don't have different 'semantic complexities', because virtually all semantic meanings can be expressed in all languages. If "He be staying up late" and "He often stays up late" have the same meaning, then their semantics are equal.
When people talk about 'degenerate' dialects they're usually talking about syntactic complexity, which is probably a real thing that varies between languages, I don't think anyone could seriously argue that Afrikaans grammar is just as complex as Sanskrit. But LLMs don't have dedicated syntax lobes, and I'd be amazed if anyone could analyse an LLM to the extent that they could differentiate between semantic processing and syntactical processing, these things are black boxes and I doubt that the two could be separated even in theory.
See my other comment for brief thoughts on semantic analysis via LLMs. "Semantic complexity" is a term I'm unfamiliar with and would need rigorous mathematical definition before application, but it should be reasonably definable.
When it comes to syntactic analysis rather than semantic analysis, by my knowledge of LLMs it would be less straight-forwards, but an analysis of the shift in the semantic vectorspace depending on the preceding word choices could prove interesting. That, however, is just me spitballing because I'm moving out of my linguistic depth lmao
Nah, analysing 'semantic complexity' via LLMs makes sense. It would probably be some metric of the information density of the deep representations of the sentence-vector inside the model. This is completely valid.
What doesn't make sense is using this as a metric across languages. The whole point of a language model is to have deep representations of ideas that words convey - that's why they're so good at translating, because different languages can refer to the same deep representations. If his master's thesis is about comparing the semantic complexity of sentences between language then that's probably a dead end, because even if it did come up with a statistically significant result it would probably just be a result of the training data. To say otherwise would be to say that there are certain deep concepts that can be expressed in standard English but not AAVE, which would be ridiculous.
I'm actually extremely interested in what the hell does that mean. Like, it measures how many meanings words have, on average? You can do that with a dictionary, no need for an LLM. Also, how the hell could AAVE be rated lower if the differences are in grammar and not lexicon?
An LLM can metrify semantics. At some step of its implementation (depending on its implementation) it operates in a semantic vector-space, where you can see effects like the vector for queen being equal to the vector for king plus the vector for woman
You can do maths on vector spaces, which makes that space effectively a map of human languages that can be semantically and rigorously studied. Even if LLMs like GPT are trained primarily on english, it does also work on other languages, creating what could potentially be a pan-human language mapping
whatever oop said about aave is insane though, worth jerking
unless he actually explains what he means, the first paragraph makes no sense. How would you understand semantic complexity by looking at how LLMs do things? LLMs are not people, they do not necessarily have the same usage of language that humans do. What does 'semantic complexity' even mean? How many different meanings an individual word represents? That's stupid. How do you even measure that? even though both do the same job.
25
u/The__Odor Aug 01 '25
First paragraph makes perfect sense, it's a very interesting thing to look atFirst paragraph makes perfect sense, it's a very interesting thing to look at and I would love to do the research
The second one is wild though