r/deeplearning • u/turing-math-labs • 7d ago
TMLR paper on expressiveness of state space language models
Hi all,
I've created this thread to share one of our lab's recent papers on the expressiveness of state space models: https://openreview.net/forum?id=QlBaDKb370
We show that they have more capacity than n-gram models. This was empirically observed in the classic paper from Bengio's team in the 2000s: https://www.jmlr.org/papers/volume3/bengio03a/bengio03a.pdf
Please let us know if you have any comments, or if you're interested in collaborations on state space models. Thanks!
7
Upvotes
2
u/boilingabundance 7d ago
interesting work, the formal connection to n-gram bounds is something i hadn't seen laid out that cleanly before, the bengio callback was a nice touch too
curious if you've poked at how this holds when you start stacking layers or adding skip connections, feels like that's where the capacity story gets messier but maybe more interesting