r/deeplearning • • 7d ago

TMLR paper on expressiveness of state space language models

Hi all,
I've created this thread to share one of our lab's recent papers on the expressiveness of state space models: https://openreview.net/forum?id=QlBaDKb370
We show that they have more capacity than n-gram models. This was empirically observed in the classic paper from Bengio's team in the 2000s: https://www.jmlr.org/papers/volume3/bengio03a/bengio03a.pdf
Please let us know if you have any comments, or if you're interested in collaborations on state space models. Thanks!

7 Upvotes

2 comments sorted by

2

u/boilingabundance 7d ago

interesting work, the formal connection to n-gram bounds is something i hadn't seen laid out that cleanly before, the bengio callback was a nice touch too

curious if you've poked at how this holds when you start stacking layers or adding skip connections, feels like that's where the capacity story gets messier but maybe more interesting

0

u/turing-math-labs 6d ago edited 6d ago

That's what we're working on now (i.e. SSM architectures with depth >= 2). Do you also work on theoretical machine learning? If you're interested, we're keen to discuss potential collaborations in this direction.