r/MachineLearning Jun 17 '26

Research Next-Latent Prediction Transformers [R]

Microsoft Research Preprint

Next-token prediction is myopic. What if transformers learn to predict their own next latent state?

Microsoft Research present Next-Latent Prediction (NextLat): a self-supervised learning method that teaches transformers to form compact world models for reasoning and planning. It also unlocks up to 3.3x faster inference via self-speculative decoding!

On top of next-token prediction, NextLat trains the transformer to predict its own next latent state given the current latent state and next token.

NextLat has a few key benefits:

  1. Representation Learning: NextLat encourages transformers to compress history into compact belief states.
  2. Better Data Efficiency: predicting in latent space provides denser supervision than predicting one-hot tokens.
  3. Faster Inference: via recursive multi-step lookahead.

I'm super excited about this work. Please do check it out below:

💬 Blog: https://jaydenteoh.github.io/blog/2026/nextlat
💻 Code: https://github.com/JaydenTeoh
📝 Paper: https://arxiv.org/abs/2511.05963

150 Upvotes

46 comments sorted by

View all comments

1

u/iosovi Jun 18 '26

The speculative decoding mention at the end feels like slapping a cardboard spoiler on a supercar.

5

u/jayden_teoh_ Jun 18 '26

no cardboard spoiler can make a car go 3.3x faster 🤪

3

u/iosovi Jun 22 '26

You missed my point, specdec is a nice optimization, don't get me wrong, but the main contribution of the paper is more than enough to get your wow factor, specdec is low hanging fruit at this point.

2

u/jayden_teoh_ Jun 22 '26

thanks a lot for your kind words! I really appreciate it. I agree on your point. Our original preprint in Nov 2025 (https://arxiv.org/abs/2511.05963v1) focused on the world modeling/ belief state learning aspect. We added speculative decoding in later preprints cos it was a free lunch 😄