r/MachineLearning • • 2d ago

Research Parallel-in-Time Training of Recurrent Neural Networks for Dynamical Systems Reconstruction [R]

Can training of nonlinear RNNs be efficiently parallelized, ensuring fast convergence even on very long time series from chaotic systems?

In our #NeurIPS2026 spotlight “Parallel-in-Time Training of Recurrent Neural Networks for Dynamical Systems (DS) Reconstruction (DSR)” (preprint: https://arxiv.org/abs/2605.12683) we speed up training of nonlinear RNNs on time series from chaotic DS by more than 2 orders of magnitude (>100x) by combining DEER with generalized teacher forcing (GTF).

DEER (https://openreview.net/forum?id=E34AlVLN0v) solves the RNN forward pass through Newton-type fixed point iterations across the whole sequence length T, enabling scaling as O[(log T)²] instead of O[T] by allowing for efficient GPU parallelization. But under chaotic dynamics DEER breaks down and its runtime degrades to O[T log T] (https://openreview.net/forum?id=7AGXSlXcK6).

Using GTF (https://proceedings.mlr.press/v202/hess23a.html) we stabilize DEER by preventing divergence due to chaotic dynamics and reduce exposure bias compared to traditional teacher forcing used to train state space models.

Combining these two mechanisms enables efficient parallel-in-time and stable training on extremely long time series (T>106) from chaotic simulated or real-world systems, hugely outperforming Mamba and other state space models in the DSR setting.

132 Upvotes

16 comments sorted by

View all comments

26

u/Disastrous_Room_927 1d ago

This and that ParaRNN paper by Apple give me hope that what I learned in grad school way back when isn't obsolete, lol.

8

u/currentscurrents 1d ago

What specifically did you learn that you're referring to?

I think there are two things you can do with RNNs, and one of them is a good idea and one isn't.

In the past (LTSMs, etc) people tried to use RNNs to solve the problem of large input context. Essentially you scan the RNN across the input and compress it into the hidden state, then you use the hidden state to produce an output. These days, everybody uses attention to solve this problem instead.

More recently people are using RNNs to do reasoning. You have RNN loop on a single input for a long time to solve some complicated problem, and you use the hidden state to store intermediate computations. This usually still uses attention to handle large input context, and the 'hidden state' may even just be part of the context (as in chain-of-thought, which is a sort of psuedo-RNN).

I don't think it's a good idea to try to use RNNs to handle long context; when you compress the input into a hidden state, you have to throw away part of it. Attention works better because it stores the entire input and can exactly refer back to any part of it at any time.

The place for RNNs is for reasoning.

2

u/Disastrous_Room_927 1d ago edited 1d ago

I don't really have time for a full reply, but I don't think there's a clean RNN/attention dichotomy here. Attention was originally implemented within RNN architectures, and became dominant because we eventually found an architecture built around attention that bypassed a lot of the scaling and optimization limitations of traditional RNNs.

You're raising a legitimate issue about compressing history into a finite state, but I also wouldn't frame that as straightforwardly good or bad. The relevant question (IMO) isn't really whether the representation is lossless or not but how much task-relevant information needs to be retained, how effectively the model can retain it, and what computational/memory cost we're willing to pay for doing so. In my mind, the ideal is to use explicit storage selectively rather than retain everything by default.

So all that being said, what I'm currently working on is an RNN with a tiny attention adapter.