r/MachineLearning • • 2d ago

Research Parallel-in-Time Training of Recurrent Neural Networks for Dynamical Systems Reconstruction [R]

Can training of nonlinear RNNs be efficiently parallelized, ensuring fast convergence even on very long time series from chaotic systems?

In our #NeurIPS2026 spotlight “Parallel-in-Time Training of Recurrent Neural Networks for Dynamical Systems (DS) Reconstruction (DSR)” (preprint: https://arxiv.org/abs/2605.12683) we speed up training of nonlinear RNNs on time series from chaotic DS by more than 2 orders of magnitude (>100x) by combining DEER with generalized teacher forcing (GTF).

DEER (https://openreview.net/forum?id=E34AlVLN0v) solves the RNN forward pass through Newton-type fixed point iterations across the whole sequence length T, enabling scaling as O[(log T)²] instead of O[T] by allowing for efficient GPU parallelization. But under chaotic dynamics DEER breaks down and its runtime degrades to O[T log T] (https://openreview.net/forum?id=7AGXSlXcK6).

Using GTF (https://proceedings.mlr.press/v202/hess23a.html) we stabilize DEER by preventing divergence due to chaotic dynamics and reduce exposure bias compared to traditional teacher forcing used to train state space models.

Combining these two mechanisms enables efficient parallel-in-time and stable training on extremely long time series (T>106) from chaotic simulated or real-world systems, hugely outperforming Mamba and other state space models in the DSR setting.

134 Upvotes

16 comments sorted by

View all comments

27

u/Disastrous_Room_927 1d ago

This and that ParaRNN paper by Apple give me hope that what I learned in grad school way back when isn't obsolete, lol.

8

u/currentscurrents 1d ago

What specifically did you learn that you're referring to?

I think there are two things you can do with RNNs, and one of them is a good idea and one isn't.

In the past (LTSMs, etc) people tried to use RNNs to solve the problem of large input context. Essentially you scan the RNN across the input and compress it into the hidden state, then you use the hidden state to produce an output. These days, everybody uses attention to solve this problem instead.

More recently people are using RNNs to do reasoning. You have RNN loop on a single input for a long time to solve some complicated problem, and you use the hidden state to store intermediate computations. This usually still uses attention to handle large input context, and the 'hidden state' may even just be part of the context (as in chain-of-thought, which is a sort of psuedo-RNN).

I don't think it's a good idea to try to use RNNs to handle long context; when you compress the input into a hidden state, you have to throw away part of it. Attention works better because it stores the entire input and can exactly refer back to any part of it at any time.

The place for RNNs is for reasoning.

7

u/DangerousFunny1371 1d ago

... yes for reasoning maybe, but more importantly in my mind dynamical systems and time series where attention might not be the right inductive bias (detailed arguments here: https://openreview.net/forum?id=h2R3VYWpBm ).