Why gradients die crossing time
A recurrent network learns by pushing an error signal backward through every earlier time step, backpropagation applied along the unrolled sequence. Hochreiter and Schmidhuber showed that signal is scaled by a derivative and a weight at each step; multiplying that scaling across hundreds of steps either explodes into oscillating weights or vanishes into nothing. [1]
With the logistic sigmoid, that scaling stays below 1.0 whenever the connecting weight's absolute value is under 4.0, true of most networks early in training, and larger weights do not help since the derivative shrinks faster than the weight grows. The network learns recent history and stays structurally blind to anything a thousand steps back. [1]
A cell built to refuse to decay
The fix is a deliberately boring unit: a linear cell with one self-connection, fixed at weight 1.0, so a stored value's error signal neither grows nor shrinks going back out. Hochreiter and Schmidhuber call this the constant error carousel, the part of LSTM that actually solves the vanishing-gradient problem. [1]
Unguarded, every input would overwrite the carousel and every downstream unit would read it too early. The 1997 architecture wraps it in two multiplicative gates: an input gate deciding when a new value may overwrite the cell, an output gate deciding when the stored value may affect the network. Guarded that way, a value sits untouched for over 1,000 time steps, on tasks earlier methods could not solve. [1]
Two gates in 1997, a third arrives in 2000
Atlas interpretation: The 1997 paper, read directly, gives the carousel an input gate and an output gate and nothing else. There is no forget gate in the architecture the paper describes. The now-standard third gate, letting the network actively erase what it stored, arrived three years later: Gers, Schmidhuber and Cummins added it in 2000, because a carousel that only accumulates grows without bound on a continuous input stream with no point to reset it. [1][2]
That the forget gate reads as original equipment is not a small mix-up. A 2017 comparison of eight LSTM variants across speech, handwriting and music found the forget gate and the output activation function to be the two components whose removal actually hurt performance. The architecture most people mean by "LSTM" today is functionally the 2000 version, not the one cited here. [3]
Two decades of quiet duty, then a replacement
With the forget gate added, LSTM ran sequence modeling for the next decade and a half. Stacked LSTM encoders and decoders, for instance, turned one sentence into another for Sutskever, Vinyals and Le in 2014, translating English to French ahead of the phrase-based system it was compared against, with no attention mechanism involved. [4]
Atlas interpretation: The Transformer that eventually displaced this design kept the encoder-decoder shape but discarded the recurrence the carousel depended on, computing relationships between positions directly instead of carrying them through a thousand sequential steps of gate-protected memory. [1]
Sources
- Long Short-Term Memory
Neural Computation · Nov 1997
- Learning to Forget: Continual Prediction with LSTM
Neural Computation · Oct 1, 2000
- LSTM: A Search Space Odyssey
arXiv · Oct 4, 2017
- Sequence to Sequence Learning with Neural Networks
arXiv · Sep 10, 2014