The problem perceptrons couldn't solve
Rosenblatt's perceptrons learned by adjusting connections between input and output units directly, closing the gap between an actual and a desired output. Rumelhart, Hinton and Williams named what that rule could not do: decide what a hidden unit, whose correct state the task never specifies, should be doing at all. Minsky and Papert had already shown a single perceptron layer cannot even learn exclusive-or. Hidden layers were the fix, if anyone could train them. [1]
The answer runs in two passes. A forward pass sets each unit's output as a nonlinear function of a weighted sum of its inputs, layer by layer. A backward pass applies the chain rule to compute how a change in each unit's input changes the total error, propagating that quantity from the output layer back through every hidden layer. Each weight moves against its share of the error, with momentum added so the descent does not oscillate. [1]
An algorithm with earlier authors
The three authors did not claim credit for inventing this: their paper credits variants to David Parker, by personal communication, and to a 1985 report by Yann Le Cun. [1]
Atlas interpretation: That modesty was earned. A later survey traces efficient backpropagation to a 1970 thesis by Seppo Linnainmaa, with no reference to neural networks, and credits Paul Werbos with applying the same calculus to networks specifically in 1974 and 1981, both years ahead of this letter. What made 1986 memorable was not the math. It was putting that math to work on a demonstration nobody could ignore. [2]
Two toy problems, not a benchmark
The case rested on two small tasks, not a leaderboard score. One network learned to detect mirror symmetry in a row of binary inputs using two hidden units, settling on an arrangement of weights the authors called elegant. Another, trained on family relationships such as who is whose aunt, organized its hidden units around distinctions never present in the raw input, English versus Italian, generation, family branch, and answered relationships it had not been trained on. [1]
A caveat, and a fix twenty years later
The authors flagged a limit of their own framing: despite describing networks of neurone-like units, they wrote plainly that the procedure, in its current form, is not a plausible model of learning in brains. [1]
Atlas interpretation: Backpropagation carried a second limit nobody here could see yet. Error signals pushed back through many layers tend to shrink or explode exponentially with depth, a problem later research would name, so deep networks trained this way stopped learning past a handful of layers. Hinton returned, twenty years on, to fix that specific failure with a network trained one layer at a time instead of end to end. The math this letter popularized still trains nearly every network on this timeline; what changed was how to hand it a usable gradient. [2]
Sources
- Learning representations by back-propagating errors
Nature · Oct 9, 1986
- Deep Learning in Neural Networks: An Overview
arXiv · Oct 8, 2014