The problem hiding inside deep belief nets
A densely connected, many-layered belief net is hard to train because inferring what its hidden units are doing, given an input, is intractable. Competing explanations for the same observation cancel each other out, a phenomenon the paper calls explaining away, and the true posterior cannot be computed exactly. [2]
Hinton, Osindero and Teh's fix was a complementary prior: extra hidden layers whose statistics cancel the explaining-away effect below, leaving a posterior that factors cleanly and is cheap to sample. They showed this construction is equivalent to a restricted Boltzmann machine, an undirected model with one visible and one hidden layer joined by a single weight matrix. [2]
Train one layer, freeze it, add another
The greedy algorithm trains each layer as its own restricted Boltzmann machine with contrastive divergence, 30 passes through the training set, then treats that layer's activities as data for the next one. The MNIST network ran 784 pixel inputs into two hidden layers of 500 units each and a top layer of 2,000 units joined to 10 label units, about 1.7 million weights. [2]
Greedy pretraining alone reached 2.49 percent errors on MNIST's 10,000 test digits. A further pass, a contrastive wake-sleep variant the paper calls up-down, adjusted every layer together and cut that to 1.25 percent, after about a week training on the full 60,000-image set. [2]
1.25 percent, and against what
The paper's own comparisons put 1.25 percent ahead of a support vector machine at 1.4 percent, and ahead of backpropagation-trained nets with no hand-crafted structure for the problem, at 1.5 percent. [2]
Atlas interpretation: That is narrower than beating the best digit classifiers of 2006: the same paper notes convolutional nets and support vector machines using domain-specific tricks already scored between 0.4 and 0.95 percent. The achievement was making an unstructured, general-purpose network competitive with tuned methods, not surpassing every published approach. [2]
Why one benchmark result changed what people worked on
By the late 1990s multilayer nets had largely been set aside, on the assumption that training deep feature extractors from scratch would get stuck in poor local minima. A group brought together by the Canadian Institute for Advanced Research, including Hinton, revived interest in deep feedforward networks around 2006 using pretraining of exactly this kind, first on digits and pedestrians, then, from 2009, on speech. [3]
Atlas interpretation: Pretraining's legacy was narrower than it looked. Later work found it mostly mattered for small datasets, and a different architecture, the convolutional network, won on images. This paper's contribution was reopening a door, not the technique that walked through it. [3]
Six years later, same author, a different architecture
Atlas interpretation: Hinton co-authored this paper at Toronto in 2006, then in 2012 co-authored AlexNet, the convolutional network whose ImageNet result ended the field's skepticism for good. Six years separate the two, the specific gap behind the claim that researchers who stayed with deep networks turned out to be ahead of everyone else. [4]
Sources
- A Fast Learning Algorithm for Deep Belief Nets
Neural Computation · Jul 2006
- A fast learning algorithm for deep belief nets
University of Toronto · Jul 2006
- Deep learning
Nature · May 28, 2015
- ImageNet Classification with Deep Convolutional Neural Networks
NeurIPS · Sep 9, 2026