Reading checks, not just digits
The 1998 paper's most concrete claim is not about MNIST. Its abstract states that the check reader “is deployed commercially and reads several million checks per day.” The body dates that deployment precisely: the system was integrated into NCR's line of check readers and had been “fielded in several banks across the United States since June 1996.” LeNet-5 supplied the character recognizer; a graph-transformer layer around it handled segmentation and amount validation. [1]
MNIST was the paper's own benchmark, not a separate release. The authors built it from NIST's handwriting samples: 60,000 training images and 10,000 test images, size-normalized and centered the same way live check images were before reaching the recognizer. The check reader stopped running years ago; the benchmark it was tested on became the default first exercise for a new architecture. [1]
What was actually inside LeNet-5
LeNet-5 itself comprised seven layers past the input: two paired stages of convolution and subsampling, then two fully connected layers into an output layer, trained end to end by backpropagation. It held roughly 60,000 trainable parameters, small enough to train in two to three days on a single 200 MHz workstation processor. [1]
Why the design waited on data and hardware
The paper is explicit about why a design that old sat still. LeCun and coauthors wrote that in 1989 “a recognizer as complex as LeNet-5 would have required several weeks' training and more data than were available and was therefore not even considered”: LeNet-1 fit the hardware and data of 1989, LeNet-5 fit 1998. [1]
Atlas interpretation: The same limit held through the next decade. Bengio, LeCun and Hinton's own retrospective credits GPUs and larger labeled datasets, not a new idea, as what let deep networks scale past that point; ImageNet supplied over a million labeled images, and Alex Krizhevsky's efficient use of multiple GPUs did the rest. [4]
AlexNet's debt, and where it stopped owing one
AlexNet, trained on that dataset, used five convolutional layers and three fully connected layers with about 60 million parameters, split across two GPUs, over five to six days of training. [3]
Atlas interpretation: That is the same lineage as LeNet-5: convolution, pooling and full connection trained by backpropagation, and the AlexNet paper cites LeCun's own earlier convolutional-network work directly. It is not the same architecture. AlexNet is roughly a thousand times larger, adds ReLU activations and dropout LeNet-5 never used, and its authors credit the dataset and the hardware, not a rediscovered blueprint, for the result. “Essentially the one that would win ImageNet” overstates a continuity of concept as a continuity of design. [3][4]
Sources
- Gradient-based learning applied to document recognition
Proceedings of the IEEE · 1998
- Gradient-based learning applied to document recognition
HAL open science · 1998
- ImageNet Classification with Deep Convolutional Neural Networks
Advances in Neural Information Processing Systems 25 (NeurIPS 2012) · Sep 9, 2026
- Deep Learning for AI
Communications of the ACM · Jun 21, 2021