Separating the inseparable
Cortes and Vapnik described a "new learning machine for two-group classification problems." The method maps input vectors non-linearly into a very high-dimensional feature space, then builds a linear decision surface inside that space. A boundary that has to bend around real data can become a straight cut once the data is lifted somewhere higher-dimensional. [1]
Atlas interpretation: A feature space rich enough to separate messy data can have enormous, even infinite, dimension, and computing coordinates there directly is not practical. A kernel function sidesteps that by returning what the dot product of two points would have been in that space, without ever constructing it. The trick is arithmetic, not geometry. [2]
Whose trick, and from when
The kernel substitution is not this paper's contribution. Boser, Guyon and Vapnik had already used it three years earlier to turn a maximum-margin hyperplane classifier into a nonlinear one, applying it to perceptrons, polynomial classifiers and radial basis function networks at COLT 1992. [2]
Atlas interpretation: What Cortes and Vapnik added in 1995 was the part that made the method usable on real data: a soft margin. Their paper is explicit that the earlier version "was previously implemented for the restricted case where the training data can be separated without errors," now extended to non-separable data with slack variables that let some points sit on the wrong side at a cost. Real data is rarely cleanly separable even after a kernel; the soft margin, not the kernel, is what turned a geometry demonstration into a general-purpose classifier. [1]
Good enough on messy data
The paper's own test came from US Postal Service zip code digits, run against established classical learning algorithms in the same benchmark. Cortes and Vapnik reported an error rate near one percent, at a rejection rate near nine percent, competitive with methods built specifically for that problem rather than adapted to it. [1]
Atlas interpretation: A single recipe that generalized across unrelated feature spaces, backed by the margin theory Vapnik had spent decades developing, suited the 2000s: labeled data existed, but rarely enough to make hand-built kernels anything but the practical ceiling. [1]
What displaced it, and why not sooner
AlexNet won the 2012 ImageNet competition with a top-5 error rate of 18.9 percent, which the authors called considerably better than the prior state of the art. Earlier winning entries there had typically paired hand-engineered image features with a conventional classifier. [3]
Atlas interpretation: A kernel still requires someone to decide what similarity between two raw inputs should mean before training starts. Once enough labeled images existed for a network to learn that similarity itself, the fixed kernel became the bottleneck, not the generalization guarantee. That is a narrow claim: it explains why the method receded first in vision and speech, not that SVMs stopped working on the smaller, structured problems they were built for. [3]
Sources
- Support-vector networks
Machine Learning · Sep 1995
- A training algorithm for optimal margin classifiers
Proceedings of the Fifth Annual Workshop on Computational Learning Theory (COLT '92), ACM Press · Sep 9, 2026
- ImageNet Classification with Deep Convolutional Neural Networks
Advances in Neural Information Processing Systems 25 (NeurIPS 2012) · Dec 3, 2012