Two smaller models instead of one slow one
Efficient Estimation of Word Representations in Vector Space appeared on arXiv on January 16, 2013, from Tomas Mikolov, Kai Chen, Greg Corrado and Jeffrey Dean at Google. It proposed two model architectures, Continuous Bag-of-Words and Skip-gram, that learn a vector for each word from surrounding context without the nonlinear hidden layer earlier neural language models used. [1]
Dropping that hidden layer was the point. The paper reports training on corpora of up to 6 billion words, drawn from Google News with the vocabulary capped at the million most frequent words, and completing training in under a day. Earlier neural network approaches to word vectors had been too slow to run at that scale. [1]
Subtracting and adding words like numbers
The paper's stated goal was vectors that captured multiple degrees of similarity at once. Its example: the vector for King, minus the vector for Man, plus the vector for Woman, lands closest to the vector for Queen. The same offset that separated man from woman also separated king from queen. [1]
When Google published the tool as open source that August, its announcement gave a second version of the same trick: the relationship between Paris and France matched the relationship between Berlin and Germany, but not the relationship between Madrid and Italy. The post described the model arriving at that structure by reading news text, with no human labeling of which words were capitals or countries. [3]
A second paper made the vectors practical to train
Distributed Representations of Words and Phrases and their Compositionality, submitted on October 16, 2013 by Mikolov with Ilya Sutskever, Chen, Corrado and Dean, replaced the first paper's hierarchical softmax with negative sampling, a cheaper way to train Skip-gram, and added subsampling of frequent words to speed training further. [2]
It also extended the method past single words. Air Canada is not the sum of Air and Canada in the way word-by-word vectors assume, so the authors added a step that first found common phrases in the text and then learned a vector for each phrase as a unit. [2]
A tool, not just a result
Atlas interpretation: Publishing runnable code, not only a benchmark table, is what let word2vec spread. Other teams did not have to reimplement the training procedure to get vectors; they downloaded the tool, ran it on their own text, and plugged the output into whatever they were building. That is a different kind of adoption than a paper getting cited. [3][1]
Sources
- Efficient Estimation of Word Representations in Vector Space
arXiv · Jan 16, 2013
- Distributed Representations of Words and Phrases and their Compositionality
arXiv · Oct 16, 2013
- Learning the meaning behind words
Google Open Source Blog · Aug 14, 2013