BERT: Bidirectional Training, Masked Words, and Google Search

Google’s BERT learned context from both sides of masked words, transferred to reading tasks, and later improved Search ranking and featured snippets.

Learning from the words on either side

Google's October 11, 2018 paper introduced BERT, a model that learned representations of text before being trained for particular tasks. The authors reported leading results on eleven language tasks. The paper still described code and pretrained models as forthcoming. [2]

Consider the difference between a river bank and a bank account. A useful representation of “bank” should change with its context. BERT allowed information from both sides of a word to influence its representation throughout the network. Google's later author explainer used this ambiguity to show why looking only at preceding words could miss useful information. [3]

Training required a way to prevent the model from simply seeing the answer. BERT selected 15% of input tokens, which can be words or word fragments, as prediction targets. Most were replaced with a mask; some were replaced randomly or left unchanged. The model learned to recover the originals. It also learned whether two text segments were consecutive or randomly paired. [2]

One starting point for different reading tasks

BERT reused the encoder side of the Transformer. The first GPT instead used left-to-right prediction. Both learned from ordinary text before supervised adaptation, but they exposed different context to each prediction. BERT's design was well suited to reading a complete passage rather than generating its next word. [2]

For question answering, fine-tuning taught BERT to identify the beginning and end of an answer in a supplied passage. For sentiment analysis, the task was to assign a label. These tasks still needed labeled examples, but developers could start from a pretrained model instead of teaching a new network language from scratch. [3]

By its November 2 announcement, Google had released TensorFlow code and pretrained models. This made the reusable starting point available to other researchers and developers; it was a separate step from publishing the October paper. [3]

Why a small word could change a search result

On October 25, 2019, Google announced BERT's use in Search ranking and featured snippets. It expected the ranking change to help with one in ten U.S. English searches. One example involved a Brazilian traveling to the United States: the old results had confused that direction with Americans traveling to Brazil. [4]

Atlas interpretation: Both searches contain similar country and travel words. Their relationship determines the answer. BERT's value was in representing that relationship, so a small word such as “to” could affect which pages matched the request. Google also showed remaining failures. The rollout demonstrated a practical use for contextual representations, not a search engine that understood every question. [4]

Sources

  1. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

    arXiv · Oct 11, 2018

  2. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

    arXiv · Oct 11, 2018

  3. Open Sourcing BERT: State-of-the-Art Pre-training for Natural Language Processing

    Google Research · Nov 2, 2018

  4. Understanding searches better than ever before

    Google · Oct 25, 2019