A translation model before it was a language-model blueprint
Attention Is All You Need appeared on June 12, 2017. The Transformer paired an encoder that represented an input sentence with a decoder that generated its translation. It removed recurrent and convolutional layers, using attention to exchange information across positions instead. [2]
The authors reported 28.4 BLEU on WMT 2014 English-to-German translation, more than two points above the best earlier result in their comparison. Their large model trained for 3.5 days on eight NVIDIA P100 GPUs. This was evidence from a specific translation benchmark, not a measurement of general intelligence or a universal training-speed record. [2]
What attention actually computes
A word's meaning depends on its neighbors, sometimes far away. In Google's explanation, whether “bank” means a financial institution or a riverbank depends on the sentence's ending. Self-attention compares a word's numerical representation with other words' representations, assigns weights, and combines their information. The model can connect “bank” to “river” directly rather than passing that information through every intervening word. [3]
Multiple attention heads learn different ways to make those comparisons. The network also contains feed-forward layers, and positional encodings supply word-order information. The title describes the replacement for recurrence, not a network containing literally nothing except attention. [2]
Parallel training does not mean instant answers
Atlas interpretation: Recurrent networks carry information through a sequence of dependent steps. Attention lets many positions be processed together, a better match for GPUs. That makes training easier to parallelize. Generating a translation still proceeds one output at a time, with each new prediction using the preceding outputs. Faster training and faster completion of an entire answer are different engineering questions. [3]
Full self-attention also compares pairs of positions. Doubling the sequence length quadruples the number of pairs. The paper's quadratic sequence-length cost helps explain why long inputs remain expensive even when the computation is parallel. [2]
GPT and BERT took different parts of the design
The first GPT used a Transformer decoder to learn from ordinary text by predicting the next token, a word or piece of a word. OpenAI then fine-tuned it on labeled tasks. Pretraining supplied a reusable starting point, so each task did not require learning language patterns from its small labeled dataset alone. [4]
BERT used a bidirectional Transformer encoder. Its masked-language-model task hid some input tokens and trained the model to recover them using context from both sides. That suited learning representations for tasks such as extracting an answer from a passage. GPT's left-to-right prediction and BERT's masked-token reconstruction were different training objectives built on related machinery. [5][4]
Atlas interpretation: The two projects turned a translation architecture into reusable language models. By changing which context was visible and what the model learned to predict, researchers could train on unlabeled text and then adapt those models to different tasks. [4][5]
Sources
- Attention Is All You Need
arXiv · Jun 12, 2017
- Attention Is All You Need
arXiv · Jun 12, 2017
- Transformer: A Novel Neural Network Architecture for Language Understanding
Google Research · Aug 31, 2017
- Improving Language Understanding by Generative Pre-Training
OpenAI · Sep 8, 2026
- BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
arXiv · Oct 11, 2018