GPT-1: How OpenAI's Pretraining and Fine-Tuning Recipe Worked

OpenAI pretrained a decoder-only Transformer on BooksCorpus, then fine-tuned one network across language tasks; the page explains the 2018 paper and results.

Pretrain on whatever text exists, then fine-tune

The paper's premise was that labeled data for any given task is scarce while unlabeled text is not. It proposed generative pretraining of a language model on unlabeled text, followed by discriminative fine-tuning on each target task. The resulting model needed only minimal architecture changes to move between tasks, and it improved on the prior state of the art in 9 of 12 tasks studied, including an 8.9 point gain on the Story Cloze Test, 5.7 points on RACE, and 1.5 points on MultiNLI. [2]

The model was a 12-layer decoder-only Transformer with 768-dimensional states and 12 attention heads, a smaller relative of the encoder-decoder architecture from the original Transformer paper. Secondary tallies put it at about 117 million parameters. It trained on BooksCorpus, described in the paper as containing over 7,000 unique unpublished books, chosen because its long, contiguous passages let the model learn from information spanning many sentences. [2][3]

What the pretraining stage actually learned

Pretraining used a standard language modeling objective: given a window of preceding tokens, predict the next one, repeated across the whole corpus. Fine-tuning then reused that same network, adding only a linear output layer and, for structured inputs like premise-hypothesis pairs, a way to serialize them into a single token sequence the pretrained model could read. [2]

Atlas interpretation: That is a narrower claim than it sounds. The paper reports gains on specific benchmarks measured against specific prior systems, not a general result about language understanding. Its own title page still carried the label "Preprint. Work in progress." [2]

The numbers behind the improvement

On the GLUE multi-task benchmark, the fine-tuned model scored 72.8 overall against a prior best of 68.9. The largest single jump was on CoLA, a test of whether a sentence is grammatical, where the score rose from 35.0 to 45.4. On the Story Cloze Test, a commonsense reasoning task, it reached 86.5 against a previous best of 77.6. [2]

Atlas interpretation: A 117-million-parameter model trained on one dataset of unpublished novels is small next to what followed. The result that mattered was not any single score but that one pretrained network, adapted with light fine-tuning, beat architectures built and tuned separately for each task. [2]

Two different sequels

BERT arrived four months later with a bidirectional encoder trained to fill in masked tokens rather than predict the next one, and it outperformed this model on the tasks the two papers shared. OpenAI's next move was to scale the same decoder-only recipe up rather than change it: GPT-2 used the same pretrain-then-adapt structure at far larger scale, and its release was staged rather than immediate. [2]

Sources

  1. Improving language understanding with unsupervised learning

    OpenAI · Jun 11, 2018

  2. Improving Language Understanding by Generative Pre-Training

    OpenAI · Sep 8, 2026

  3. GPT-1

    Wikipedia · Sep 8, 2026