GPT-3: 175B Parameters, Few-Shot Prompts, and API Access

OpenAI’s GPT-3 performed tasks from instructions and examples without weight updates, while private API access followed the paper in June 2020.

Examples became part of the input

OpenAI's May 28, 2020 paper evaluated GPT-3 on different tasks while keeping its 175 billion parameters fixed. Instructions and examples went into the input; the model generated a continuation. [2]

Atlas interpretation: An illustrative sentiment prompt might contain “A delightful film → positive” and “A tedious film → negative,” followed by a new review and an arrow. The examples specify both the task and the expected answer format. They condition this completion; they do not permanently train a new classifier. A later request needs its own context. [2]

The paper compared zero-shot evaluation with no demonstrations, one-shot with one, and few-shot with several. Prompt content was part of the experimental setup. [2]

What changed from the first GPT

The first GPT had already separated learning from unlabeled text from fine-tuning on a labeled task. That let different applications reuse a trained language model instead of building a separate language-learning system for each dataset. [3]

Atlas interpretation: The 2018 work also explored performing tasks directly with the underlying language model, although results were often weaker than supervised alternatives. GPT-3 therefore developed an existing research direction: make the pretrained model useful through examples at inference time. It did not invent the idea that language models could do more than complete prose. [3][2]

A useful interface with uneven results

Results varied by task. The authors identified weaknesses in reading comprehension and sentence comparison, repetitive text, and benchmark material appearing in training data. Examples did not guarantee reliable performance. [2]

The June 11 OpenAI API launch made GPT-3-family models accessible through a private beta. Developers sent text and received a completion, rather than downloading weights. OpenAI controlled admission and operated the models; the May research paper and the June service launch were different events. [4]

Atlas interpretation: For an application developer, this moved some experimentation into the prompt: change the examples or wording, then inspect the results. It reduced the need to run a separate training job for every experiment. Evaluation still mattered, and using a hosted API introduced a separate dependency on the provider's access and operating decisions. [4]

Sources

  1. Language Models are Few-Shot Learners

    arXiv · May 28, 2020

  2. Language Models are Few-Shot Learners

    arXiv · May 28, 2020

  3. Improving language understanding with unsupervised learning

    OpenAI · Jun 11, 2018

  4. OpenAI API

    OpenAI · Jun 11, 2020