OpenAI created the GPT-1 language model released in 2018. GPT-1 introduced the “generative pre-trained transformer” approach, in which a language model first learns broad statistical patterns from large text collections and is then adapted for specific tasks.
The model contained 117 million parameters and used the Transformer architecture introduced by Google researchers in 2017. OpenAI’s 2018 paper showed that unsupervised pretraining could improve performance on several natural-language-processing tasks after supervised fine-tuning.
GPT-1 is sometimes confused with BERT, another influential 2018 language model created at Google. GPT-1 generated text autoregressively by predicting the next token, whereas BERT was designed around bidirectional representations and masked-word prediction. GPT-1 was much smaller than later GPT-2 and GPT-3 systems, but it established the naming and pretraining line that later became widely known.