Which 2019 paper introduced the BERT language model at Google?
Answer
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Answer
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
The paper that introduced Google’s BERT language model was “BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.”
Jacob Devlin and colleagues published the work in 2018, and the paper is commonly associated with Google’s BERT release that year rather than 2019. BERT stands for Bidirectional Encoder Representations from Transformers. It showed how a transformer encoder could be pretrained on large text collections and then fine-tuned for many language tasks.
A central technique was masked-language modeling: selected words were hidden, and the model learned to predict them using context from both directions. BERT also used next-sentence prediction in its original pretraining setup. These methods helped it perform strongly on question answering, sentence classification, and other benchmarks.
BERT influenced search, language understanding, and later transformer systems. It is an encoder model, so it differs from autoregressive text generators that predict the next token from left to right. That distinction is a common source of confusion when comparing BERT with generative language models.
Source: Wikipedia · fact-checked Sept. 2026