Which 2018 Google language model introduced bidirectional pretraining using masked words?

The story behind the answer

Google’s 2018 language model that introduced bidirectional pretraining using masked words was BERT.

BERT stands for Bidirectional Encoder Representations from Transformers. Google researchers introduced it in a 2018 paper, and the model was designed to learn contextual representations from large collections of unlabeled text before being adapted to specific tasks.

During pretraining, BERT used masked-language modeling: some words were hidden, and the model learned to predict them from surrounding context. It also used a next-sentence prediction objective in its original training setup. Unlike left-to-right generation models, BERT’s encoder could use information from both directions at once.

BERT became influential in question answering, classification, named-entity recognition, and other language tasks. It is often confused with GPT because both use Transformer components, but BERT is primarily an encoder-style model for understanding text rather than a standard autoregressive text generator.

Source: Wikipedia · fact-checked Sept. 2026

Add question to a list

Choose a list to keep this question in: