Google’s 2018 language model that introduced bidirectional pretraining using masked words was BERT.
BERT stands for Bidirectional Encoder Representations from Transformers. Google researchers introduced it in a 2018 paper, and the model was designed to learn contextual representations from large collections of unlabeled text before being adapted to specific tasks.
During pretraining, BERT used masked-language modeling: some words were hidden, and the model learned to predict them from surrounding context. It also used a next-sentence prediction objective in its original training setup. Unlike left-to-right generation models, BERT’s encoder could use information from both directions at once.
BERT became influential in question answering, classification, named-entity recognition, and other language tasks. It is often confused with GPT because both use Transformer components, but BERT is primarily an encoder-style model for understanding text rather than a standard autoregressive text generator.