Which 2018 language model introduced bidirectional Transformer pretraining and became known as BERT?

The story behind the answer

BERT was the 2018 language model that introduced bidirectional Transformer pretraining in a widely influential form.

BERT stands for Bidirectional Encoder Representations from Transformers. Google researchers trained it to use context from both the left and right sides of a word, allowing the model to build richer representations of language. Its pretraining included predicting masked words and learning relationships between sentence pairs.

After pretraining, BERT could be fine-tuned for tasks such as question answering, sentiment classification, and named-entity recognition. This separation between broad pretraining and task-specific fine-tuning became a standard pattern in natural-language processing.

BERT is sometimes confused with GPT because both use Transformer technology. BERT uses the Transformer encoder and was designed primarily to understand text representations, whereas the original GPT family used a decoder-style architecture aimed at generating text from preceding context. BERT also does not stand for “Bidirectional Encoding Representation Transformer”; its full expansion is the phrase above.

Source: Wikipedia · fact-checked Sept. 2026

Add question to a list

Choose a list to keep this question in: