The Transformer neural-network architecture was introduced in the 2017 paper “Attention Is All You Need.”
The paper was written by researchers at Google and the University of Toronto. It proposed using attention mechanisms to relate words or other tokens to one another, rather than relying primarily on recurrent processing. This made it easier to process sequences in parallel during training and to model relationships across long distances.
Transformers were first presented for machine translation, but the architecture soon spread to language modeling, image processing, speech, and multimodal systems. Large language models such as BERT and GPT use Transformer-based designs, though they differ in training objectives and detailed structure.
The name can cause confusion because “attention” is a mechanism, while the Transformer is the complete architecture described in the paper. Transformers also do not inherently understand language like a person; they learn statistical patterns from training data and can still produce errors or unsupported claims.