What is the name of the 2017 Google research paper that introduced the Transformer architecture?

The story behind the answer

The 2017 Google research paper that introduced the Transformer architecture is titled “Attention Is All You Need.”

The paper was written by Ashish Vaswani and colleagues and presented a neural-network architecture based entirely on attention mechanisms. Unlike earlier sequence models, it did not rely on recurrent or convolutional layers to process tokens in order. This made training more parallelizable and helped the model capture relationships between distant words.

The Transformer was initially demonstrated for machine translation, especially English-to-German and English-to-French tasks. Its encoder–decoder design later became the foundation for many influential systems, including BERT, GPT-style language models, and modern large language models.

A common mix-up is to credit the architecture to BERT or GPT. Those are later model families that use Transformer designs; “Attention Is All You Need” is the original paper that introduced the architecture itself.

Source: Wikipedia · fact-checked Sept. 2026

Add question to a list

Choose a list to keep this question in: