A visual, step-by-step guide to the paper that introduced BERT — the deep bidirectionally pre-trained Transformer that reached state-of-the-art on 11 NLP tasks with a single model and one extra output layer.
BERT arrived at the end of a five-year sprint in transfer learning for NLP. Here is the road that led there.
Language understanding is deeply bidirectional. To predict the masked word in "I went to the ___ to deposit money", you need both the left context ("went to the") and the right context ("to deposit money"). Models that only read left-to-right are solving the task half-blind.
Before BERT, the best pre-trained language models read text strictly left-to-right — or stitched together shallow halves. Both approaches crippled context.
BERT corrupts the input by masking words, then learns to reconstruct them. It trains on unlabeled text — a task that generates itself, at any scale.
A cloze test, at scale. BERT sees both directions for every masked position, so every layer of the encoder can exchange information across the whole sentence. MLM turns the entire web into a supervised dataset for free.
Is sentence B the actual next sentence after A, or a random one? 50/50 construction. NSP teaches sentence-pair relationships for QA and natural language inference. Later research (RoBERTa) showed its value is limited — but it was part of the original recipe.
The Transformer encoder from "Attention Is All You Need", stripped of the decoder and scaled into two sizes: Base and Large.
Every token is the sum of three embeddings. The [CLS] token's final hidden state becomes the sentence (or sentence-pair) representation for classification tasks; token-level outputs feed span tasks like QA. That's it — the same encoder, no task-specific architecture.
Two objectives, two corpora, no labels. Everything BERT "knows" about language comes from this single unlabeled pass.
The 80/10/10 split keeps the model from believing every position could be masked, and keeps representations sharp on unmasked tokens. Fine-tuning inputs never contain [MASK] — this mismatch is the cost of deep bidirectionality.
Swap in a single linear layer, fine-tune for a few hours on a TPU, and beat every task-specific architecture of the era.
| System | Score | Task |
|---|---|---|
| Previous best (ensemble era) | 72.8 | GLUE benchmark |
| OpenAI GPT | 72.8 | GLUE benchmark |
| BERT-Base | 79.6 | GLUE benchmark |
| BERT-Large | 80.5 | GLUE benchmark (+7.7 points) |
| ELMo / BiLSTM era | ~85.8 F1 | SQuAD 1.1 (QA) |
| BERT-Large (single) | 90.9 F1 | SQuAD 1.1 (QA) |
| BERT-Large (ensemble) | 93.2 F1 | SQuAD 1.1 — human level |
| BERT-Large | 86.3 | SWAG (commonsense) |
| BERT-Large | 92.8 F1 | CoNLL-2003 NER |
BERT reached a new state of the art on all 11 tasks evaluated in the paper — sentence-level and token-level, with the same pre-trained weights.
BERT reframed NLP: pre-train bidirectionally once, fine-tune everywhere. An entire research literature grew from it.
Check your understanding of the key concepts from the BERT paper.
Everything you need to remember about this paper.