History Problem Core Idea Architecture Pre-Training Results Impact Quiz
Interactive Paper Explainer

Bidirectional Transformers
BERT

A visual, step-by-step guide to the paper that introduced BERT — the deep bidirectionally pre-trained Transformer that reached state-of-the-art on 11 NLP tasks with a single model and one extra output layer.

Start Learning Read the Paper ↗
340M
Parameters (Large)
11
NLP Tasks Conquered
3.3B
Pre-Training Words
2018
Year Published
History

From Word Vectors to BERT

BERT arrived at the end of a five-year sprint in transfer learning for NLP. Here is the road that led there.

2013
word2vec (Mikolov et al.)
Static word embeddings — one vector per word, learned with shallow networks. "bank" gets one meaning.
2017
Transformer (Vaswani et al.)
Attention-only architecture. Fast, parallel, and expressive — but used as an encoder-decoder for translation.
2018 · Feb
ELMo (Peters et al.)
Contextual embeddings from a bidirectional LSTM. "bank" finally changed meaning with the sentence.
2018 · Jun
GPT-1 (Radford et al.)
Generative pre-training on a left-to-right Transformer decoder + task-specific fine-tuning. Unidirectional.
2018 · Oct
🚀 BERT (Devlin et al.)
Pre-train a Transformer encoder with masked language modeling. Deep bidirectionality at scale.
2019 →
RoBERTa, ALBERT, ELECTRA…
The "BERTology" era: better training recipes, smaller and faster variants, and BERT in Google Search.
Key Insight

Language understanding is deeply bidirectional. To predict the masked word in "I went to the ___ to deposit money", you need both the left context ("went to the") and the right context ("to deposit money"). Models that only read left-to-right are solving the task half-blind.

CONTEXT NEEDED FOR THE MASK
I went to the [MASK] to deposit money
Blue = left context · Teal = right context · BERT uses both.
Chapter 01

The Problem with One-Way Reading

Before BERT, the best pre-trained language models read text strictly left-to-right — or stitched together shallow halves. Both approaches crippled context.

⬅️
Unidirectional Pre-Training
  • GPT-style models condition only on previous tokens
  • Masked words can never use future context
  • Left-to-right "right-context" predictions are weak
  • Fine-tuning is suboptimal for token-level tasks like QA
  • One direction = half the information, always
↔️
BERT's Solution
  • Mask random tokens, then predict them from BOTH sides
  • Deep Transformer encoder, fully bidirectional
  • One pre-trained model, one added layer per task
  • State-of-the-art on 11 tasks with minimal task engineering
  • Pre-training learns from unlabeled text at internet scale
Interactive Demo — Why Direction Matters

The word "bank" is ambiguous. Toggle what a model is allowed to see and watch the prediction flip.

Chapter 02

The Core Idea — Masked Language Modeling

BERT corrupts the input by masking words, then learns to reconstruct them. It trains on unlabeled text — a task that generates itself, at any scale.

Input: "The capital of France is [MASK]."
Target: "[MASK] → Paris"
[MASK]
Masked Token
15% of tokens are selected for corruption in each sequence.
80%
→ [MASK]
80% of selected tokens are replaced with the mask token.
10%
→ random
10% are replaced with a random vocabulary word.
10%
→ keep
10% stay unchanged — the model must stay calibrated on real text.
Task 1 — Masked LM (MLM)

A cloze test, at scale. BERT sees both directions for every masked position, so every layer of the encoder can exchange information across the whole sentence. MLM turns the entire web into a supervised dataset for free.

Task 2 — Next Sentence Prediction (NSP)

Is sentence B the actual next sentence after A, or a random one? 50/50 construction. NSP teaches sentence-pair relationships for QA and natural language inference. Later research (RoBERTa) showed its value is limited — but it was part of the original recipe.

Interactive Demo — Masked Token Prediction

Pick a sentence, run the "prediction", and see the candidate distribution a bidirectional model would produce.

Chapter 03

BERT Architecture

The Transformer encoder from "Attention Is All You Need", stripped of the decoder and scaled into two sizes: Base and Large.

Architecture Overview
🧱
Encoder-Only
Full self-attention, no causal mask, no decoder.
⬛
[CLS] + [SEP]
Special tokens for classification and sentence pairs.
📉
WordPiece
30,000-token subword vocabulary from GPT-style BPE era.
📏
512 Tokens
Fixed maximum input length, position + segment embeddings.
⚡
GELU
Gaussian Error Linear Unit activations, layer norm everywhere.
BERT-BASE
110M
parameters
12 layers · 768 hidden · 12 heads
BERT-LARGE
340M
parameters
24 layers · 1024 hidden · 16 heads
GPT-1 (reference)
117M
parameters
12 layers · 768 hidden · 12 heads (decoder)
Input Representation
Input = TokenEmbedding + PositionEmbedding + SegmentEmbedding
[CLS] the capital of france is [MASK] [SEP] my cat is cute [SEP]
E    E E E E E E E E E   E        E     E E E E E E   E    — segment A = 0s, segment B = 1s

Every token is the sum of three embeddings. The [CLS] token's final hidden state becomes the sentence (or sentence-pair) representation for classification tasks; token-level outputs feed span tasks like QA. That's it — the same encoder, no task-specific architecture.

Chapter 04

Pre-Training at Scale

Two objectives, two corpora, no labels. Everything BERT "knows" about language comes from this single unlabeled pass.

📚 BooksCorpus
800M words of clean, continuous narrative text — 11,038 unpublished books.
📖 English Wikipedia
2,500M words — lists, tables, and headers stripped, prose kept.
🔢 ~3.3B words total
Roughly 40 epochs over the corpus during training — about 131B tokens seen.
⚙️ 4 TPU days (Base)
1M steps, batch 256 sequences of 512 tokens, Adam 1e-4, warmup 10K. Large: 4 days on 16 TPUs.
The 80 / 10 / 10 Rule

The 80/10/10 split keeps the model from believing every position could be masked, and keeps representations sharp on unmasked tokens. Fine-tuning inputs never contain [MASK] — this mismatch is the cost of deep bidirectionality.

Ablations — What Actually Mattered
  • NSP on/off: removing NSP hurt QQP, MNLI, and SQuAD scores modestly (RoBERTa later removed it entirely with more data — and improved).
  • Longer training: more steps, more data, longer sequences consistently helped — gains had not saturated.
  • Model size: BERT-Large beat Base on every single task; capacity was still the binding constraint.
  • MLM vs LTR: left-to-right-only pre-training was far worse on token-level tasks (SQuAD), and even hurt sentence-level tasks.
Chapter 05

Fine-Tuning — One Model, Eleven Tasks

Swap in a single linear layer, fine-tune for a few hours on a TPU, and beat every task-specific architecture of the era.

GLUE & Benchmark Results (single models unless noted)
SystemScoreTask
Previous best (ensemble era)72.8GLUE benchmark
OpenAI GPT72.8GLUE benchmark
BERT-Base79.6GLUE benchmark
BERT-Large80.5GLUE benchmark (+7.7 points)
ELMo / BiLSTM era~85.8 F1SQuAD 1.1 (QA)
BERT-Large (single)90.9 F1SQuAD 1.1 (QA)
BERT-Large (ensemble)93.2 F1SQuAD 1.1 — human level
BERT-Large86.3SWAG (commonsense)
BERT-Large92.8 F1CoNLL-2003 NER

BERT reached a new state of the art on all 11 tasks evaluated in the paper — sentence-level and token-level, with the same pre-trained weights.

FINE-TUNING COST
hours
on a single TPU vs months of task engineering
TASK-SPECIFIC PARAMS
1
linear output layer per task
STEPS TO BEAT SOTA
2–4
epochs of fine-tuning, typically
MNLI ACCURACY
86.7
BERT-Large, vs 82.1 for previous best
Legacy

Impact — BERTology

BERT reframed NLP: pre-train bidirectionally once, fine-tune everywhere. An entire research literature grew from it.

🚀 RoBERTa (2019)
Same architecture, better recipe: more data, longer training, dynamic masking, no NSP.
🔍 ALBERT (2019)
Parameter sharing across layers cut parameters drastically at similar quality.
⚡ DistilBERT (2019)
Knowledge distillation made BERT 40% smaller, 60% faster, retaining ~97% quality.
🎯 ELECTRA (2020)
Replaced masking with "replaced-token detection" — learn from every position, not 15%.
🌐 Google Search (2019)
BERT was deployed in Google Search — reportedly touching ~1 in 10 English US queries at launch.
🖼️ Beyond text
The masked-modeling recipe jumped to vision (ViT, MAE), audio, and protein sequences.
Test Yourself

Quick Quiz

Check your understanding of the key concepts from the BERT paper.

Reference

Key Takeaways

Everything you need to remember about this paper.

✅ Deep bidirectionality: masked LM lets every layer see both left and right context.
✅ Pre-training = masked LM + next sentence prediction on 3.3B words of unlabeled text.
✅ One Transformer encoder, two sizes: Base (110M) and Large (340M).
✅ Fine-tuning adds a single output layer — hours of TPU time per task.
✅ State of the art on all 11 tasks evaluated, including GLUE 80.5 and SQuAD 93.2 (F1, ensemble).
✅ The 80/10/10 masking rule, [CLS] token, and segment embeddings became standard vocabulary.