Understanding Transformers
A 16-part intuition-first tour of how GPT-style models work — from basic math to next-token generation.
Articles
1
The Big Idea: Learning by Prediction
2
The Math You Actually Need
optional
3
How Learning Happens
optional
4
Neurons, Weights, Bias, and Activations
5
Depth, Residuals, and Normalization
6
Tokens, Embeddings, and Position
7
Why Attention Was Needed
8
Self-Attention, Step by Step
9
Multi-Head Attention
10
Logits and Softmax
11
Feedforward Networks
12
How X, Q, K, V Evolve Layer by Layer
13
The Complete Transformer Block
14
KV Caching
15
How GPT Trains and Generates
16
End-to-End Walkthrough and Cheat Sheet
Back matter
§
References