Previous
← Colab Agent Practice
Understanding Transformers
A 16-part intuition-first tour of how GPT-style models work — from prediction to next-token generation, then to the hardware it runs on. Pair it with the companion AI Core Math Review series if the underlying math is rusty, and Build Your Own Transformer for a hands-on notebook.
Lessons
1
Learning by Prediction
2
How Learning Happens
3
Neurons, Weights, Bias, and Activations
4
Depth, Residuals, and Normalization
5
Tokens, Embeddings, and Position
6
Why Attention Was Needed
7
Self-Attention, Step by Step
8
Multi-Head Attention
9
Logits and Softmax
10
Feedforward Networks
11
How X, Q, K, V Evolve Layer by Layer
12
The Complete Transformer Block
13
KV Caching
14
How GPT Trains and Generates
15
End-to-End Walkthrough and Cheat Sheet
16
Hardware Accelerators: GPUs, TPUs, and JAX
Back matter
§
References
Previous
← Colab Agent Practice