Tutorials › Understanding Transformers
Previous ← Colab Agent Practice
Next Build Your Own Transformer →

Understanding Transformers

A 16-part intuition-first tour of how GPT-style models work — from prediction to next-token generation, then to the hardware it runs on. Pair it with the companion AI Core Math Review series if the underlying math is rusty, and Build Your Own Transformer for a hands-on notebook.


Lessons
1 Learning by Prediction 2 How Learning Happens 3 Neurons, Weights, Bias, and Activations 4 Depth, Residuals, and Normalization 5 Tokens, Embeddings, and Position 6 Why Attention Was Needed 7 Self-Attention, Step by Step 8 Multi-Head Attention 9 Logits and Softmax 10 Feedforward Networks 11 How X, Q, K, V Evolve Layer by Layer 12 The Complete Transformer Block 13 KV Caching 14 How GPT Trains and Generates 15 End-to-End Walkthrough and Cheat Sheet 16 Hardware Accelerators: GPUs, TPUs, and JAX
Back matter
§ References
Previous ← Colab Agent Practice
Next Build Your Own Transformer →