TutorialsUnderstanding Transformers › References

Understanding Transformers

References

Works cited across the series. In-text citations use author and year; the part numbers after each entry show where it appears.

Ainslie, J., Lee-Thorp, J., de Jong, M., Zemlyanskiy, Y., Lebrón, F., & Sanghai, S. (2023). GQA: Training generalized multi-query transformer models from multi-head checkpoints. EMNLP. arXiv:2305.13245 — Part 14

Ba, J. L., Kiros, J. R., & Hinton, G. E. (2016). Layer normalization. arXiv:1607.06450 — Part 5

Bahdanau, D., Cho, K., & Bengio, Y. (2015). Neural machine translation by jointly learning to align and translate. ICLR. arXiv:1409.0473 — Part 7

He, K., Zhang, X., Ren, S., & Sun, J. (2016). Deep residual learning for image recognition. CVPR. arXiv:1512.03385 — Part 5

Hendrycks, D., & Gimpel, K. (2016). Gaussian error linear units (GELUs). arXiv:1606.08415 — Part 4

Holtzman, A., Buys, J., Du, L., Forbes, M., & Choi, Y. (2020). The curious case of neural text degeneration. ICLR. arXiv:1904.09751 — Part 15

Kingma, D. P., & Ba, J. (2015). Adam: A method for stochastic optimization. ICLR. arXiv:1412.6980 — Part 3

Meng, K., Bau, D., Andonian, A., & Belinkov, Y. (2022). Locating and editing factual associations in GPT. NeurIPS. arXiv:2202.05262 — Part 1

Michel, P., Levy, O., & Neubig, G. (2019). Are sixteen heads really better than one? NeurIPS. arXiv:1905.10650 — Part 9

Press, O., & Wolf, L. (2017). Using the output embedding to improve language models. EACL. arXiv:1608.05859 — Part 10

Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., & Sutskever, I. (2019). Language models are unsupervised multitask learners. OpenAI. — Parts 1, 6

Rumelhart, D. E., Hinton, G. E., & Williams, R. J. (1986). Learning representations by back-propagating errors. Nature, 323, 533–536. — Part 3

Sennrich, R., Haddow, B., & Birch, A. (2016). Neural machine translation of rare words with subword units. ACL. arXiv:1508.07909 — Part 6

Su, J., Lu, Y., Pan, S., Murtadha, A., Wen, B., & Liu, Y. (2021). RoFormer: Enhanced transformer with rotary position embedding. arXiv:2104.09864 — Part 6

Touvron, H., Lavril, T., Izacard, G., et al. (2023). LLaMA: Open and efficient foundation language models. arXiv:2302.13971 — Parts 3, 5, 6

Uszkoreit, J. (2017). Transformer: A novel neural network architecture for language understanding. Google Research Blog. — Part 7

Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., & Polosukhin, I. (2017). Attention is all you need. NeurIPS. arXiv:1706.03762 — Parts 6–11, 13

Xiong, R., Yang, Y., He, D., et al. (2020). On layer normalization in the transformer architecture. ICML. arXiv:2002.04745 — Parts 13, 16

Zhang, B., & Sennrich, R. (2019). Root mean square layer normalization. NeurIPS. arXiv:1910.07467 — Part 5