01
Attention Is All You Need
Vaswani, Shazeer, Parmar, et al. — Google Brain
The transformer paper. My summary walks the attention mechanism with a spreadsheet analogy, and the glossary decodes “multi-head”, “positional encoding” and friends.
The white papers that actually moved the field, each linked to the original and paired with my plain-language summary, a glossary of the terms, and what it changes in practice.
01
Vaswani, Shazeer, Parmar, et al. — Google Brain
The transformer paper. My summary walks the attention mechanism with a spreadsheet analogy, and the glossary decodes “multi-head”, “positional encoding” and friends.
02
Kaplan, McCandlish, et al. — OpenAI
Why bigger kept winning — until it didn’t. Essential background for any budget conversation about training vs. buying.
Summary
03
Hoffmann, Borgeaud, et al. — DeepMind
The “Chinchilla” correction: most large models were undertrained on data. Changed how every lab spends its compute.
Summary · Glossary
04
Wei, Wang, et al. — Google Research
The paper behind “think step by step”. Where CoT genuinely helps, where it just burns tokens, and what it means for eval design.
Summary
05
Bai, Kadavath, et al. — Anthropic
Alignment without an army of human labelers. Worth reading for the governance pattern alone.
Summary · Glossary
06
Rafailov, Sharma, et al. — Stanford
RLHF without the RL. The math is dense; the summary isn’t. If you fine-tune in production, this is the one to understand.
Summary
All original papers remain the work of their authors and link to the official source (arXiv or publisher). Summaries and glossaries on this site are my own commentary.