Skip to content
The paper library

Read the papers. Skip the jargon.

The white papers that actually moved the field, each linked to the original and paired with my plain-language summary, a glossary of the terms, and what it changes in practice.

AllArchitectureScalingReasoningAlignment

01

Architecture2017

Attention Is All You Need

Vaswani, Shazeer, Parmar, et al. — Google Brain

The transformer paper. My summary walks the attention mechanism with a spreadsheet analogy, and the glossary decodes “multi-head”, “positional encoding” and friends.

02

Scaling2020

Scaling Laws for Neural Language Models

Kaplan, McCandlish, et al. — OpenAI

Why bigger kept winning — until it didn’t. Essential background for any budget conversation about training vs. buying.

03

Scaling2022

Training Compute-Optimal Large Language Models

Hoffmann, Borgeaud, et al. — DeepMind

The “Chinchilla” correction: most large models were undertrained on data. Changed how every lab spends its compute.

Original paper ↗

Summary · Glossary

04

Reasoning2022

Chain-of-Thought Prompting Elicits Reasoning in Large Language Models

Wei, Wang, et al. — Google Research

The paper behind “think step by step”. Where CoT genuinely helps, where it just burns tokens, and what it means for eval design.

05

Alignment2022

Constitutional AI: Harmlessness from AI Feedback

Bai, Kadavath, et al. — Anthropic

Alignment without an army of human labelers. Worth reading for the governance pattern alone.

Original paper ↗

Summary · Glossary

06

Alignment2023

Direct Preference Optimization

Rafailov, Sharma, et al. — Stanford

RLHF without the RL. The math is dense; the summary isn’t. If you fine-tune in production, this is the one to understand.

All original papers remain the work of their authors and link to the official source (arXiv or publisher). Summaries and glossaries on this site are my own commentary.