01
Scaling2022
Training Compute-Optimal Large Language Models
Hoffmann et al.
The field had been building models too large and training them on too little data.
The white papers that actually moved the field, each linked to the original and paired with my plain-language summary, a glossary of the terms, and what it changes in practice.
01
Hoffmann et al.
The field had been building models too large and training them on too little data.
All original papers remain the work of their authors and link to the official source (arXiv or publisher). Summaries and glossaries on this site are my own commentary.