01
Alignment2023
Direct Preference Optimization: Your Language Model is Secretly a Reward Model
Rafailov et al.
Collapsed the two-stage RLHF pipeline into a single loss, with no reward model to train.
The white papers that actually moved the field, each linked to the original and paired with my plain-language summary, a glossary of the terms, and what it changes in practice.
01
Rafailov et al.
Collapsed the two-stage RLHF pipeline into a single loss, with no reward model to train.
02
Ouyang et al.
Turned a next-token predictor into an assistant, and found a far smaller aligned model was preferred over a much larger unaligned one.
Summary
All original papers remain the work of their authors and link to the official source (arXiv or publisher). Summaries and glossaries on this site are my own commentary.