Skip to content
The blog

Written after the deployment, not before.

Essays on AI architecture, strategy and the org problems in between. Heart what’s useful, argue with me in the comments.

Tagged llm-evaluationclear ✕


Feb 10, 2026

llm-as-a-judge13 min read

Your LLM Judge's Secret Drift: Why 80% Agreement Isn't Enough

LLM-as-a-judge provides scalable evaluation but suffers from silent drift as models and rubrics change. This guide explains how to use a human-labeled holdout set to calibrate your judge, diagnose regressions, and prevent your evaluation system from misleading you.

Abstract scales with a judge icon and a downward trending graph.

1 reads

Oct 29, 2025

llm evaluation10 min read

Your First Eval: Build a 30-Example Test Set This Afternoon

Stop building elaborate evaluation frameworks and start measuring model quality today. Learn how to build a small, effective 30-example evaluation set from real traffic in a single afternoon.

A person's hands typing on a laptop with a graph on the screen.

2 reads