Skip to content
The blog

Written after the deployment, not before.

Essays on AI architecture, strategy and the org problems in between. Heart what’s useful, argue with me in the comments.

Tagged llm-as-a-judgeclear ✕


Feb 10, 2026

llm-as-a-judge13 min read

Your LLM Judge's Secret Drift: Why 80% Agreement Isn't Enough

LLM-as-a-judge provides scalable evaluation but suffers from silent drift as models and rubrics change. This guide explains how to use a human-labeled holdout set to calibrate your judge, diagnose regressions, and prevent your evaluation system from misleading you.

Abstract scales with a judge icon and a downward trending graph.

1 reads