Skip to content
Model notes

Benchmarks lie. Deployments don’t.

Everything I write about models — single-model reviews, performance notes from real workloads, and head-to-head matchups. Scored on what matters in production, updated as models ship.

Model families compared on strengths, long-context handling, cost tier and a one-line verdict.
Model familyStrongest atLong contextCost tierMy one-line verdict
ClaudeComplex reasoning, code, agentic workExcellent$$$The default for high-stakes reasoning and agents.
GPTBreadth, ecosystem, multimodalVery good$$$Widest tooling surface; strong generalist.
GeminiNative multimodal, huge contextExcellent$$Best when video/context volume is the problem.
LlamaOpen weights, on-prem controlGood$The compliance-friendly self-host answer.
DeepSeek / MistralCost-efficient reasoningGood$Where the price-performance frontier lives.

Qualitative tiers from my own deployment evals, not vendor benchmarks. Snapshot: July 2026.

Recent notes