Skip to content
Model reviews

Benchmarks lie. Deployments don’t.

Everything I write about models — single-model reviews, performance notes from real workloads, and head-to-head matchups. Scored on what matters in production, updated as models ship.

The scoreboard
Model families compared on strengths, long-context handling, cost tier and a one-line verdict.
Model familyStrongest atLong contextCost tierMy one-line verdict
ClaudeComplex reasoning, code, agentic workExcellent$$$The default for high-stakes reasoning and agents.
GPTBreadth, ecosystem, multimodalVery good$$$Widest tooling surface; strong generalist.
GeminiNative multimodal, huge contextExcellent$$Best when video/context volume is the problem.
LlamaOpen weights, on-prem controlGood$The compliance-friendly self-host answer.
DeepSeek / MistralCost-efficient reasoningGood$Where the price-performance frontier lives.

Qualitative tiers from my own deployment evals, not vendor benchmarks. Snapshot: July 2026.

All reviews
Cover card: Kimi K3, 2.8T open weights and a 1 million token context window, set in white type on khetwal's vermilion red.

Review

2.8 Trillion Parameters, Free to Download: Inside Moonshot AI's Kimi K3

Kimi K3 packs 2.8 trillion parameters but activates only 104 billion per token — and Moonshot AI put the full weights on Hugging Face eleven days after announcing it.

Jul 29, 2026 · 10 min

3 reads

A stylized brain with glowing circuits and a question mark.

Matchup

Qualifying Claude Opus 4.8: From Shadowing to Go/No-Go Decision

Anthropic's Claude Opus 4.8 is out, but with no official benchmarks or release notes, upgrading is a gamble. This article provides a complete framework for safely qualifying the new model using shadow traffic, custom evaluation metrics, and a data-driven go/no-go decision.

Jun 04, 2026 · 12 min

2 reads

Abstract blue and white graphic with "Gemini 3.1 Pro" text.

Performance note

Gemini 3.1 Pro: Is a 1M Token Window Worth a Blind Upgrade?

Google's Gemini 3.1 Pro was released in February 2026 with a 1M token context window but few performance details. This article provides a production-focused framework for deciding whether to upgrade your AI stack to a new model when vendor benchmarks are missing.

Mar 19, 2026 · 13 min

1 reads

Claude AI logo with large text indicating 1M tokens and 90% recall.

Performance note

Claude's 1M Tokens & 90% Recall: What It Solves, What It Doesn't

Anthropic's Claude Opus 4.6 offers a 1M token context window, but its real value lies in its claimed 90% recall. This article explores when this massive context replaces RAG and when it's an expensive distraction.

Feb 24, 2026 · 11 min

1 reads

A computer screen displays lines of code with a red error message.

Performance note

Codex-Max and the $100+ Mistake 24 Hours In

GPT-5.1-Codex-Max can code autonomously for over 24 hours, but this power introduces new risks. This article explores the economics of long-horizon tasks and how to architect systems that prevent costly, deep-rooted errors.

Jan 13, 2026 · 9 min

3 reads

Two AI chatbots, Gemini 3 Pro and Claude 4.5, face off in a comparison.

Matchup

Opus 4.5 vs Gemini 3 Pro: Which Model for Which Workload

Released six days apart in November 2025, Google's Gemini 3 Pro and Anthropic's Claude Opus 4.5 are good at different things. A comparison on the axes that actually decide a stack: coding and agentic work, context handling, cost per unit of work, and provider concentration risk.

Dec 09, 2025 · 12 min

2 reads

Claude Opus 4.5 logo with a stylized brain and circuit board.

Review

Claude Opus 4.5 at $5/Mtok: When to Upgrade Your Agent's Brain

Anthropic's Claude Opus 4.5, released November 24, 2025, changes the economics of using frontier models. This architectural review explores how its $5/$25 per million token price point forces a redesign of model routing logic, especially for complex agentic workflows.

Nov 26, 2025 · 12 min

2 reads