Skip to content

What are you curious about?

One search across the wire, the paper library, model reviews, the blog and off hours.

TryagentsRAGcontext windowfine-tuning


Aug 25, 2026

Debate and Decompose: When a Second Agent Is Worth the Cost

Multi-agent systems add cost and complexity, but are justified for specific use cases like independent verification (debate) or breaking down tasks that exceed a single model's context window (decomposition). This article provides a framework for deciding when to add a second agent.

Two abstract figures debate, one decomposing into smaller parts.

11 min

Aug 25, 2026

The Model Upgrade Is Not a Drop-In: A Production Migration Guide

A major new foundation model release is not a simple drop-in upgrade. This guide provides a framework for migrating your production pipeline safely, covering re-qualification, prompt auditing, and sequenced rollouts to avoid breaking changes.

Abstract diagram showing a complex pipeline with interconnected nodes and arrows.

11 min

The paper libraryAll papers

Jun 18, 2026

FlashAttention: How a Memory Trick Unlocked Today's Giant AI Models

The attention mechanism in AI models was once limited by a memory bottleneck that scaled quadratically with input length. We explain FlashAttention, the IO-aware algorithm that solved this by reorganizing the calculation, enabling the massive context windows now common in large language models.

Abstract visualization of interconnected nodes and data flow.

13 min

1 reads

May 22, 2026

Your LLM Is a Reward Model: How DPO Cut Alignment Costs

Direct Preference Optimization (DPO) simplifies the complex RLHF pipeline by eliminating the need for a separate reward model. This article explains how DPO works, why it made alignment more accessible, and the trade-offs involved in this more direct approach.

Abstract diagram showing a simplified alignment process with DPO.

12 min

1 reads

Model reviewsAll reviews

Jun 04, 2026

Qualifying Claude Opus 4.8: From Shadowing to Go/No-Go Decision

Anthropic's Claude Opus 4.8 is out, but with no official benchmarks or release notes, upgrading is a gamble. This article provides a complete framework for safely qualifying the new model using shadow traffic, custom evaluation metrics, and a data-driven go/no-go decision.

A stylized brain with glowing circuits and a question mark.

12 min

2 reads

Mar 19, 2026

Gemini 3.1 Pro: Is a 1M Token Window Worth a Blind Upgrade?

Google's Gemini 3.1 Pro was released in February 2026 with a 1M token context window but few performance details. This article provides a production-focused framework for deciding whether to upgrade your AI stack to a new model when vendor benchmarks are missing.

Abstract blue and white graphic with "Gemini 3.1 Pro" text.

13 min

1 reads


Solution architect · technology enthusiast · lifelong learner

Field notes on AI, architecture and strategy.

I'm Sunder. I help teams work out where AI fits, adopt it for real, and architect the whole path from ideation to production. Beyond the work, I'm happiest breaking dense white papers down until they read plainly, comparing how the newest models really hold up, and writing about what I find — with a bit of photography and the odd drawing when there's time.

Portrait of Sunder K

What lives here

01

The Wire

What actually happened in AI this week — short, dated, sourced, and linked to whoever reported it first.

02

The paper library

Foundational and current white papers, each with a plain-language summary, a glossary of the jargon, and my take on what it changes in practice.

03

Model reviews

Everything I write about models: single-model reviews, hands-on notes, and the occasional head-to-head matchup.

04

The blog

Essays on AI — architecture, models, and whatever's got my attention that week. Readers can react and comment.

05

Off hours

Photographs and drawings. Proof that not everything needs a GPU.

Working through an AI decision? Happy to talk it through.