The WireAI moves faster than it changes. Short notes on what landed this week, what it actually means, and a link to the source so you can disagree with me.
July 2026
Jul 23, 2026
ResearchPublication
The write-up covers the harness, the sampling temperature and the seeds — the three things that make a benchmark reproducible and the three things nobody publishes. Worth reading as a template for your own internal evals, not for the leaderboard position.
Jul 13, 2026
ResearchResearch
The task set overlapped with a public repository that appeared in pretraining data. Scores dropped several points on the cleaned set. A useful reminder that an agent benchmark is a software supply chain, not a measurement.
Every item links to its original source, which remains the work of its publisher. Headlines and summaries here are my own paraphrase and commentary, not reproductions.