The evaluation crisis nobody wants to own — that, at least, is the claim. AI Ledger spent the week testing it against the evidence, talking to the people closest to the work, and separating what has shipped from what has merely been announced.

The story matters because the stakes are no longer theoretical. Decisions taken now in research will compound for years, shaping who builds, who pays, and who is held to account. The temptation is to treat each new release as a verdict. The discipline is to treat it as a data point.

What the demonstrations show is real progress. What they obscure is the gap between a controlled setting and a deployment that has to survive contact with messy, adversarial reality. That gap is where most of the value — and most of the risk — actually lives.

Follow the incentives and the picture sharpens. Capital is abundant; conviction is not. The organisations that win will be the ones that can tell the difference between a capability and a product, and that resist the pull to confuse motion with progress.

None of this is a reason for cynicism. It is a reason for rigour. The work is genuinely hard, the people doing it are mostly serious, and the direction of travel is clear. The job of this publication is to keep the claims honest and the evidence in view.

We will keep reporting it the way we report everything: evidence over bias, balance over noise, clarity over volume. — Lena Brandt.