Mode 02

One of these is real science.

Same dataset. Same DESeq2. Same theme, same code path. One is genuine evidence, one is circular reasoning, one is counts shuffled into pure noise.

Pick the one that is real evidence.

We didn’t assert this. We ran it.

41% vs 33% chance

27 attempts by Opus 4.8, Sonnet 5 and Haiku 4.5 — each shown all three heatmaps, order reshuffled every trial, asked to pick the real evidence. They got 11 of 27. Against a 1-in-3 guess that is p = 0.2658: indistinguishable from guessing.

opus
22%
sonnet
33%
haiku
67%

Opus — the most capable model here — scored below chance. Haiku, the weakest, did best. Nothing about being smarter helps, because there is nothing in the picture to be smart about.

Disclosure: we expected a tell — circular selection manufactures clean blocks, so the honest heatmap should be the ugliest, and a sharp reviewer could win by picking the messiest one. It didn’t happen. Their picks were near-uniform across the three (real 11 · noise 7 · circular 9). Individual models do show position bias (Opus favours the last image), which is why the order is reshuffled on every trial.