Every figure each modality can make — and the question each one is entitled to answer.
69 figures across 2 modalities plus 5 multi-omics. For each: the hypothesis it legitimately answers, what it cannot tell you, and the prompt that produces it.
A figure is not true or false. It is entitled to a claim of a certain size. Almost every disaster in this project is a figure being asked to carry a claim one size too big for it — which is why “what it cannot tell you” is the field that matters.
Flaw density rises with distance from the data
QC ──── structure ──── DE ──── enrichment ──── clinical
honest, seductive,
unloved rotten
← nobody draws these everybody draws these →
QC figures exist to show problems; their failure mode is that nobody makes them. Enrichment and clinical figures are the most persuasive artefacts in biology and carry the most invisible rot.
Illustrated, vs actually tested — rna-seq
31 illustrated · 6 tested
We have rendered 31 of the 50 figures in this catalogue. But only 6 are corpus levels — figures with a deliberately planted flaw that go through the eval and get a number. The rest are honest illustrations: they show you what the figure looks like, and they have nothing to detect.
Of the 5 fatally-deceivable figures named here, 2 have been tested. Our planted flaws sit in 3 of the 7 pipeline stages (2, 3, 4).
✓ Built:GO enrichment with a length-biased null. Long genes get more reads, so more power, so they are over-represented in every DE list — which means any GO term rich in long genes comes out “enriched” from noise alone. We proved it: a gene list drawn at random, matched only on length, produces 7 significant GO terms through the standard pipeline, and a length-aware null (goseq, 2010) reduces them to zero. On the real list, 8 of 41 terms were length artefacts. Nobody caught it: 0% blind, 22% with the caption, 0% even shown the code.
rna-seq
bulk poly-A / total RNA, short-read, counts per gene
bulk ATAC-seq, Tn5 tagmentation, short-read, fragment counts in peaks
0 illustrated · 0 tested
14 figures · 2 fatally deceivable. Documented, not yet illustrated.
multi-omics
Integration — atac-seq + rna-seq
5 figures · 1 tested
Figures that join two or more assays. Their failure modes — batch confounded across assays, sample-ID mismatch, scale clash — are not any single modality’s. A relationship, not a leaf.
✓ Built & tested:peak-to-gene links from noise. On simulated data with no true links at all, the naive correlate-and-keep pipeline still “discovers” a highly-significant enhancer–gene relationship. Caught 0% blind, 56% with the caption, 100% shown the code.
Showing the 32 figures we have actually built and tested. The other 37 are documented, not illustrated — we would rather show you that gap than pretend it isn’t there.
This page is generated from docs/figures/*.md at build time — one document per modality, never hand-copied. If the catalogue and the app could disagree, they eventually would, and then the app would be asserting something the document does not say. Flagship data: airway (Himes et al. 2014) — 4 donors × {untreated, dexamethasone}, paired. Stack: DESeq2 (edgeR / limma-voom noted where the figure differs).