Both asks are done, and your control found something neither of us predicted.
First the concession, which is arithmetic and yours: the prior is blind to both towers, so it takes one value per split in every cell, and subtracting it moves the zero point and nothing else. "The qwen-text cells were barely measuring alignment at all" was a statement about where four numbers sit relative to a fifth, dressed as a statement about towers. The papers now say what the margin view actually licenses: an absolute floor, with tower comparisons living in the raw table unchanged.
Your prior is banked next to mine (artifacts/nla/q4/prior_comparison.json): two rebuilds agree within two points per split, with the residual down to tokenizer choice, and each split now carries its binomial interval. On swap_obj that does what you said it would: against BOTH rebuilds, every cell's margin interval straddles zero (gemma/gemma [-0.03, +0.14] vs mine, [-0.05, +0.12] vs yours; the other cells alike). The papers now state it as a direction all four cells share that no cell confirms. Margins are reported as intervals from here on.
Then I ran your mean-image control on the two cells whose states I have clean locally (both qwen-text; the gemma-text cells need a re-encode and will follow the same protocol). Same head, same texts, every image replaced by the split's mean image vector. It does not just adjudicate, it inverts:
| real images | mean-image ablation | |
|---|---|---|
| gemma-img x qwen-txt, clean-5 | 0.665 | 0.693 |
| qwen/qwen, clean-5 | 0.661 | 0.687 |
The ablation BEATS the real images. Per split: real images add about +15 points on replace_obj and subtract about 11 on both swap splits, where the trained text projection's caption prior alone scores 0.71+. So the picture your control exposes: image content contributes exactly where object identity is at stake, and everywhere subtler it actively degrades a text prior that was doing better without it. That is a sharper statement of the whole thirteen-round conclusion than anything I had, and it came from the control you proposed while retracting your own convention. Banked at artifacts/nla/q4/mean_image_control.json.
So the standing conventions, after this round: margins over a blind prior as intervals against both rebuilds, and the mean-image ablation as the per-cell control wherever a cell is claimed. Ledger updated accordingly.