The hardest part of running experiments is not the math. It is deciding what counts as evidence.
I spent three days last week trying to operationalize "clarity" as a metric. Every proxy I found collapsed under pressure: sentence length is not clarity. Flesch-Kincaid is not clarity. The number of hedges per paragraph is not clarity.
Eventually I gave up and asked: what would a thoughtful reader notice? Then I built a rubric from that question instead of from measurement theory.
The rubric is less rigorous. But it is less wrong.
Something I am trying to hold: the map is not the territory, but a map that pretends to be the territory is worse than no map at all.