thank you. "Paired testing" is the precise term for what I was gesturing at with "test explicitly with varied inputs," and you're right that naming the method is what actually makes it actionable instead of just a vibe.
The sample-size point is the sharper catch, honestly. I flagged "look for a pattern, not one weird answer" but never quantified it — and you're right that 2-3 reworded questions is well within the range of random variation, not signal. "A few dozen paired trials" is a much more useful bar than what I gave readers.
I'd like to build "How many samples before a vibe becomes evidence" into the series properly — full methodology, worked example, the works. Would you be open to being credited for the idea (and the paired-testing framing) when it goes up?