The point that most of these prompts cannot be verified is the one people skip. A repo with 63K stars gets treated as ground truth even when nobody can confirm a single prompt is real. I treat scraped prompts as untrusted input and diff them against observed behaviour. How do you separate a genuine leak from a plausible reconstruction?