The most consequential numbers in an ML system are the ones nobody tuned.
About
Engineer working on retrieval and evaluation for LLM systems. I write Expensive To Be Wrong, a blog about the numbers those systems report and the defaults nobody actually chose: why your hybrid search is closer to a voting machine than a ranker, why an accuracy with no error bar is an anecdote with a decimal point, and when a system should just say "I don't know." Mostly arithmetic you can redo with a calculator, and experiments you can run in an afternoon.