Why AI Benchmarks Are Broken — MMLU, HumanEval, and Chatbot Arena
MMLU, HumanEval, ChatBot Arena — the tests we use to rank AI models are deeply flawed. Here’s exactly how, and what actually matters instead.
In 2023, a model scored 90% on MMLU — the benchmark most c
silicontosoftmax.hashnode.dev7 min read