MMLU
Why it matters
MMLU is a clean lens on a trap I see constantly: confusing a leaderboard with a fitness test. A high MMLU score says a model has broad multiple-choice knowledge; it says nothing about whether it will hallucinate on your documents, follow your output format, or stay grounded in your retrieval context. The benchmark is also saturating and partly contaminated (its questions leak into training data), so the top models now cluster within a point or two, which makes the number nearly useless for choosing between them. It is a screening signal, not an evaluation of your use case. The real evaluation is a harness on your own task.
Where you’ll encounter it
Three contexts. A model-launch announcement leading with an MMLU score, which tells you about general knowledge and almost nothing about your workload. A procurement or model-selection conversation where MMLU is used as a proxy it cannot support (pick on a task harness, not a leaderboard). And an evaluation-design discussion where MMLU’s known weaknesses (saturation, contamination, the multiple-choice format) are exactly why teams build their own held-out sets.
Part of the 7wData AI Glossary. Tracking how concepts like this move in the expert conversation: daily signals at ins7ghts.com.