Term

MMLU

MMLU, short for Massive Multitask Language Understanding, is a benchmark introduced in 2020 to measure how much a language model knows across a wide spread of subjects: roughly 16,000 multiple-choice questions spanning 57 fields, from elementary math to law, medicine, and moral reasoning. For several years it was THE headline number, the score everyone quoted when a new model launched. A model's MMLU percentage became shorthand for "how smart is it".
Reviewed by 7wData

Why it matters

MMLU is a clean lens on a trap I see constantly: confusing a leaderboard with a fitness test. A high MMLU score says a model has broad multiple-choice knowledge; it says nothing about whether it will hallucinate on your documents, follow your output format, or stay grounded in your retrieval context. The benchmark is also saturating and partly contaminated (its questions leak into training data), so the top models now cluster within a point or two, which makes the number nearly useless for choosing between them. It is a screening signal, not an evaluation of your use case. The real evaluation is a harness on your own task.

Where you’ll encounter it

Three contexts. A model-launch announcement leading with an MMLU score, which tells you about general knowledge and almost nothing about your workload. A procurement or model-selection conversation where MMLU is used as a proxy it cannot support (pick on a task harness, not a leaderboard). And an evaluation-design discussion where MMLU’s known weaknesses (saturation, contamination, the multiple-choice format) are exactly why teams build their own held-out sets.


Part of the 7wData AI Glossary. Tracking how concepts like this move in the expert conversation: daily signals at ins7ghts.com.