AI Terminology
/Intermediate
Benchmarking
Definition
The process of running a series of standardized, universally recognized tests on an AI model to evaluate its intelligence, reasoning, and coding capabilities compared to competitor models.
Explain Like I'm New
The SATs or Final Exams for Artificial Intelligence. Just like students take the SATs to prove how smart they are to colleges, tech companies force their AIs to take standard tests to prove who built the smartest model.
Real World Example
The MMLU (Massive Multitask Language Understanding) is a famous benchmark consisting of 16,000 multiple-choice questions across 57 subjects (Math, Law, Medicine). GPT-4 scored ~86% on the MMLU, beating earlier models.
Common Use Cases
- •Model selection
- •Academic research
- •Marketing AI
Interview Questions
basic
- What does the MMLU test measure?
intermediate
- What is 'Data Contamination' in AI Benchmarking?