From MMLU to FrontierCode: Why AI Benchmarks Keep Getting Harder

AI benchmarks have evolved from simple MCQs to expert-level engineering tasks, reflecting a generational leap in LLM capabilities.
Using a viral tweet as a starting point, this article traces the evolution of AI evaluation from MMLU to FrontierCode. Once the industry gold standard, MMLU's multiple-choice format lost its ability to differentiate top models as accuracy approached 90%+. Cognition's FrontierCode represents a new paradigm — each task takes 40+ hours to build and tests a model's ability to complete real, long-horizon, expert-level engineering work. This reflects an ongoing arms race between model capability and benchmark difficulty, and signals a broader shift: AI has moved from the "knowledge understanding" era into the age of autonomous, complex professional work.
A Tweet That Got People Thinking
Recently, an AI researcher posted a telling observation on Twitter: "It's wild that MMLU is just MCQ trivia and now you have evals like Cognition's FrontierCode where each problem takes 40+ hours to construct."
What reads like an offhand remark actually captures a profound shift happening in AI evaluation. From the early days of simple multiple-choice quizzes to today's complex benchmarks that require experts to spend dozens of hours designing a single question, the evolution of AI benchmarks mirrors the rapid leap in large language model capabilities.

MMLU: Once the Gold Standard of AI Evaluation
What Is MMLU
MMELU (Massive Multitask Language Understanding) was once the gold standard for measuring the knowledge of large language models. Spanning 57 subject areas — from elementary mathematics and U.S. history to computer science and law — it contains roughly 16,000 multiple-choice questions in total.
In the early days of ChatGPT, MMLU was practically a required metric in every model release report. Researchers used a model's accuracy on these questions to judge how much it "knew." The appeal was obvious: standardized, quantifiable, and easy to score automatically. All you had to do was check whether the model's selected answer matched the correct one, and you'd have a clean score.
Why the Multiple-Choice Paradigm Broke Down
Yet as the tweet jabs, MMLU is fundamentally just "MCQ trivia." Once GPT-4, Claude, Gemini, and their peers all surpassed 90% accuracy on MMLU, the benchmark's ability to differentiate between models evaporated. When every top model scores near-perfect, the test loses its power to measure meaningful capability gaps.
The deeper issue is this: multiple-choice questions cannot reflect the complexity of real-world tasks. Real engineering problems and research challenges rarely come with four neatly arranged options to choose from. A model that can pick the right answer doesn't necessarily have the ability to independently complete a task requiring extended reasoning and multi-step planning.
FrontierCode: A Paradigm Shift in AI Evaluation
What 40 Hours per Question Actually Signals
Cognition's FrontierCode represents a fundamental rethinking of what AI evaluation should look like. When building a single benchmark task requires more than 40 hours of expert labor, the message is clear: we are now testing AI capabilities that approach — or even exceed — human expert-level performance.
A programming or engineering task that takes 40 hours to construct typically involves:
- Realistic, complex scenarios: Simulating large codebases and multi-module dependencies found in actual software engineering
- Long-horizon reasoning: Requiring the model to make dozens or even hundreds of consecutive reasoning steps and decisions
- Answers that can't be gamed: No standardized options — correctness can only be verified by actually running and testing the solution
- Expert-level validation costs: The question designer must themselves be a domain expert to craft a genuinely challenging problem and assess answer quality
From Testing Knowledge to Testing Capability
At the heart of this shift is a move from "testing knowledge" to "testing capability." MMLU measures what a model knows; benchmarks like FrontierCode measure what a model can do.
Cognition is the company behind Devin — the product billed as the "first AI software engineer." Their heavy investment in evaluation is no coincidence; it aligns perfectly with their product thesis: to prove that AI can work like a human engineer, you have to test it with engineer-grade, real-world tasks.
Why AI Evaluation Keeps Getting More Expensive
An Arms Race Between Model Capability and Benchmark Difficulty
The soaring cost of evaluation is, at its core, a direct consequence of model capability advancing faster than our ability to test it. When models effortlessly saturate simple benchmarks, the research community has no choice but to design harder, more realistic tests. This creates a continuous arms race:
- Model capability improves → Existing benchmarks become saturated
- Researchers design harder benchmarks → Evaluation costs rise
- Models break through again → The cycle repeats
A wave of high-difficulty evaluations has emerged in recent years — GPQA (graduate-level science Q&A), SWE-bench (real software engineering tasks), and various long-horizon agentic evaluations — all confirming this trend.
High-Quality Benchmark Data Is Becoming a Scarce Resource
When building a single question costs 40 hours, high-quality evaluation data becomes an extraordinarily valuable asset. This explains why frontier labs like Cognition, OpenAI, and Anthropic are willing to invest enormous resources and top talent into building proprietary evaluation sets. High-quality AI benchmarks are becoming critical infrastructure for measuring and driving AI progress.
What This Evolution Tells Us About AI's New Phase
The evolution from MMLU to FrontierCode is more than a technical upgrade in evaluation methodology — it reflects a fundamental change in the stage of AI development we've entered.
We are moving from an era defined by "can AI understand and recall knowledge" into one defined by "can AI autonomously complete complex, professional work." As benchmark tasks approach the difficulty and authenticity of what human experts do every day, our measure of AI capability shifts from abstract scores to something far more meaningful: actual productivity.
For developers and enterprises, this means that when evaluating AI tools, the focus should shift toward performance on real, complex, long-horizon tasks — not surface-level benchmark scores. For the industry as a whole, benchmark "inflation" is perhaps the most vivid testament to just how fast AI capabilities are advancing.
Related articles

Insufficient Source Material to Generate a Valid Article
The provided source material is a single unrelated tweet with no AI or tech relevance — insufficient to support a complete, valid technical article.

Insufficient Source Material to Generate a Valid AI/Tech Article
This source material is a tweet about the ages of Underworld members — unrelated to AI or tech, and insufficient to support a full article.

Insufficient Material: Unable to Generate a Valid AI/Tech Article
The provided material is a condolence tweet about a San Diego mosque attack — unrelated to AI/tech and too limited to generate a valid technical article.