ArXiv: 2311.16502

🎯 Pitch

To assess how close AI is to expert-level competence, the authors present 11,500 college exam questions spanning art, medicine, and engineering that stump even top multimodal models—GPT-4V and Gemini Ultra score just 56% and 59%. The benchmark reveals a stark collapse from multiple-choice to open-ended settings, with accuracy sometimes halving even for proprietary systems, highlighting how brittle current models remain when forced to generate structured answers rather than guess.


1. Executive Summary

This paper introduces MMMU, a massive multi-discipline multimodal understanding and reasoning benchmark designed to evaluate expert-level AGI capabilities in large multimodal models. MMMU consists of 11.5K college-level questions spanning 30 subjects and 183 subfields across six disciplines, featuring 30 heterogeneous image types—from diagrams and charts to medical scans and sheet music—that demand both expert-level visual perception (identifying anatomical structures in an MRI) and deliberate reasoning with domain-specific knowledge (applying Fourier Transform or Equilibrium Theory to derive solutions). Even the most advanced proprietary LMMs—GPT-4V and Gemini Ultra—achieve only 56% and 59% accuracy respectively on the validation set, while leading open-source models reach approximately 34%, establishing that current models fall substantially short of expert human performance (88.6% best expert accuracy) and that strong performance on MMMU should be a necessary—though not sufficient—criterion for Expert AGI systems.

2. Context and Motivation

The Core Problem: We Don't Have Benchmarks That Test Expert-Level Multimodal Intelligence

The fundamental gap this paper addresses is deceptively straightforward: existing multimodal benchmarks evaluate basic perception and commonsense knowledge, not the kind of expert-level understanding and reasoning that skilled adults demonstrate in professional domains. This gap matters because it leaves the field without a meaningful yardstick for measuring progress toward Expert AGI—the level at which AI systems begin to substitute for human labor across many industries, triggering significant economic disruption (Section 1).

The paper anchors this concern to a specific operationalized definition. Drawing on Morris et al.'s (2023) leveled taxonomy for AGI, the authors identify Level 3, Expert AGI, as a critical milestone: an AI system that reaches "at least 90th percentile of skilled adults" in a broad range of tasks. Quoting directly:

"Candid and constructive discussions on AGI have been challenging due to a lack of shared operationalizable definitions."

The taxonomy provides that shared definition, but without benchmarks calibrated to it, monitoring progress becomes impossible. The authors argue that a natural starting point for measuring Expert AGI is college-level exams across different disciplines, since these are precisely the instruments designed to evaluate skilled adults specialized in each field. This strategy has precedent—benchmarks like MMLU and AGIEval have successfully adopted it for text-only evaluation—but critically, human experts solve multimodal problems, and existing multimodal benchmarks fail to capture this dimension.

Why This Gap Matters Now

The paper identifies several converging trends that make this gap urgent to address:

Rapid progress in large multimodal models (LMMs) without corresponding evaluation depth. Models like CogVLM achieve 85% on VQA-v2, 92% on ScienceQA-IMG, and 93% on RefCOCO (Section 1). These numbers suggest near-saturation on existing benchmarks, but the benchmarks themselves test relatively shallow capabilities. The paper argues that this creates a false sense of progress: models appear competent because we are not testing them on tasks that genuinely require expert-level perception and reasoning. Without harder benchmarks, the field cannot distinguish between models that have merely memorized common visual patterns and those that can perform the kind of deliberate, knowledge-intensive reasoning that experts do.

The specific challenge of multimodal expert reasoning is qualitatively different from unimodal text reasoning or basic visual QA. A radiologist interpreting an MRI doesn't just recognize shapes—they apply domain-specific knowledge about anatomy, pathology, and imaging physics to reason through diagnostic possibilities. An engineer analyzing a circuit diagram doesn't just identify components—they apply principles of electronics to calculate voltages and currents. Existing benchmarks do not require this integration of visual perception + domain knowledge + multi-step reasoning, and consequently cannot tell us whether models possess it.

The field lacks a shared benchmark for monitoring progress toward Expert AGI. The authors are explicit that MMMU is "not a sufficient test for Expert AGI" and that "there lacks a direct mapping between performance on MMMU and '90th percentile of skilled adults,' nor are college exams the only tasks an AGI shall tackle." However, they argue that "it should be necessary for an Expert AGI to achieve strong performance on MMMU to demonstrate their broad and deep subject knowledge as well as expert-level understanding and reasoning capabilities." This necessity-but-not-sufficiency framing positions MMMU as a filter: systems that cannot perform at human-expert levels on college-level multimodal problems across disciplines cannot credibly claim Expert AGI status.

Prior Benchmarks: Where They Fall Short

The paper systematically identifies limitations in existing multimodal benchmarks along two dimensions: breadth (knowledge coverage) and depth (reasoning complexity).

Limited knowledge breadth. The paper provides a detailed comparison in Figure 3 and Section 3.3. Prior benchmarks like VQA, GQA, VisWiz, OKVQA, and TextVQA focus overwhelmingly on commonsense and daily knowledge. The image formats are correspondingly narrow—typically photographs and scenes with optical character recognition (OCR) elements. ScienceQA, the benchmark closest in spirit to MMMU, covers diverse disciplines but operates at the elementary to middle school level, not college level. The paper states: "While it covers diverse disciplines (breadth), the majority of the questions are at the elementary to the middle school level, thus falling short in depth for benchmarking Expert AGI."

Limited reasoning depth. The paper draws a sharp distinction (Figure 3) between the kind of reasoning required by existing benchmarks and what MMMU demands. Prior benchmarks "normally require commonsense knowledge or simple physical or temporal reasoning." Even the more comprehensive holistic evaluations like MMBench and MM-Vet "still largely focus on relatively basic perception abilities without requiring expert-level domain knowledge and deliberate reasoning." The paper notes a recent effort, MathVista, that presents visually challenging mathematical questions, but its scope is "limited exclusively to the mathematical domain"—it cannot speak to expert understanding across the breadth of disciplines that would characterize an AGI.

Specific limitations of ScienceQA. As the closest predecessor, ScienceQA receives particular attention because its strengths (diverse disciplines, multimodal format) and weaknesses (grade-level difficulty) directly motivate MMMU's design. The authors use ScienceQA to illustrate the gap: a benchmark can have breadth without depth, and depth matters precisely because expert performance requires it. A model scoring 92% on middle-school science questions tells us little about whether it can handle the college-level biology, chemistry, and physics problems that would challenge a human major in those fields.

The challenge of interleaved text and images and heterogeneous image types. Beyond knowledge and reasoning depth, the paper identifies two technical challenges absent from existing benchmarks (Section 3.1, Figure 1). First, heterogeneous image types: MMMU includes 30 different image formats—diagrams, tables, charts, chemical structures, photographs, paintings, geometric shapes, music sheets, medical images, microscopic images, comics, mathematical notations, and more—rather than the limited palette of natural photographs that dominates prior datasets. Second, interleaved text-image inputs: questions often embed images within text (at the beginning, middle, or end), requiring models to jointly understand both modalities and connect them to domain knowledge. These design choices directly test whether models can generalize across visual formats rather than specializing in the natural images that dominate web-based pretraining data.

How MMMU Positions Itself

The paper's positioning is methodical and transparent:

MMMU is a benchmark, not a method. This is a pure evaluation contribution. The paper's ambition is to create the shared yardstick that the field currently lacks for measuring expert-level multimodal intelligence. The contributions are: (1) a carefully curated 11.5K question dataset spanning 6 disciplines, 30 subjects, and 183 subfields; (2) a comprehensive evaluation of 28 open-source and proprietary LMMs establishing a performance baseline; (3) a detailed error analysis that identifies the specific failure modes (perceptual errors, knowledge gaps, reasoning errors) that current models exhibit; and (4) a public leaderboard to track community progress.

MMMU aims for both breadth and depth simultaneously. The paper explicitly contrasts this dual ambition with prior benchmarks that achieved one or the other. The breadth comes from covering 30 subjects across 6 disciplines—from Art & Design and Music to Clinical Medicine and Mechanical Engineering—chosen based on the principle that "visual inputs should be commonly adopted in the subjects to provide valuable information" (Section 3.2). The depth comes from sourcing college-level exam, quiz, and textbook problems that require applying expert knowledge through multi-step reasoning.

MMMU intentionally creates a hard benchmark to expose model limitations. The paper's error analysis (Section 5) is deliberately diagnostic: by categorizing 150 GPT-4V errors into perceptual (35%), knowledge (29%), and reasoning (26%) failures, the authors map out the specific capabilities that current models lack. This positions MMMU not merely as a leaderboard but as a research tool for identifying where to invest effort—whether improving visual perception for rare image types, expanding domain knowledge coverage, or strengthening multi-step reasoning capabilities.

MMMU builds directly on the Expert AGI taxonomy. The connection to Morris et al.'s Levels of AGI framework gives the benchmark a clear theoretical anchor. This distinguishes MMMU from benchmarks that simply aggregate hard questions without a principled argument for what "hard enough" means. The Expert AGI framing provides that principle: the benchmark should be calibrated to test capabilities at the level of skilled human adults in their domains of expertise. The human expert validation—90 college seniors achieving 88.6% best accuracy—confirms that the benchmark indeed measures expert-level performance rather than impossibly obscure knowledge.

MMMU is positioned as necessary but not sufficient for Expert AGI. The authors are careful to bound their claims. They explicitly state that MMMU cannot single-handedly certify Expert AGI—college exams are not the only tasks an AGI must tackle, and there is no direct mapping from benchmark accuracy to the "90th percentile of skilled adults" threshold. This intellectual honesty strengthens rather than weakens the contribution: it establishes the benchmark as one important signal among many, while making a clear case for why strong multimodal expert reasoning is an indispensable component of any system claiming general intelligence.

3. Technical Approach

3.1 Reader Orientation (Approachable Technical Breakdown)

This paper constructs MMMU, a benchmark dataset and evaluation framework for measuring whether large multimodal models can perform college-level expert reasoning across 30 subjects spanning art, business, science, medicine, humanities, and engineering. The system solves the problem of assessing expert-level multimodal intelligence by collecting 11.5K carefully curated questions from college exams, quizzes, and textbooks that demand three simultaneous capabilities: accurate visual perception of 30 heterogeneous image types (from chemical structures to medical scans), recall of deep domain-specific knowledge (from art history to thermodynamics), and application of deliberate multi-step reasoning to derive correct answers—then evaluating 28 open-source and proprietary LMMs against this benchmark to establish a performance baseline that reveals substantial gaps between current models (GPT-4V at ~56%) and human experts (~89%).

3.2 Big-Picture Architecture (Diagram in Words)

The MMMU system has four major components arranged in a sequential pipeline:

  1. Data Collection Engine — a human-driven curation process involving 50 college students who source 13K multimodal questions from textbooks, online resources, and lecture materials across 183 subfields, following a principled subject selection based on whether "visual inputs should be commonly adopted in the subjects to provide valuable information."

  2. Quality Control Pipeline — a three-stage cleaning process that (a) identifies duplicates via lexical overlap and URL similarity, (b) validates formatting and corrects typos through cross-author review, and (c) categorizes questions into difficulty levels (easy/medium/hard/very easy), filtering out approximately 10% classified as "very easy" for not meeting college-level standards, yielding the final 11.5K questions.

  3. Benchmark Partition — a structured split into a few-shot development set (150 questions, 5 per subject), a validation set (~900 questions for hyperparameter selection), and a test set (~10.5K questions), with questions distributed across six disciplines (Art & Design, Business, Science, Health & Medicine, Humanities & Social Science, Tech & Engineering) and 30 subjects.

  4. Evaluation Framework — a zero-shot evaluation protocol that processes model outputs through rule-based answer extraction (regex patterns and response-processing workflows to handle long generations with intermediate reasoning steps), computes micro-averaged accuracy as the primary metric, and compares performance against multiple baselines (random choice, frequent choice, text-only LLMs with OCR/captioning augmentation, and human experts).

Information flows as follows: annotated questions enter the quality control pipeline → surviving questions are partitioned into dev/validation/test splits → models generate zero-shot answers using their default prompts for multiple-choice or open-ended QA → answer extraction pipelines parse the responses → accuracy scores are computed overall, per-discipline, per-subject, per-image-type, and per-difficulty-level → a detailed error analysis categorizes model failures into perceptual errors (35%), knowledge gaps (29%), and reasoning errors (26%).

3.3 Roadmap for the Deep Dive

  1. The subject selection and coverage rationale, because understanding what is measured determines what the benchmark actually tests and why 30 specific subjects were chosen over alternatives.

  2. The data collection protocol and annotator workflow, because the quality of a benchmark depends fundamentally on how questions are sourced, selected, and formatted—this determines whether the benchmark genuinely tests expert-level capabilities or merely aggregates trivia.

  3. The quality control and difficulty calibration process, because without rigorous filtering and difficulty categorization, the benchmark cannot serve its purpose of distinguishing model capabilities across easy, medium, and hard expert-level problems.

  4. The benchmark structure and partition design, because how questions are split into dev/validation/test sets and distributed across disciplines and question types (multiple-choice vs. open-ended) affects both evaluation validity and the practical utility of the benchmark for model development.

  5. The evaluation protocol and answer extraction, because extracting correct answers from long model outputs with intermediate reasoning is a non-trivial engineering problem that directly impacts reported accuracy scores.

  6. The human expert baseline methodology, because establishing that the benchmark indeed measures expert-level capability requires calibrating against actual human experts in each discipline.

  7. The error analysis framework, because beyond aggregate scores, understanding why models fail—and in what proportions—provides the diagnostic signal needed to guide future research.

3.4 Detailed, Sentence-Based Technical Breakdown

This is primarily a benchmark construction and evaluation paper whose core technical contribution is a carefully curated dataset and a rigorous evaluation protocol designed to measure expert-level multimodal understanding and reasoning.


Subject Selection and Coverage Rationale

The benchmark's subject coverage is not arbitrary—it follows an explicit design principle stated in Section 3.2:

"The selection is based on the principle that visual inputs should be commonly adopted in the subjects to provide valuable information."

This principle serves a dual purpose. First, it ensures that every question in the benchmark genuinely tests multimodal understanding rather than being a text-only question with a decorative image attached. Second, it excludes subjects where visual reasoning is not central—the authors explicitly note ruling out "law and linguistics because it is difficult to find enough relevant multimodal problems in these subjects." This negative selection is important: it means the benchmark does not attempt to cover all possible academic disciplines, but rather covers the subset where visual expertise is essential to professional practice.

The resulting coverage spans six broad disciplines, each comprising multiple subjects:

  • Art & Design (11% of questions): Art (Fine Arts, Drawing, Painting, Photography), Design (Design History, Graphic Design, Industrial Design), Music, and Art Theory (Art History, Art Criticism, Aesthetics). Total: 1,163 questions.
  • Business (14%): Accounting (Financial Accounting, Investment), Economics (Macroeconomics, Econometrics), Finance (Corporate Finance, Financial Marketing), Management (Cost Management, Strategic Management), Marketing (Market Research). Total: 1,428 questions.
  • Science (23%): Biology (Genetics, Cell Biology, Evolution), Chemistry (Organic Chemistry, Inorganic Chemistry, Physical Chemistry), Geography (Physical Geography, Human Geography), Math (Calculus, Linear Algebra, Graph Theory), Physics (Classical Mechanics, Optics, Electromagnetism). Total: 2,426 questions.
  • Health & Medicine (17%): Basic Medical Science (Anatomy, Neurosciences, Pathophysiology), Clinical Medicine (Pathology, Radiology, Cardiology), Diagnostics & Laboratory Medicine (Electrocardiography, Neuropathology), Pharmacy (Pharmacology, Medicinal Chemistry), Public Health (Epidemiology, Biostatistics). Total: 1,752 questions.
  • Humanities & Social Science (9%): History (World History, Modern History), Literature (Poetry, Comparative Literature), Psychology (Cognitive Psychology, Clinical Psychology, Social Psychology), Sociology (Sociology Theory, Political Economics). Total: 947 questions.
  • Tech & Engineering (26%): Agriculture (Plant Physiology, Animal Science), Architecture & Engineering (Structural Engineering, Surveying), Computer Science (Data Structures, Operating Systems, Compiler Design), Electronics (Signal Processing, Analog Electronics), Energy & Power (Thermodynamics, Fluid Mechanics), Materials (Materials Science, Mechanics of Materials), Mechanical Engineering (Engineering Dynamics, Control Systems). Total: 2,784 questions.

Each of the 30 subjects further breaks down into subfields—183 total across the benchmark. The complete taxonomy is provided in Appendix D (Table 12). This hierarchical structure (discipline → subject → subfield) is not merely organizational; it enables fine-grained analysis of model strengths and weaknesses. A model might perform well on Computer Science overall but fail specifically on Compiler Principle questions involving deterministic finite automata—the subfield granularity makes this visible.

What the distribution reveals about design intent. The largest discipline is Tech & Engineering at 26% (2,784 questions), followed by Science at 23% (2,426). The smallest are Humanities & Social Science at 9% (947) and Art & Design at 11% (1,163). This weighting reflects the availability of genuinely multimodal college-level questions—subjects like engineering and science naturally have more visual problems (circuit diagrams, chemical structures, geological maps) than subjects like literature or sociology. The authors do not attempt to artificially balance the distribution, which is a defensible choice: forcing a 50/50 split between quantitative and humanities questions would require inventing contrived multimodal problems for text-heavy subjects, potentially diluting the benchmark's quality.


Data Collection Protocol and Annotator Workflow

The data collection process is documented in Section 3.2 and elaborated extensively in Appendix H. It proceeds through three stages, each with specific procedures and quality controls.

Stage 1: Subject selection and annotator recruitment. The authors first "go through the common university majors to decide what subjects should be included" (Section 3.2), applying the visual-input centrality principle described above. They then recruit "over 50 university students, including co-authors, specializing in these majors as annotators to assist in question collection." The inclusion of co-authors among annotators is notable—it means some portion of the data collection was performed by researchers with direct domain expertise, while the remainder was distributed across a broader pool of domain-specialist students.

Stage 2: Question sourcing and creation. Annotators are instructed (Appendix H.1) to collect questions from "free online resources, quizzes, textbooks, and other study materials" while strictly adhering to copyright and licensing regulations—data from sources that prohibit copying or redistribution must be explicitly avoided. The protocol emphasizes diversity of sources: "the annotators should try to find diverse sources instead of collecting questions from a single source." When suitable questions cannot be found from existing materials, annotators are permitted to create new questions "based on their expertise where necessary" (Section 3.2). This hybrid approach—collecting most questions but authoring some—is a practical compromise that ensures subject coverage while maintaining difficulty standards.

Stage 3: Data contamination mitigation. A specific instruction in the annotation protocol (Appendix H.7) addresses the risk that foundation models have memorized answers from their training data:

"Annotators should be tasked with carefully selecting questions that go beyond straightforward queries with easily accessible answers. Instead, the focus should be on questions whose answers are tucked away in less obvious locations, such as in separate documents or hidden in the concluding sections of extensive textbooks."

This is a deliberate design choice. Rather than attempting to verify that every question is novel relative to model training corpora (which is essentially impossible to guarantee), the authors bias selection toward questions that are harder to memorize—those whose answers are not trivially retrievable from the question text alone. The underlying assumption is that models that solve these questions are more likely to be reasoning from principles than recalling memorized answers.

Question types and formats. The annotation protocol (Appendix H.2) specifies two question types: multiple-choice (including standard multiple-choice and true/false) and open-ended (including factoid, fill-in-the-blank, calculation-based, and short descriptive responses). Open-ended questions with "very long answers" are explicitly excluded to keep evaluation tractable. Each annotated sample must include: a question type classification, the question text, answer options for multiple-choice questions, the correct answer, image types, question difficulty, and an explanation if available from the source. Questions must "meet the college-level difficulty" and must not be ambiguous—they "can be answered with one of the given options or a short answer."

Data format and interleaving. The annotation protocol specifies a structured JSON format with fields for number, question type, question text, answer options, correct answer, difficulty, and explanation. Images are stored as separate PNG files following the naming convention image_{QuesNum}_{ImageNum}.png and are "inserted as a file path in the question/options/explanations" (Appendix H.3). This enables interleaved text-image inputs at arbitrary positions—the benchmark includes questions with images at the beginning (17.81%), in the middle (36.92%), and at the end (50.42%) of the question text (Table 1). This positional variability tests whether models can attend to visual information regardless of where it appears in the input sequence.

Resulting scale. This process yields approximately 13K questions initially, before quality control filtering reduces the count to the final 11,550.


Quality Control and Difficulty Calibration

The quality control pipeline operates in three sequential stages (Section 3.2), each addressing a different aspect of benchmark quality.

Stage 1: Duplicate detection. Lexical overlap (measuring word-level similarity between questions) and source URL similarity are employed to identify potential duplicates. Suspected duplicates are then manually reviewed by the authors to confirm whether they represent the same question from the same source. This is a standard deduplication step that prevents a benchmark from over-representing particular question patterns.

Stage 2: Format and typo verification. Questions are distributed among co-authors for cross-review. This is not a mechanical check—reviewers ensure that all questions "adhere to a standardized format" and make corrections where formatting deviations or typographical errors are found. Standardizing the format is particularly important for consistent answer extraction during evaluation: if different questions use different notation for answer options or embed images differently, the extraction pipeline must handle these edge cases, increasing the risk of evaluation errors.

Stage 3: Difficulty categorization. Authors classify problems into four difficulty levels: very easy, easy, medium, and hard. The protocol states that approximately 10% of problems, classified as very easy, are removed from the benchmark because they do not align with the design criteria of college-level difficulty. This filtering yields 11.5K questions distributed across three difficulty levels: easy (28%), medium (45%), and hard (27%) (Table 1). The medium-dominant distribution reflects the natural difficulty profile of college exam and textbook problems—most fall in the middle range, with some easy introductory problems and some challenging advanced problems.

Why this multi-stage approach matters. Each stage addresses a different failure mode. Deduplication prevents benchmark inflation from near-identical questions. Format verification ensures evaluation reliability. Difficulty filtering ensures the benchmark maintains its intended expert-level character. Without the third stage, the benchmark would include a tail of trivial questions that artificially inflate model scores while contributing nothing to the differentiation between models—a model could achieve high accuracy by solving easy questions while failing on the genuinely hard ones that test expert reasoning.

What "difficulty" means in this context. The difficulty labels are assigned by human annotators based on their domain expertise, not derived from model performance. This is an important design choice: difficulty is defined relative to human expert expectations (what a skilled college student in the subject would find challenging), not relative to what current models find difficult. This means the difficulty labels represent a fixed, interpretable standard rather than a moving target that shifts as models improve. It also means that analyzing model performance by difficulty level reveals whether models and humans find the same questions hard—or whether models struggle on problem types that humans find easy, which would indicate fundamentally different capability profiles.


Benchmark Structure and Partition Design

The 11.5K questions are divided into three sets with clearly defined purposes (Section 3.1, Table 1):

  • Few-shot development set: 150 questions (5 per subject). This set provides in-context examples for models that support few-shot learning. The small size (only 5 per subject) ensures that the development set covers all subjects while remaining manageable for prompt construction. The per-subject sampling guarantees that few-shot evaluation is possible for any individual subject without the few-shot examples being drawn from a different domain.
  • Validation set: approximately 900 questions. This set is explicitly designated "useful for hyperparameter selection" (Section 3.1). Its primary function is to enable researchers to tune model prompts, select the best-performing checkpoint, or choose among decoding strategies without contaminating the test set. The validation set size—roughly 8% of total questions—is small enough to keep the test set large but large enough to provide statistically meaningful signal for hyperparameter optimization.
  • Test set: approximately 10.5K questions. This is the primary evaluation split. With an average of approximately 350 questions per subject (varying by discipline distribution), the test set provides enough samples per subject for meaningful per-discipline and per-subject comparisons (Table 4 in Appendix B provides per-subject sample counts).

Question type distribution. Multiple-choice questions dominate at 94.03% (10,861 questions), with open-ended questions comprising 5.97% (689 questions) (Table 1). This heavy skew toward multiple-choice is a deliberate practical choice. The authors state in the conclusion that MMMU "combines multiple-choice questions with concise open-ended questions, enabling the assessment of diverse subjects while addressing the challenges associated with evaluating open-ended responses." The implicit reasoning: evaluating open-ended responses at scale requires either human judgment (expensive and slow) or automated metrics that may be unreliable, particularly for expert-level content where surface-level matching can be misleading (a correct physics answer expressed in different notation would fail exact-match evaluation). Multiple-choice questions eliminate this ambiguity—answer correctness is binary and evaluation is fully automated.

Questions with explanations. 2,035 questions (17.62%) include explanations (Table 1). These explanations, sourced from the original exam materials or textbooks, describe the reasoning steps and domain knowledge required to reach the correct answer. They serve as training signal for researchers developing models that generate reasoning chains, and as diagnostic material for understanding what kinds of reasoning particular questions demand. The fraction with explanations is a result of availability—questions were only annotated with explanations when the source material provided them—rather than a targeted sampling strategy.

Image characteristics. 97.52% of questions (11,264) contain at least one image in the question text, approximately 3.37% (389) contain images in the answer options, and 7.39% (854) contain multiple images (Table 1). The small fraction of questions without images (2.48%) appears to be an edge case—likely text-only questions that were included because they are part of subject areas where most questions are multimodal and the text-only examples provide useful context or coverage.

Image positional distribution. Table 1 reports three image positions: images at the beginning of the question (17.81%), in the middle (36.92%), and at the end (50.42%). These percentages sum to more than 100% because some questions contain multiple images at different positions. The distribution reveals that images are not uniformly positioned—half appear at the end of the question text—which means models that process sequences left-to-right must retain question context before seeing the most common image location. This design choice tests whether models can integrate text and image information regardless of presentation order, a capability that is not tested by benchmarks where images always appear in a fixed position.

Discipline and subject sampling. The test set distribution across disciplines (Table 2) is: Art & Design (1,163), Business (1,428), Science (2,426), Health & Medicine (1,752), Humanities & Social Science (947), Tech & Engineering (2,784). Within each discipline, subjects vary substantially in sample count—for example, within Business, Accounting has 380 questions while Marketing has 181. This reflects the natural availability of multimodal questions in each subfield rather than an attempt to enforce uniform coverage. A subject like Market Research (the Marketing subfield) simply has fewer college-level questions that centrally involve images compared to Financial Accounting.


Evaluation Protocol and Answer Extraction

The evaluation framework is designed for zero-shot assessment, meaning models are evaluated without fine-tuning on MMMU data and without few-shot demonstrations from the benchmark itself (Section 4). This choice maximizes practical relevance: it tests how well models can generalize their pretrained knowledge to new expert-level problems, which is closer to how an AGI system would need to operate than a model fine-tuned specifically for the benchmark.

Model invocation. For each model, the authors "use the default prompt provided by each model for multi-choice or open QA, if available. If models do not provide prompts for task types in MMMU, we conduct prompt engineering on the validation set and use the most effective prompt for the zero-shot setup in the main experiments" (Section 4). This hybrid approach—using model-provided defaults where they exist and validation-set-optimized prompts otherwise—aims to give each model its best fair chance while avoiding test-set leakage through prompt optimization. The validation set serves its intended purpose here: it absorbs the cost of prompt engineering so the test set remains uncontaminated.

Answer extraction from long responses. This is a non-trivial engineering challenge, particularly for models (like GPT-4V) that generate detailed reasoning chains, intermediate calculations, and natural language commentary before producing a final answer. The paper's approach (Section 4) is:

"To mitigate the potential influence of any intermediate generations (e.g., reasoning steps, calculations) in the long response, we construct robust regular expressions and develop response-processing workflows. These are employed to extract key phrases, such as numbers and conclusion phrases, from the long responses for accurate answer matching."

The extraction pipeline operates in two modes. For multiple-choice questions, it searches for answer indicators—option letters (A, B, C, D), the text of the chosen option, or explicit answer declarations (e.g., "The correct answer is: (B)"). For open-ended questions, it searches for numerical answers, key phrases, or fill-in-the-blank completions. The "robust regular expressions" are designed to handle the variability in how models format their responses: some output "Answer: B", others output "(B) Option text", and others embed the choice in prose like "I believe the answer is B because...".

Fallback for unparseable responses. If the extraction pipeline finds no valid answer—either because the model refused to answer, produced an incoherent response, or formatted its answer in an unexpected way—the protocol specifies:

"If there is no valid answer in the model's response, we perform random selection as a remedy for multiple-choice questions or consider the response incorrect for open questions."

This fallback is conservative but necessary. For multiple-choice questions, random selection ensures that unparseable responses contribute the expected chance-level accuracy rather than being treated as correct (which would inflate scores) or incorrect (which would penalize models for formatting issues rather than capability failures). For open-ended questions, treating unparseable responses as incorrect is the only defensible choice—there is no random baseline for a fill-in-the-blank or calculation question.

Evaluation metric. The primary metric is micro-averaged accuracy: the fraction of all test questions for which the extracted answer matches the ground truth. "Micro-averaged" means accuracy is computed over all questions equally, regardless of which discipline or subject they belong to—each question contributes one count to the numerator (if correct) and one count to the denominator. This is distinct from macro-averaged accuracy, which would compute accuracy per subject and then average the per-subject accuracies. Micro-averaging weights subjects proportionally to their sample counts, so a model's performance on Tech & Engineering (2,784 questions) influences the overall score more than its performance on Humanities & Social Science (947 questions).

Baseline computations. The paper includes two naïve baselines for calibration:

  • Random Choice: selects an option uniformly at random for multiple-choice questions. The expected accuracy depends on the number of options but averages approximately 22–25% across the benchmark (varying by discipline because of different option counts).
  • Frequent Choice: selects the most frequent answer option within each subject on the validation set, based on its observed frequency. This baseline captures whether models are simply exploiting answer distribution biases (e.g., if option C is disproportionately common) rather than understanding the questions.

These baselines establish a performance floor. Any model that performs at or below frequent choice is effectively not engaging with the questions' content.

Cross-validation for strategy selection. When testing few-shot models like OpenFlamingo and Otter (Appendix G), the authors use the development set as the source of in-context examples. The reported few-shot results (Table 14) show that performance actually degrades with more shots for Otter (from 0.291 at 0-shot to 0.258 at 3-shot and 5-shot) and remains essentially flat for OpenFlamingo (0.263 at 0-shot, 0.264 at 5-shot). This negative result is itself informative: it suggests that the questions are too complex for current few-shot learning mechanisms to extract useful patterns from a small number of examples.


Human Expert Baseline Methodology

To establish that MMMU genuinely measures expert-level capability, the authors conduct a human evaluation with a carefully specified protocol (Section 4.1):

Participant selection. "90 college senior students, selected to represent a wide range of experts in the corresponding 30 subjects (3 student experts per subject)" are recruited. The use of senior students (rather than professors or practitioners with decades of experience) is deliberate: it matches the benchmark's college-level difficulty—these are the "skilled adults" that college exams are designed to evaluate. Three experts per subject provides minimal redundancy (enough to capture variance across individuals while keeping the total participant count manageable).

Task assignment. Each student is "tasked with completing the 30 questions in their corresponding subjects (900 validation questions in total)." The assignment is within-subject: each student only answers questions from their own discipline. This ensures that the human baseline reflects genuine expertise rather than general intelligence—a physics major answering art history questions would not represent expert performance.

Testing conditions. Students are "allowed to consult their textbooks but were prohibited from searching the Internet for answers." This condition mirrors the open-book nature of the benchmark for models: models have access to their pretrained knowledge (analogous to consulting a textbook) but cannot perform real-time web search (unless explicitly designed to do so). "Prohibited from searching the Internet" closes a loophole where a human could simply look up the exact question online, defeating the purpose of testing expert knowledge.

Results. The validation set results (Table 2) are: worst expert 76.2%, median expert 82.6%, best expert 88.6%. These numbers establish the performance ceiling against which models are compared. The gap between best expert (88.6%) and GPT-4V (56.8% on validation) is approximately 32 percentage points—a substantial deficit that cannot be attributed to measurement noise given the validation set size (900 questions). The inter-expert variance (worst to best spread of 12.4 percentage points) reflects genuine differences in individual expertise even among senior students in their field, and provides a sense of the range within which a model achieving "expert-level" performance would need to fall.

Why 30 questions per subject is sufficient. The per-subject human evaluation sample (30 questions × 3 experts = 90 responses per subject) is small relative to the per-subject test set size (typically hundreds of questions). However, for calibration purposes—establishing that the questions are answerable by experts and measuring the performance ceiling—30 questions per subject is adequate. A larger human evaluation would be more statistically robust but would have required either more expert participants (expensive) or each expert answering more questions (increasing burden and potentially reducing answer quality due to fatigue).


Error Analysis Framework

The error analysis (Section 5) is designed to be diagnostic rather than merely descriptive. The authors randomly sample 150 error instances from GPT-4V's predictions and have "expert annotators who identify the root causes of mispredictions based on their knowledge and the golden explanations if available." This annotation process categorizes each error into one of five types (Figure 5):

Perceptual Errors (35%, 53 out of 150 errors). These are errors where the model correctly understands the task and possesses the necessary domain knowledge, but fails at the basic level of interpreting visual information. The paper further subdivides this category:

  • Basic perceptual errors: the model "accurately processes and understands the given information but fails in elementary visual interpretation, such as misjudging the sequence described as 'from left to right, top to bottom'" (Section 5). The illustrative example in Figure 6 shows GPT-4V correctly reasoning about egoism vs. other-isms in an oxygen mask scenario but failing to map picture IDs to the correct illustrations because the order is only described in text, not visually marked in the figure. This subtype reveals that even sophisticated models struggle with spatial reasoning tasks that human adults perform effortlessly.

  • Domain-specific perceptual errors: errors that combine perceptual failure with knowledge gaps. The authors classify these under "Lack of Knowledge" for root cause analysis because the perceptual failure is downstream of missing domain expertise—the model cannot perceive what it does not know to look for. An example from Appendix Figure 83: GPT-4V fails to interpret double circles in a deterministic finite automaton diagram as denoting "accept states" because it lacks the computer science knowledge that defines this convention.

The paper also notes a specific systematic bias: "GPT-4V often exhibits a bias towards text, prioritizing textual information over visual inputs, a trend noted in recent studies." The example in Figure 67 (Appendix) shows the model incorrectly prioritizing its text-based interpretation of "imperialism" over the visual narrative in a political cartoon that actually depicts the United States as a "Savior."

Lack of Knowledge (29%, ~44 out of 150 errors). These errors occur when the model correctly perceives visual information but lacks the domain-specific knowledge to interpret it. The error is not about what the model sees but about what it knows. The medical example in Appendix Figure 54 illustrates this: GPT-4V correctly reads the hemodynamic values from a cardiac catheterization table but fails to recognize that the combination of low aortic diastolic pressure and high left ventricular diastolic pressure is classic for aortic valve regurgitation—knowledge that a cardiology resident would have internalized. Similarly, Appendix Figure 68 shows a history question where the model knows the time period but does not know that British economic interactions with India during 1600–1870 were "chiefly concerned with cotton" rather than opium.

Reasoning Errors (26%, ~39 out of 150 errors). These errors occur when the model correctly perceives visual information, correctly recalls relevant domain knowledge, but fails to apply logical or mathematical reasoning to derive the correct conclusion. Appendix Figure 45 provides a concrete example: GPT-4V correctly identifies the formula for the volume of a cone and correctly identifies the optimization task (minimize surface area given fixed volume), but makes algebraic errors during the computation, leading to an incorrect conclusion. Another example in Appendix Figure 45 shows the model missing a necessary step in a chain of mathematical reasoning.

Textual Understanding Errors (6%, ~9 out of 150). These errors involve misinterpreting the text of the question itself—the linguistic content—rather than the visual content. The boundary between textual understanding and reasoning errors is somewhat fuzzy, but the category captures cases where the model misreads the question's intent or misinterprets technical notation. Appendix Figure 63 shows an example where the model extracts the wrong data field from a table (reading "EV71-related hand, foot, and mouth disease or herpangina" when the question asked only for "hand, foot, and mouth disease").

Answer Extraction Errors (1%), Annotation Errors (2%), Refusal to Answer (3%). The remaining small fraction of errors are attributed to technical failures (the answer extraction pipeline failing to parse a correctly reasoned response), ground-truth errors (rare annotation mistakes in the benchmark), or the model explicitly declining to answer (e.g., Appendix Figure 57 where GPT-4V responds to an ophthalmic pathology question with "I cannot make a definitive choice. However, it's crucial to collaborate with a pathologist or an ocular oncologist for a comprehensive evaluation.")

Why this categorization matters. The error taxonomy is not merely descriptive—it has direct implications for research prioritization. If perceptual errors dominate, the focus should be on improving visual encoders and multimodal fusion. If knowledge gaps dominate, the priority is expanding domain-specific training data. If reasoning errors dominate, the bottleneck is the model's chain-of-thought and multi-step inference capabilities. The finding that errors are roughly equally distributed across perceptual (35%), knowledge (29%), and reasoning (26%) categories suggests that all three dimensions are simultaneously bottlenecks for current models—there is no single dominant failure mode, and progress toward Expert AGI requires improvements across perception, knowledge, and reasoning.


Summary of Key Design Choices and Their Justifications

  • College-level difficulty calibration via human experts rather than model-based difficulty estimation: ensures the benchmark measures expert capability relative to a fixed human standard, not relative to current model weaknesses. The human expert validation (88.6% best accuracy) confirms the benchmark is answerable by domain experts.

  • Overwhelmingly multiple-choice format (94%) rather than open-ended: enables fully automated, unambiguous evaluation at scale while eliminating inter-annotator disagreement in answer grading. The small open-ended component (6%) provides a signal for models' ability to generate answers without being prompted with options.

  • Subject selection via visual-input centrality rather than exhaustive coverage: focuses the benchmark on domains where multimodal understanding is genuinely essential to expertise, avoiding contrived questions in text-heavy fields that would not test multimodal capability.

  • Zero-shot evaluation as primary protocol rather than fine-tuning or few-shot: tests generalization from pretrained knowledge, which is closer to AGI-like capability than narrow task-specific adaptation. The few-shot results (Appendix G) confirm that current open-source models do not benefit from in-context examples on this benchmark, further justifying the zero-shot default.

  • Rule-based answer extraction rather than LLM-based grading: eliminates the risk of evaluation models hallucinating grades or introducing their own biases. The engineering cost (building robust regex patterns) is traded off against evaluation reliability.

  • Frequent choice baseline alongside random choice: detects whether models are exploiting answer distribution biases—a model that performs near frequent choice is effectively not engaging with question content. This is a simple but important calibration check.

  • Difficulty categorization by human annotators rather than by model pass@1 rates: anchors difficulty to a fixed, interpretable standard (what skilled students find challenging) that does not shift as models improve. This enables meaningful longitudinal tracking of progress.

  • Two-stage annotation (student collection + author review) with explicit deduplication and format verification: balances the scale achievable with distributed annotation against the quality standards required for a benchmark that will be used to evaluate frontier models.

4. Key Insights and Innovations

Innovation 1: Difficulty as a Diagnostic Lens, Not Just a Metric

The paper's most conceptually distinctive move is using question difficulty not merely as a reporting axis (how models perform on easy vs. hard problems) but as a diagnostic lens that reveals how model capability degrades. This goes well beyond the standard practice in benchmark papers of showing a table with "Easy / Medium / Hard" accuracy columns.

Prior benchmarks typically report difficulty-stratified results as a sanity check—verifying that harder questions indeed yield lower scores. MMMU's innovation is to treat difficulty as a probe into the qualitative nature of model failure. The analysis in Section 4.3 (Table 3) shows something striking: GPT-4V achieves 76.1% on easy questions but only 31.2% on hard ones—a 45-point gap. Open-source models like LLaVA-1.5 drop from 41.3% to 26.7%—a 15-point gap. The gap itself is not the insight. The insight is that the gap reveals fundamentally different scaling behavior: open-source models hover near random guessing on hard questions regardless of architecture (all models in Table 3 cluster within ~5 points of each other on "Hard"), while GPT-4V's advantage on easy and medium questions almost completely evaporates at the hard level.

This framing—difficulty exposes the ceiling of current reasoning capabilities rather than just ranking models—is conceptually distinct from prior work. ScienceQA reported difficulty breakdowns but as an attribute of the data, not as a diagnostic instrument for understanding where model reasoning collapses. MMMU's difficulty analysis transforms "hard questions" from a score bucket into a signal that current models, even GPT-4V, plateau near random performance on problems requiring the most sophisticated domain-specific reasoning—and that this plateau is nearly uniform across model families. This suggests a fundamental bottleneck, not a gradual slope that more scale will naturally climb.

The evidence that anchors this: Table 3 shows GPT-4V at 31.2% on hard questions versus random choice at ~23.9%—a gain of only ~7 points over chance, compared to a ~52-point gain on easy questions. Open-source models gain essentially zero over random on hard questions. This pattern cannot be explained by "harder questions are harder" alone—it indicates a regime where current architectures fail to engage with the reasoning demands at all.


Innovation 2: Categorizing Multimodal Failure into Perceptual, Knowledge, and Reasoning Errors

The error taxonomy introduced in Section 5 is a diagnostic framework with implications beyond the paper's own results. It decomposes model failure into three root causes—perceptual errors (35%), lack of knowledge (29%), and reasoning errors (26%)—that correspond to three distinct architectural bottlenecks in LMMs.

Prior benchmark papers typically report aggregate error rates and provide qualitative examples, but do not systematically categorize why models fail. ScienceQA, MM-Vet, and MMBench report per-skill or per-category scores, but these categories are defined by the task type (OCR, spatial reasoning, etc.), not by the underlying cause of failure. A model might answer a spatial reasoning question incorrectly because it failed to perceive the spatial layout (perceptual error), because it didn't know the relevant geometry principles (knowledge gap), or because it could not chain the spatial and mathematical inferences (reasoning error)—but prior benchmarks collapse these into a single "spatial reasoning accuracy" number that conflates distinct failure modes.

MMMU's taxonomy breaks this conflation. The evidence is concrete: Figure 6 shows a clear perceptual error (model correctly knows the ethical framework and correctly reasons about reconciliation of egoism vs. other-isms, but fails to count pictures left-to-right, top-to-bottom as instructed in text). Appendix Figure 54 shows a clear knowledge error (model reads hemodynamic values correctly but doesn't know the diagnostic criteria for aortic valve regurgitation). Appendix Figure 45 shows a clear reasoning error (model correctly identifies the optimization problem and formulas but makes algebraic mistakes in the computation). These are not just examples—they demonstrate that the three failure categories are separable and diagnosable empirically, not merely conceptual distinctions.

The significance of this innovation is that it maps directly to research investment decisions. If perceptual errors dominate, the priority is visual encoders and multimodal fusion. If knowledge gaps dominate, the priority is domain-specific pretraining or retrieval augmentation. If reasoning errors dominate, the priority is chain-of-thought and multi-step inference architectures. That the errors are roughly equally distributed (~1/3 each) implies all three fronts must advance simultaneously—there is no single bottleneck whose removal would dramatically improve performance. This is a genuinely diagnostic contribution that the paper's own experimental design makes possible.


Innovation 3: The "Necessary but Not Sufficient" Framing for Expert AGI Evaluation

The paper's explicit positioning—that strong MMMU performance should be necessary but not sufficient for Expert AGI—represents a conceptual contribution to how the field thinks about benchmarking general intelligence. This is more than careful language in the conclusion; it is a principled methodological stance that addresses a recurring tension in AI evaluation.

The tension is this: benchmarks that claim to measure "AGI" or "general intelligence" invite criticism that they reduce a multifaceted concept to a single score or set of tasks. The typical response is to either overclaim (treating the benchmark as a sufficient test for AGI) or underclaim (presenting the benchmark as just another dataset without connecting it to broader intelligence questions). The former invites debunking; the latter forfeits the benchmark's relevance to the AGI conversation that motivates much of the field's research.

MMMU navigates this tension by anchoring to an external framework—Morris et al.'s Levels of AGI taxonomy—and then making a modular claim: the taxonomy defines what Expert AGI means (90th percentile of skilled adults); MMMU is one concrete, operationalized component of what an Expert AGI must be able to do (solve college-level multimodal problems); but MMMU is not the only component, and performance on it does not by itself certify Expert AGI status. The connection to college exams provides the necessary grounding—these are real instruments designed to evaluate skilled adults—while the modular framing prevents overclaiming.

This is significant because it provides a template for future benchmark construction. Rather than each benchmark implicitly competing to be the "AGI test," benchmarks can position themselves as measuring specific necessary conditions for AGI at specific levels. MMMU's "necessary but not sufficient" framing, combined with its explicit connection to the Morris et al. taxonomy, makes this methodological contribution transferable to other benchmarks targeting different dimensions of general intelligence.

The evidence for this innovation is not in a single figure or table but in the architecture of the paper's argument: Section 1 establishes the Expert AGI definition and positioning; Section 3.2 describes a data collection process calibrated to college-level difficulty; Section 4.1 anchors the benchmark to human expert performance (88.6% best); and the conclusion explicitly bounds the claims. This integrated argument—connecting a theoretical AGI framework to a concrete evaluation methodology with transparent limitations—is genuinely distinctive relative to prior benchmarks that either avoid the AGI framing entirely or make implicit claims without the modular structure.


Innovation 4: Heterogeneous Image Types as a Stress Test for Visual Generalization

MMMU's inclusion of 30 distinct image types—from sheet music and chemical structures to medical scans and political cartoons—constitutes a deliberate stress test for visual generalization rather than merely a reporting axis. This is fundamentally different from prior benchmarks that either restricted themselves to natural photographs (VQA, GQA, VisWiz) or included a small number of chart/diagram types alongside photographs (MMBench, MM-Vet).

The intellectual move is recognizing that visual generalization across image types is a distinct capability from visual perception within a single type—and that it is a capability professional experts demonstrably possess (a radiologist can read an MRI, a schematic diagram, and a photograph, all fluently) but current LMMs lack. The evidence is in Figure 4 and Table 13 (Appendix F): GPT-4V achieves 64.2% on photographs but only 38.8% on sheet music and 40.2% on geometric shapes. Open-source models show even starker drops—some image types see performance near random guessing. The paper notes this explicitly: "for less common image categories like Geometric shapes, Music sheets and Chemical structures, all models obtain very low scores (some are close to random guesses). This indicates that the existing models are generalizing poorly towards these image types."

Prior benchmarks treated image type diversity as coverage breadth—more types means a more comprehensive benchmark. MMMU reframes it as a capability test: can models transfer visual understanding across the range of representations that human experts fluently navigate? The answer, based on Figure 4, is no—and the pattern of failure is not uniform. Models perform relatively well on photographs and paintings (likely well-represented in pretraining data) and poorly on specialized notation systems (sheet music, chemical structures, geometric diagrams) that require learned conventions for interpretation. This is evidence of a training-data bias rather than a fundamental visual reasoning limitation, but identifying it as such—and providing the benchmark to measure it—is the contribution.

The performance breakdown by image type in Table 13 (Appendix F), covering all 30 types across six models including GPT-4V, provides the fine-grained evidence that makes this insight actionable: researchers can now identify exactly which visual domains their models fail on, rather than knowing only that "image understanding needs improvement." This transforms image-type reporting from a descriptive statistic into a diagnostic tool for visual generalization.


Innovation 5: The OCR and Captioning Augmentation Negative Result

The paper includes a carefully designed negative result that carries conceptual weight beyond its numerical values. Section 4.2 reports that augmenting text-only LLMs with OCR (optical character recognition) or generated captions (from LLaVA-1.5) yields essentially no improvement on MMMU. FLAN-T5-XXL achieves 31.2% on the test set without augmentation, 31.9% with OCR, and 31.9% with LLaVA captions. Vicuna-13B achieves 31.0% without augmentation, 31.9% with OCR, and 32.7% with captions—the largest gain being a marginal 1.7 percentage points.

This is a negative result with substantial implications. The dominant assumption in some multimodal research is that text-only LLMs can be made "multimodal enough" by converting images to text descriptions—if an OCR system extracts the text in a chart and a captioning model describes the visual content, the LLM can reason over this textual representation without needing native visual understanding. MMMU's negative result directly challenges this assumption for expert-level tasks. The paper states the implication clearly:

"This finding suggests that the MMMU benchmark requires models that can effectively interpret and integrate both textual and visual information, underscoring the complexity of the multimodal tasks it presents."

The significance is that this result establishes a lower bound on the necessity of genuine multimodal architectures for expert-level reasoning. OCR and captioning work reasonably well for tasks where the visual content is simple enough to be fully captured in a text description (reading text from signs, identifying common objects). MMMU's expert-level questions—which involve interpreting nuanced visual patterns in medical images, reasoning about spatial relationships in engineering diagrams, or reading domain-specific notation in chemical structures—cannot be reduced to text descriptions without losing the information needed for correct reasoning. The negative result proves this empirically rather than assuming it.

This finding also serves as a validity check on the benchmark itself: if OCR-augmented text-only models performed comparably to native multimodal models, it would suggest MMMU's visual content is superficial or decorative rather than genuinely necessary for answering the questions. That OCR and captioning provide no meaningful boost confirms that MMMU's images carry information essential to solving the problems—information that cannot be adequately captured by off-the-shelf vision-to-text pipelines. This is an important calibration that strengthens confidence in the benchmark as a test of multimodal understanding specifically, not just text-based reasoning on multimodal-formatted questions.

5. Experimental Analysis

Evaluation Methodology

  • Dataset. All experiments use the MMMU benchmark itself: 11.5K college-level multimodal questions split into a few-shot development set (150 questions, 5 per subject), a validation set (~900 questions for hyperparameter selection and prompt engineering), and a test set (~10.5K questions for final reported results). Questions span 30 subjects across six disciplines, with 94.03% multiple-choice and 5.97% open-ended formats.

  • Base model(s). The evaluation covers a broad spectrum of both proprietary and open-source models to establish a comprehensive performance baseline. Proprietary LMMs include GPT-4V(ision) (Playground version), Gemini 1.0 Ultra, Gemini 1.5 Pro, Claude 3 Opus, and GPT-4o. Open-source LMMs span multiple model families at various scales: BLIP-2 FLAN-T5 (XL and XXL), InstructBLIP (T5-XL and T5-XXL), LLaVA-1.5 (13B), LLaVA-1.6 (34B), CogVLM, Qwen-VL (7B-Chat, PLUS, MAX), InternVL-Chat (V1.1 and V1.2), InternLM-XComposer2-VL, Yi-VL (6B and 34B), VILA1.5, OpenFlamingo (9B), Fuyu (8B), and others. Text-only baselines include GPT-4 (text), Llama2-7B, FLAN-T5-XXL, and Vicuna-13B. The models are chosen to represent the state-of-the-art across both closed and open ecosystems, with model scale ranging from ~1.6B parameters (Kosmos2) to GPT-4V/Gemini-class systems. Evaluations use the latest, largest available checkpoint for each model family by default.

  • Metrics. The primary metric is micro-averaged accuracy: the fraction of all test questions for which the model's extracted answer matches the ground-truth answer exactly. Micro-averaging means each question contributes equally to the overall score regardless of which discipline or subject it belongs to, so disciplines with more questions (e.g., Tech & Engineering at 2,784) influence the aggregate more than those with fewer (e.g., Humanities & Social Science at 947). For both multiple-choice and open-ended questions, a systematic rule-based evaluation pipeline extracts key phrases (option letters, numbers, conclusion phrases) from long model responses using robust regular expressions. If no valid answer is extracted, multiple-choice questions receive a random guess, and open-ended questions are marked incorrect. The paper also reports per-discipline, per-subject, per-image-type, and per-difficulty-level accuracy to enable fine-grained diagnostic analysis.

  • Baselines. The paper employs multiple baselines to calibrate interpretation. Random Choice selects an option uniformly at random, establishing a chance-level floor that varies by discipline (averaging 22.1% on validation, 23.9% on test). Frequent Choice selects the most common correct answer option within each subject on the validation set based on its observed frequency—any model performing near this baseline is exploiting answer distribution biases rather than question content. Text-only LLMs (GPT-4, Llama2-7B, FLAN-T5-XXL, Vicuna-13B) test whether textual reasoning without visual input suffices. OCR-augmented text-only LLMs add text recognized by MMOCR to the input. Caption-augmented text-only LLMs add image descriptions generated by LLaVA-1.5. Human Experts (90 college senior students, 3 per subject, each answering the 30 validation questions in their specialty) establish the expert performance ceiling, with results reported as worst (76.2%), median (82.6%), and best (88.6%) expert accuracy on the validation set.

  • Generation budget / compute accounting. The paper does not compare methods under a common FLOPs or inference-time compute budget, as this is a pure benchmark evaluation rather than a methods comparison. Each model is evaluated under its standard inference protocol (zero-shot generation with default or validation-set-optimized prompts). The primary fairness mechanism is evaluating all models identically—same questions, same answer extraction pipeline, same zero-shot setting—rather than controlling for computational cost. This means comparisons between models of vastly different scales (e.g., Kosmos2 at 1.6B vs. GPT-4V) reflect genuine capability differences but conflate model scale, training data, and architecture.

  • Cross-validation / statistical protocol. For models that do not provide default prompts for the task types in MMMU, prompt engineering is conducted on the validation set to select the best-performing prompt, which is then used for zero-shot evaluation on the test set. This keeps test-set results uncontaminated by prompt optimization. For few-shot evaluation (OpenFlamingo and Otter, Appendix G), the development set provides in-context examples, with results reported at 0, 1, 3, and 5 shots. The paper reports results from two sources: authors' own evaluations (conducted with NVIDIA A100 GPUs) and author-provided results (indicated with an asterisk in all tables, representing models where the model developers ran the evaluation and reported results). Human expert evaluation uses 3 experts per subject × 30 subjects = 90 total participants, each answering 30 validation questions in their specialty.


Main Quantitative Results

Overall Benchmark Difficulty

The headline finding from Table 2 is that MMMU poses a severe challenge to all current models. On the validation set (900 questions):

  • Random Choice achieves 22.1% and Frequent Choice achieves 26.8%.
  • The best human expert achieves 88.6%, with the median expert at 82.6% and the worst at 76.2%.
  • GPT-4V achieves 56.8% on validation and 55.7% on the test set (10,500 questions)—approximately 32 points below the best human expert.
  • Gemini 1.0 Ultra achieves 59.4% on validation.
  • GPT-4o achieves 69.1% on validation—a notable improvement over GPT-4V but still approximately 19 points below the best expert.
  • The highest-performing open-source model at the time of paper submission, BLIP-2 FLAN-T5-XXL, achieves 35.4% on validation and 34.0% on the test set—roughly 22 points below GPT-4V.

On the test set (Table 2), the performance ordering among representative models is:

  • SenseChat-Vision-0423-Preview: 50.3%
  • GPT-4V(ision): 55.7%
  • InternVL-Chat-V1.2: 46.2%
  • VILA1.5: 46.9%
  • LLaVA-1.6-34B: 44.7%
  • BLIP-2 FLAN-T5-XXL: 34.0%
  • Random Choice: 23.9%

The key takeaway is that even the most advanced proprietary models are approximately 30 percentage points below expert human performance, and open-source models lag an additional 20+ points behind GPT-4V. This validates the paper's central claim that MMMU is substantially harder than prior multimodal benchmarks—where models routinely achieve 80–90% accuracy—and that significant room for improvement exists.


Performance Disparities Across Disciplines

The per-discipline breakdown (Table 2, test set) reveals a consistent and large performance gap between proprietary and open-source models, but also shows that this gap varies dramatically by discipline. For GPT-4V versus the best open-source model (BLIP-2 FLAN-T5-XXL at submission time):

  • Art & Design: GPT-4V achieves 65.3% vs. 49.2% for BLIP-2 FLAN-T5-XXL (a gap of ~16 points). Open-source models perform relatively well here.
  • Business: GPT-4V achieves 64.3% vs. 28.6% (a gap of ~36 points). Open-source models struggle disproportionately.
  • Science: GPT-4V achieves 48.4% vs. 27.3% (a gap of ~21 points). Both proprietary and open-source models find Science challenging.
  • Health & Medicine: GPT-4V achieves 63.5% vs. 33.7% (a gap of ~30 points). Large disparity.
  • Humanities & Social Science: GPT-4V achieves 76.3% vs. 51.5% (a gap of ~25 points). This is the highest-performing category for both model types, suggesting these questions rely on image types (paintings, photographs, comics) that are better represented in pretraining data or require less specialized domain knowledge.
  • Tech & Engineering: GPT-4V achieves 41.7% vs. 30.4% (a gap of ~11 points). This is both the lowest-performing category for GPT-4V and the category where the proprietary-open-source gap is narrowest—both model types struggle substantially.

The pattern is revealing: disciplines where visual data is more "natural" (Art & Design, Humanities & Social Science) show stronger model performance and narrower gaps, while disciplines requiring specialized visual conventions (Science with chemical structures and mathematical notation, Tech & Engineering with circuit diagrams and technical blueprints, Health & Medicine with medical images) show lower performance and wider or narrower gaps depending on the specific visual demands. The authors summarize this in Section 4.2:

"In disciplines such as Art & Design and Humanities & Social Sciences, where the images tends to be more 'natural' and questions involve relatively less reasoning, models demonstrate relatively higher performance. Conversely, in fields like Science, Health & Medicine, and Technology & Engineering, where tasks often involve intricate perception and complex reasoning, models exhibit lower performance."

The per-subject breakdown (Appendix B, Tables 5–10) provides even finer granularity, showing that within a discipline like Science, GPT-4V achieves 56.4% on Physics but only 44.8% on Geography and 45.0% on Math—the variation within a discipline can exceed the average variation across disciplines.


Performance on Different Image Types

Figure 4 and Appendix F (Table 13) report model accuracy across 30 distinct image types, revealing substantial heterogeneity in model capabilities:

  • GPT-4V consistently outperforms all open-source models across every image type. Its strongest categories include Advertisements (100.0%), Logos and Branding (85.7%), Poster (80.7%), Icons and Symbols (78.6%), Sculpture (76.1%), and Paintings (75.9%). Its weakest categories include Sheet Music (38.8%), Geometric Shapes (40.2%), Technical Blueprints (38.9%), Mathematical Notations (45.9%), and Diagrams (46.8%).

  • Open-source models show similar patterns but at lower absolute performance. For example, BLIP-2 FLAN-T5-XXL achieves 52.1% on Paintings, 62.5% on Landscapes, and 70.0% on Advertisements, but only 25.5% on Chemical Structures, 21.1% on Mathematical Notations, and 28.3% on Geometric Shapes.

  • The pattern across image types reveals a clear divide: image types commonly found in web-scale pretraining data—photographs, paintings, portraits, posters, advertisements—yield substantially higher accuracy across all models. Image types that use specialized, domain-specific notational systems—sheet music, chemical structures, geometric shapes, mathematical notations, technical blueprints—yield dramatically lower accuracy, with many open-source models performing near or below random chance.

The authors explicitly connect this finding to the training data distribution:

"Open-source models demonstrate relatively strong performance in categories like Photos and Paintings, which are more frequently seen during training. However, for less common image categories like Geometric shapes, Music sheets and Chemical structures, all models obtain very low scores (some are close to random guesses). This indicates that the existing models are generalizing poorly towards these image types."

Table 13 provides the complete breakdown: for example, on Chemical Structures (573 samples), Fuyu-8B achieves 25.0%, Qwen-VL-7B achieves 27.2%, InstructBLIP-T5-XXL achieves 27.1%, and BLIP-2 FLAN-T5-XXL achieves 25.5%—all within a few points of random choice (which averages ~22–25% across the benchmark). Even GPT-4V only reaches 50.6% on Chemical Structures, 38.8% on Sheet Music, and 40.2% on Geometric Shapes.


Performance Across Difficulty Levels

Table 3 (Section 4.3) decomposes test-set performance into three difficulty levels: Easy (2,946 questions), Medium (4,917 questions), and Hard (2,637 questions). The results reveal a qualitatively different scaling behavior across model tiers:

For GPT-4V:

  • Easy: 76.1%
  • Medium: 55.6%
  • Hard: 31.2%

The drop from Easy to Hard is approximately 45 percentage points—a massive degradation. Critically, GPT-4V achieves 31.2% on Hard questions versus Random Choice at approximately 23.9%—a gain of only ~7 points over random sampling, compared to a ~52-point gain on Easy questions.

For open-source models (e.g., BLIP-2 FLAN-T5-XXL):

  • Easy: 41.0%
  • Medium: 32.7%
  • Hard: 28.5%

The drop here is only ~13 points—much smaller than GPT-4V's—but this is because the Easy performance is already low. On Hard questions, BLIP-2 FLAN-T5-XXL achieves 28.5%, which is only ~5 points above random—effectively at chance level. The same pattern holds across all reported open-source models: Fuyu-8B at 26.4% on Hard, Qwen-VL-7B at 27.6%, LLaVA-1.5-13B at 26.7%, InstructBLIP-T5-XXL at 29.4%.

The critical insight is that all models converge to near-random performance on Hard questions. The paper notes:

"The further diminishing performance gap in the 'Hard' category across models indicates that as the complexity of tasks increases, the advantage of more advanced models like GPT-4V almost disappears. This might reflect a current limitation in handling expert-level challenging queries even for the most advanced models."

The convergence at the hard level—where GPT-4V (31.2%), BLIP-2 FLAN-T5-XXL (28.5%), and Fuyu-8B (26.4%) all cluster within ~5 points of each other—suggests a fundamental ceiling on current architectures' reasoning capabilities that is not surmounted by additional scale or training data. If scaling alone solved expert reasoning, we would expect GPT-4V's advantage to persist or even widen on harder questions (as it does on many text-only benchmarks). The fact that it narrows instead implies that the hardest MMMU questions probe reasoning capabilities beyond what current approaches can deliver, regardless of model size.


OCR and Captioning Augmentation: A Negative Result

Section 4.2 and Table 2 report a carefully designed negative result: augmenting text-only LLMs with OCR-extracted text or LLaVA-1.5-generated image captions yields essentially no meaningful improvement. On the test set:

  • FLAN-T5-XXL: 31.2% (text only) → 31.9% (+OCR) → 31.9% (+Captions). Net gain: 0.7 percentage points.
  • Vicuna-13B: 31.0% (text only) → 31.9% (+OCR) → 32.7% (+Captions). Net gain: at most 1.7 points.

These gains are negligible—within the range of what could result from random variation or minor differences in answer extraction behavior. For context, the gap between FLAN-T5-XXL (31.9% with OCR) and BLIP-2 FLAN-T5-XXL (34.0% as a true multimodal model) is 2.1 points—also small in absolute terms but representing a different architecture rather than an input augmentation. The key interpretation is not that multimodal models marginally outperform text+OCR models, but that OCR and captioning fail to close any meaningful fraction of the gap between text-only and multimodal capabilities, which is approximately 3-5 percentage points on aggregate but substantially larger in image-dependent disciplines.

The per-discipline breakdown (Tables 5–10) reveals where the augmentation is most (in)effective. For FLAN-T5-XXL on Tech & Engineering, accuracy is 28.3% text-only, 29.7% with OCR, and 28.7% with captions—essentially flat. On Science: 26.7% text-only, 26.2% with OCR, 27.0% with captions—also flat. The augmentation provides a small boost on Humanities & Social Science (44.8% → 50.5% with OCR) and Health & Medicine (32.8% → 32.6% with OCR—actually a slight decrease), but these are disciplines where text-only performance is already relatively high, suggesting that OCR helps when the visual information is primarily textual (labels, captions, OCR-heavy images) and fails when visual information is genuinely non-textual (spatial relationships, visual patterns, domain-specific notation).


Human Expert Comparison

The human expert evaluation (Section 4.1, Table 2 validation set) involved 90 college senior students, 3 per subject, each answering 30 validation questions in their specialty. Results:

  • Worst expert: 76.2%
  • Median expert: 82.6%
  • Best expert: 88.6%

The best human expert outperforms GPT-4V (56.8% validation) by 31.8 percentage points. Even the worst human expert (76.2%) exceeds GPT-4V by 19.4 points. This gap is large and, given the validation set size of 900 questions, cannot be explained by sampling noise. It establishes that MMMU questions are answerable by domain experts at high accuracy—the 88.6% ceiling is not 100%, reflecting that some questions are genuinely difficult even for skilled students—and that current models fall substantially short of expert-level performance.

The inter-expert variance (12.4 points from worst to best) provides a useful band: an AI system achieving "expert-level" performance would need to score within or above this range. GPT-4V at 56.8% is clearly below this band, while GPT-4o at 69.1% is approaching but still ~7 points below the worst human expert. This calibration is the empirical basis for the paper's claim that strong MMMU performance should be a "necessary" criterion for Expert AGI.


Few-Shot Learning Results

Appendix G (Table 14) reports few-shot results for two models that support in-context learning: OpenFlamingo (9B) and Otter. Using the development set (5 examples per subject) as in-context examples:

  • OpenFlamingo: 0-shot: 0.263 → 1-shot: 0.256 → 3-shot: 0.259 → 5-shot: 0.264. Essentially flat across all shot counts—providing examples neither helps nor hurts.
  • Otter: 0-shot: 0.291 → 1-shot: 0.276 → 3-shot: 0.258 → 5-shot: 0.258. Performance actually degrades with more shots, dropping 3.3 points from 0-shot to 5-shot.

The paper interprets this negative result cautiously but clearly:

"This trend suggests that existing open-source models' few-shot learning ability is very weak. And it additionally shows that our data samples might be too hard for these models to understand the underlying patterns or context."

The degradation for Otter is particularly notable—it implies that in-context examples are not merely unhelpful but actively harmful, perhaps because the model attempts to pattern-match on surface features of the examples that do not generalize, or because the examples consume context window space that would otherwise be used for the model's own reasoning. This result also validates the paper's decision to use zero-shot evaluation as the default: few-shot learning does not provide a viable path to improved performance on this benchmark with current open-source models.


Ablation Studies and Robustness Checks

MMMU is primarily a benchmark construction and evaluation paper, not a methods paper with standard architectural ablations. However, several robustness checks and comparative analyses are embedded in the evaluation design:

Validation set as a hyperparameter tuning buffer without test-set contamination: The paper uses the validation set (~900 questions) for prompt engineering and hyperparameter selection, keeping the test set (~10.5K questions) untouched for final reported results. This split ratio (test set is ~12× larger than validation) ensures that prompt optimization has ample signal without consuming the evaluation resource. For models where default prompts are available (GPT-4V, GPT-4, most proprietary models), those defaults are used directly, eliminating any risk of validation-set overfitting. The large test set size (~350 questions per subject on average) provides statistically stable per-subject accuracy estimates.

OCR and captioning augmentation as a negative control: As discussed above, the flat performance of text-only LLMs with OCR and captioning serves as an implicit control study. It demonstrates that the benchmark's images contain information not recoverable through standard vision-to-text pipelines, validating the claim that MMMU tests genuine multimodal understanding rather than text-based reasoning on images that happen to contain text.

Human expert calibration across disciplines: The human evaluation with 90 college seniors (3 per subject) provides per-subject difficulty calibration. Since the validation set covers all 30 subjects at 30 questions each, the human expert performance establishes that the questions are answerable by domain specialists. The fact that the worst expert (76.2%) substantially outperforms GPT-4V (56.8%) confirms that the performance gap is not an artifact of questions being ambiguous, unanswerable, or incorrectly annotated.

Few-shot vs. zero-shot comparison: The Appendix G results (Table 14) serve as an ablation of the in-context learning assumption. The finding that few-shot examples do not improve performance—and in Otter's case, actively degrade it—justifies the paper's decision to default to zero-shot evaluation and suggests that MMMU's difficulty profile is fundamentally different from benchmarks where few-shot learning provides clear benefits.

Multiple model families and scales: The evaluation of 28 open-source models spanning different architectures (BLIP-2, LLaVA, InstructBLIP, CogVLM, Qwen-VL, OpenFlamingo, Fuyu, InternVL, Yi-VL, etc.), scales (~1.6B to ~34B+), and training paradigms (frozen visual encoders + LLM, end-to-end trained, instruction-tuned) provides a form of robustness check on the benchmark's discriminative power. The consistent performance ordering—proprietary models dominating, specific open-source models clustering in tiers—across a diverse set of architectures suggests that MMMU scores reflect genuine capability rather than idiosyncratic fit to a particular model family.

Answer extraction pipeline validation: While not an explicit ablation study, the paper's description of "robust regular expressions and response-processing workflows" that extract answers from long model outputs implies engineering investment in ensuring that models are not penalized for formatting differences. The frequent choice baseline (26.8% validation) serves as a sanity check: models performing near this level are effectively not using question content. The random choice baseline (22.1% validation) establishes the floor for multiple-choice questions.

Per-image-type analysis with sufficient granularity: Appendix F (Table 13) reports performance for 6 models across all 30 image types, with sample counts ranging from 10 (Advertisements) to 3,184 (Diagrams). While the smallest categories (Advertisements at 10 samples, Logos and Branding at 14) are too small for statistically reliable per-type conclusions, the major categories (Diagrams at 3,184, Tables at 2,267, Plots and Charts at 840, Photographs at 770) provide robust signal. The consistent pattern across all models—strong on natural images, weak on specialized notation—persists even within the well-sampled categories (e.g., Chemical Structures at 573 samples, Sheet Music at 335, Geometric Shapes at 336).


Critical Assessment

Claim 1: MMMU is substantially harder than prior multimodal benchmarks.

Strongly supported. The evidence is overwhelming. Models that achieve 80–93% on VQA-v2, ScienceQA-IMG, and RefCOCO (CogVLM: 85%, 92%, 93% respectively; cited in Section 1) achieve 30–56% on MMMU. GPT-4V, the strongest publicly available multimodal model, achieves 55.7%—roughly 30–40 points below its performance on standard VQA benchmarks. Expert humans achieve 88.6%, confirming the questions are answerable. The 900-question validation set and 10.5K-question test set provide large enough samples for reliable aggregate comparisons.

Caveat: While the evidence strongly supports that MMMU is harder on average, the paper does not individually compare MMMU questions to matched-difficulty questions from prior benchmarks. The claim that MMMU is harder "than existing benchmarks" is supported by aggregate score differences, but these could partly reflect differences in answer format, evaluation protocol, or question design philosophy rather than pure difficulty. A direct study where models take both MMMU and a prior benchmark's questions matched on topic and format would isolate the difficulty effect, but such a study is not practical and is not standard for benchmark papers.

Claim 2: Expert human performance substantially exceeds current models.

Supported with qualification. The human expert evaluation design is straightforward but has limitations. The key evidence: best human expert at 88.6% vs. GPT-4V at 56.8% on validation, a ~32-point gap. The worst human at 76.2% still outpaces all models except GPT-4o (69.1% validation). The sample is adequate for calibration (90 experts × 30 questions = 2,700 total human answers).

Qualification: The human experts answered only the validation set (900 questions), not the full test set (10.5K). Their performance on the validation set may not perfectly generalize to the test set if the two differ in difficulty distribution—the paper does not report per-difficulty human performance, so we cannot verify that validation-set difficulty matches test-set difficulty. Additionally, 30 questions per subject is a small sample for measuring individual expert accuracy—a student might score 27/30 (90%) or 23/30 (77%) based partly on which specific questions they receive. The inter-expert variance (12.4 points from worst to best) may partly reflect sampling variance rather than purely expertise differences. A larger per-subject sample for humans (e.g., 50–100 questions) would provide tighter estimates of the expert performance distribution.

Claim 3: Models fail in diagnostically distinct ways (perceptual, knowledge, reasoning).

Supported with careful qualification. The error analysis on 150 GPT-4V errors categorizes failures as perceptual (35%), knowledge (29%), and reasoning (26%). The paper provides concrete examples of each category (Figures 6, 13, 45, 54, etc.) that are genuinely illustrative. The tripartite split—roughly 1/3 each—suggests no single failure mode dominates.

Qualification: The 150-error sample, while carefully annotated, represents only ~3% of GPT-4V's total errors on the test set (GPT-4V gets ~44.3% wrong on 10,500 questions ≈ 4,650 errors). A sample of 150 from 4,650 is small and may not be representative of the full error distribution. The annotation was done by "expert annotators" (Section 5), but inter-annotator agreement is not reported—classifying an error as "reasoning" vs. "knowledge" vs. "perceptual" requires subjective judgment, and different annotators might disagree on borderline cases. The paper acknowledges this boundary fuzziness (noting that "domain-specific perceptual errors occur due to the lack of knowledge" and are classified under knowledge for root cause analysis), but does not quantify how often such borderline cases occur or how they affect the distribution. A more rigorous error analysis would report inter-annotator agreement (e.g., Cohen's kappa) and analyze how sensitive the 35/29/26 split is to annotation guidelines.

Claim 4: Performance varies dramatically across disciplines and image types.

Strongly supported. The evidence from Tables 2 and 5–10 (per-discipline and per-subject results) and Figure 4/Table 13 (per-image-type results) is comprehensive. For GPT-4V, the range across disciplines is 41.7% (Tech & Engineering) to 76.3% (Humanities & Social Science)—a 35-point spread. For image types, the range is 38.8% (Sheet Music) to 100.0% (Advertisements, though n=10). The patterns are consistent across multiple model families, ruling out model-specific artifacts. Per-discipline sample sizes are adequate (947 to 2,784 per discipline). Per-image-type sample sizes vary substantially (10 to 3,184) but the major categories are well-sampled.

Caveat: Some per-image-type results are based on very small samples (Advertisements: 10, Logos and Branding: 14, Landscapes: 16, DNA Sequences: 20). The 100% GPT-4V accuracy on Advertisements and 85.7% on Logos and Branding, while suggestive, are not statistically robust—a single additional error would drop these by 10 and 7 percentage points respectively. The paper does not note this in the main text, though the sample counts are available in Table 13 for attentive readers.

Claim 5: OCR and captioning provide no meaningful improvement, proving MMMU requires genuine multimodal understanding.

Supported with important nuance. The near-zero gains from OCR (+0.7 points for FLAN-T5-XXL, +0.9 points for Vicuna-13B) and captions (+0.7 points for FLAN-T5-XXL, +1.7 points for Vicuna-13B) on the test set are convincing evidence that standard vision-to-text pipelines do not capture the visual information needed for MMMU questions. The per-discipline breakdown confirms that even in text-heavy disciplines, the gains are minimal.

Caveat: The paper tests only one OCR system (MMOCR) and one captioning model (LLaVA-1.5). It is possible that a more advanced OCR system or a stronger captioning model (e.g., GPT-4V itself as a captioner) would yield larger gains—the paper acknowledges this implicitly by not claiming that OCR/captioning is fundamentally insufficient, only that the tested configurations are insufficient. Additionally, the augmentation approach tested is simple (prepend OCR text or caption to the question). More sophisticated approaches—interleaving OCR text with the original image regions, using multiple captions, or applying chain-of-thought reasoning over captions—might extract more value from text-only representations. The conclusion should be read as "standard OCR and captioning as implemented here provide no benefit" rather than "no text-only representation could ever suffice."

Claim 6: All models converge to near-random performance on Hard questions.

Supported but interpret with care. Table 3 shows GPT-4V at 31.2% on Hard vs. 23.9% random—a ~7-point gap. BLIP-2 FLAN-T5-XXL is at 28.5%—a ~5-point gap. These are positive (above chance) but small, and the model-to-model variance on Hard questions (26.4% to 31.2%) is much narrower than on Easy (28.9% to 76.1%).

Caveat: "Near-random" is a judgment call. A 7-point gap on a 2,637-question subset is statistically significant (a binomial test would confirm this). The convergence of model scores does not necessarily mean models are "failing" in the same way—GPT-4V might get a different subset of Hard questions correct than BLIP-2, so their similar aggregate scores could mask complementary strengths. Additionally, "Hard" is defined by human annotator judgment, not by model performance. It is possible that some Hard questions test obscure factual knowledge that no model happens to have memorized, rather than testing reasoning per se—a model could in principle be a perfect reasoner and still fail these questions due to knowledge gaps. The paper's error analysis partly addresses this by identifying knowledge gaps as 29% of errors, but does not break this down by difficulty level.

Missing Experiments and Analyses

Several additional evaluations would have strengthened the paper:

  1. Inter-annotator agreement for error analysis. As noted above, reporting agreement metrics (e.g., Cohen's kappa) for the 150-error annotation would strengthen confidence in the error taxonomy and the reported proportions. Without this, the 35/29/26 split might partly reflect annotator biases rather than true underlying error distributions.

  2. Per-difficulty human expert performance. Reporting how human experts perform on Easy, Medium, and Hard questions would calibrate whether the difficulty labels align with human experience—do experts also find "Hard" questions difficult? If experts achieve 95% on Easy and 70% on Hard, that's a very different calibration than if they achieve 95% across the board.

  3. Model calibration analysis. Beyond accuracy, reporting whether models are calibrated (do their confidence scores predict correctness?) and whether calibration degrades on Hard questions would provide additional diagnostic signal about model failure modes.

  4. Error overlap analysis. Do GPT-4V and open-source models fail on the same Hard questions, or do they fail on different ones? The convergence of aggregate scores near random could mask complementary strengths—if GPT-4V solves a different subset of Hard questions than BLIP-2, an ensemble might substantially outperform either individually. This analysis would inform whether the "ceiling" on Hard questions is fundamental or reflects different models hitting different knowledge/perception/reasoning gaps.

  5. Fine-grained per-subfield results. The paper reports per-subject results (Appendix B) but not per-subfield. With 183 subfields, the per-subfield sample sizes are small (average ~57 questions), but even coarse reporting—which subfields are models strongest/weakest on—would provide more actionable diagnostic information for model developers.

6. Limitations and Trade-offs

6.1 College-Level Exams Are an Incomplete Proxy for Expert AGI

The assumption or constraint. The paper anchors its benchmark to college-level exams, quizzes, and textbooks, explicitly positioning strong MMMU performance as "a necessary criterion for an Expert AGI system." Crucially, the authors themselves acknowledge the limitation of this framing. Section 6 states:

"The focus on college-level subjects might not be sufficient for testing Expert AGI. However, we argue that strong performance on this benchmark should be a necessary criterion for an Expert AGI system."

The core issue is a potential construct validity gap: college exams measure a specific slice of expert capability—the ability to answer curated, well-defined problems under exam conditions—while real-world expertise involves open-ended problem formulation, integration of multiple knowledge sources, handling of ambiguity, and adaptive reasoning in novel situations. A system that achieves human-expert accuracy on MMMU might still fail at the kind of unstructured, context-rich reasoning that skilled professionals routinely perform.

The consequence. The benchmark risks creating a false sense of progress toward Expert AGI. If models achieve high MMMU scores but cannot transfer that capability to real-world expert tasks (diagnosing a patient when the presentation is atypical, designing an engineering solution when constraints are partially specified, evaluating a historical argument when evidence is conflicting), then MMMU would be measuring exam-taking ability rather than expert understanding. The "necessary but not sufficient" framing partially hedges against this, but does not guard against the risk that the research community—and the public discourse the paper aims to influence (Section 1: "significant risks of job displacement and economic disruption")—over-interprets benchmark progress as genuine expert capability.

What evidence exists in the paper. The paper does not provide evidence that MMMU performance correlates with real-world expert capability. The human validation establishes that college seniors perform well on the benchmark, confirming that the questions are answerable by domain specialists. But no experiment tests whether MMMU accuracy predicts performance on more ecologically valid expert tasks. The paper's error analysis (Section 5) identifies reasoning errors, knowledge gaps, and perceptual failures, but does not distinguish between errors that would also occur in real-world expert practice and errors that are artifacts of the multiple-choice exam format.

Mitigation status. The paper addresses this limitation transparently but without resolution. The "necessary but not sufficient" framing in the conclusion is an explicit acknowledgment, and the connection to Morris et al.'s AGI taxonomy provides theoretical grounding for why college exams are a reasonable starting point. However, the paper does not suggest concrete follow-up work for validating MMMU's ecological validity—for example, correlating benchmark performance with expert performance on case-based assessments, open-ended design tasks, or professional certification exams that involve more realistic problem formats. The limitation is effectively flagged but left for future work to address.


6.2 Human Expert Evaluation Uses Only the Validation Set with a Small Per-Subject Sample

The assumption or constraint. The human expert calibration that establishes the benchmark's performance ceiling involves 90 college seniors, each answering only 30 validation questions in their specialty (Section 4.1). This yields 3 experts × 30 subject = 90 total participants, with each expert contributing answers to only the ~30 validation questions in their subject. The validation set totals approximately 900 questions. The paper reports expert performance on the validation set (best: 88.6%, median: 82.6%, worst: 76.2%) but provides no human expert results on the 10,500-question test set, nor does it break down human performance by difficulty level (Easy/Medium/Hard), by question type (multiple-choice vs. open-ended), or by individual subject.

The consequence. Two distinct validity concerns arise:

First, the test set may have a different difficulty distribution than the validation set. If the validation set is systematically easier or harder than the test set, the human ceiling estimated from the validation set does not accurately represent the ceiling for the test set—which is where all primary model results are reported. This means we cannot confidently say, for example, that GPT-4V's 55.7% test accuracy is "31.8 points below the best human expert" because the best human expert score (88.6%) was measured on a different set of questions. The gap might be larger or smaller on the test set.

Second, the 30-questions-per-subject sample is too small for reliable per-subject human accuracy estimates. A student scoring 27/30 (90%) vs. 24/30 (80%) may differ by 10 percentage points due to which specific 30 questions they received rather than genuine expertise differences. The reported inter-expert spread of 12.4 percentage points (76.2% to 88.6%) partly reflects this sampling variance. The paper cannot determine whether this spread represents true individual differences, question sampling noise, or both.

What evidence exists in the paper. The paper reports only aggregate expert results on the validation set (Table 2). Per-subject, per-difficulty, and per-question-type human results are absent. The paper does not discuss the sampling variance of the 30-question estimate, does not report confidence intervals for expert scores, and does not provide an analysis of whether validation-set difficulty matches test-set difficulty. The overall test-set difficulty distribution (28% Easy, 45% Medium, 27% Hard; Table 1) is known, but the validation-set difficulty distribution is not separately reported, preventing a direct comparison.

Mitigation status. The paper does not address this limitation. The human evaluation design is described straightforwardly without caveats about sample size or generalizability. This is a missed opportunity: a larger human evaluation (e.g., 50–100 questions per subject) or a design where human experts answer questions drawn from both validation and test sets would provide a more reliable ceiling estimate. The absence of per-difficulty human expert scores is particularly notable given that the paper's diagnostic claims about model capability degradation on Hard questions (Section 4.3) rely on an implicit comparison to what humans would achieve on those same difficulty levels—a comparison that is never empirically established.


6.3 Manual Curation Carries Scalability and Bias Risks

The assumption or constraint. MMMU's data collection relies entirely on manual curation: "over 50 university students, including co-authors, specializing in these majors as annotators" collect questions from online sources, textbooks, and lecture materials, creating new questions "based on their expertise where necessary" (Section 3.2). The 11.5K questions were manually sourced, formatted, and quality-controlled through a pipeline of deduplication, format verification, and difficulty categorization. The paper acknowledges potential biases explicitly:

"MMMU, like any benchmark, has limitations despite its comprehensive nature. The manual curation process may carry biases" (Section 6).

The consequence. Manual curation introduces several risks that differ from automated or crowdsourced collection at larger scale:

First, scale limits. 11.5K questions is substantial for a manually curated benchmark but small relative to the diversity of expert knowledge. With 183 subfields, the average is approximately 63 questions per subfield, and many subfields likely have far fewer. A model might fail on a particular subfield not because it lacks that knowledge but because the few sampled questions happen to probe knowledge it has not memorized—or succeed on a subfield because the questions coincidentally align with its training data. Manual curation cannot economically scale to the tens or hundreds of thousands of questions that would provide robust per-subfield coverage.

Second, annotator bias. The 50 student annotators, despite being domain specialists, bring individual perspectives that shape question selection. Questions that one annotator considers "college-level" another might consider too easy or too obscure. The paper's difficulty filtering removed approximately 10% of questions as "very easy" (Section 3.2), but this filtering itself reflects annotator judgment, not an objective standard. If annotators systematically select questions that are more "textbook-like" (with clean answers, standard problem formats) and avoid questions that are ambiguous or require creative reasoning, the benchmark would inadvertently over-represent well-structured problems and under-represent the messier reasoning that characterizes real expertise.

Third, copyright and source bias. The annotation protocol instructs collectors to "strictly adhere to copyright and licensing regulations" and avoid data from "sites prohibiting copy and redistribution" (Appendix H.1). This constraint—while ethically necessary—biases question selection toward sources that permit redistribution, which may not be representative of college-level exam content in aggregate. High-quality proprietary textbooks and exam banks that forbid redistribution are excluded; openly licensed or freely available materials (which may be older, less rigorously curated, or differently distributed across disciplines) are over-represented.

What evidence exists in the paper. The paper does not provide an analysis of annotator agreement during difficulty categorization, of the source distribution (what fraction from textbooks vs. online resources vs. annotator-authored), or of the demographic/institutional distribution of annotators. The quality control pipeline description (Section 3.2) focuses on deduplication and format checking rather than quantifying collection biases. The paper includes interleaved images at the beginning, middle, and end of questions (Table 1), suggesting some intentional diversity in question formatting, but this is not a measure of content diversity.

Mitigation status. The paper acknowledges the bias risk in the conclusion but does not attempt to measure or mitigate it. Compared to benchmarks constructed through automated data mining (which can achieve scale at the cost of noise and potential contamination), MMMU prioritizes quality and controllability over scale—a deliberate trade-off. However, without quantification of selection biases (e.g., comparing the difficulty distribution of MMMU questions to a random sample from the source textbooks, or measuring inter-annotator agreement on difficulty and subject classification), the magnitude and direction of manual curation biases remain unknown. The limitation is flagged but unaddressed.


6.4 The Multiple-Choice Format Dominance Limits Diagnostic Power

The assumption or constraint. MMMU consists of 94.03% multiple-choice questions (10,861 out of 11,550), with only 5.97% open-ended (Table 1). The paper defends this choice pragmatically:

"MMMU combines multiple-choice questions with concise open-ended questions, enabling the assessment of diverse subjects while addressing the challenges associated with evaluating open-ended responses" (Section 6).

The underlying assumption is that multiple-choice questions can adequately probe expert-level understanding and reasoning—that selecting the correct option from a set demonstrates the same cognitive capability as generating the correct answer from scratch.

The consequence. The multiple-choice format imposes specific limitations on what the benchmark can measure:

First, guessing and elimination strategies. A model can answer a multiple-choice question correctly without fully solving it—by eliminating obviously wrong options, exploiting answer distribution patterns, or making educated guesses based on surface-level cues. The frequent choice baseline (26.8% validation, 25.8% test) and random choice baseline (22.1% validation, 23.9% test) establish that simple strategies outperform random guessing, but do not capture more sophisticated elimination or pattern-matching. A model achieving 56% accuracy might genuinely solve only a fraction of those questions while guessing or eliminating on the remainder. The benchmark cannot distinguish between a fully reasoned correct answer and a correct guess based on partial elimination.

Second, inability to assess process and reasoning depth. For multiple-choice questions, MMMU evaluates only whether the final answer matches the ground truth—not whether the reasoning that led to it was correct. The error analysis (Section 5) manually examines 150 GPT-4V responses and finds 26% reasoning errors, but this requires human annotation and cannot be automated. For the remaining 10,350 test questions, we do not know whether correct answers reflect sound reasoning or lucky guesses, or whether incorrect answers reflect reasoning failures or knowledge gaps. This substantially limits the diagnostic value of aggregate accuracy scores.

Third, format mismatch with real expertise. Skilled professionals rarely encounter multiple-choice versions of their real-world tasks. A radiologist does not choose among four diagnostic options—they generate a differential diagnosis from their knowledge base. A mechanical engineer does not select among four design configurations—they derive specifications from principles and constraints. The multiple-choice format abstracts away the generation component of expertise—the ability to produce the correct answer from the space of all possible answers—and replaces it with a recognition task (identifying the correct answer among presented options). Recognition is typically easier than generation, meaning MMMU may overestimate models' expert capabilities.

What evidence exists in the paper. The paper does not report comparative performance on multiple-choice vs. open-ended subsets (the 689 open-ended questions may be too few for reliable comparison). The per-subject breakdown tables (Appendix B) do not separate these question types. The human expert evaluation used only the validation set, which includes both multiple-choice and open-ended questions, but does not report which type contributed to human errors. The error analysis (150 GPT-4V errors) might incidentally capture format effects, but the paper does not analyze whether error types differ between multiple-choice and open-ended questions.

Mitigation status. The paper acknowledges the trade-off in the conclusion—"addressing the challenges associated with evaluating open-ended responses"—but does not attempt to quantify the gap between multiple-choice and open-ended performance or to include more open-ended questions. The 5.97% open-ended component is too small to serve as a meaningful calibration of the multiple-choice score inflation. This is a deliberate practical choice (automated evaluation of open-ended expert responses is genuinely difficult) rather than an oversight, but it means the benchmark's primary metric likely overestimates true expert-level understanding.


6.5 Data Contamination Risk Is Acknowledged but Unmeasured

The assumption or constraint. MMMU collects questions from publicly available online sources, textbooks, and lecture materials. Foundation models—particularly the proprietary models at the top of the leaderboard—are trained on large-scale web corpora that may include these exact sources. The paper acknowledges this concern in the annotation protocol (Appendix H.7):

"In the construction of benchmarks for evaluating foundation models, it is essential to consider the risk of data contamination. To address this, annotators should be tasked with carefully selecting questions that go beyond straightforward queries with easily accessible answers. Instead, the focus should be on questions whose answers are tucked away in less obvious locations, such as in separate documents or hidden in the concluding sections of extensive textbooks."

The mitigation strategy relies on annotator judgment—selecting questions whose answers are not trivially retrievable from the question text alone—rather than on empirical measurement of contamination.

The consequence. Contamination creates a fundamental confound in interpreting benchmark results. If GPT-4V achieves 55.7% accuracy on MMMU, it is impossible to determine what fraction of correct answers reflect genuine multimodal expert reasoning versus memorization of specific questions seen during pretraining. The confound is asymmetric: contamination inflates scores but cannot be detected from benchmark results alone. A model trained on MMMU's source materials might answer questions correctly through pattern completion rather than understanding, yet appear to possess expert-level reasoning capabilities.

The authors' mitigation strategy—selecting questions with less accessible answers—addresses a subset of contamination risk (questions where the answer appears alongside the question in the source) but does not address the case where the model's pretraining data includes the source textbook in full. A language model trained on a textbook containing both the question and its answer (even if separated by chapters) could learn the association without understanding the underlying concepts. The mitigation also relies on annotators' ability to judge answer accessibility, which is subjective and may not align with what models actually memorize.

What evidence exists in the paper. The paper provides no empirical measurement of contamination. There is no analysis of how many MMMU questions can be found verbatim in web search results, no measurement of whether models exhibit differential performance on questions from sources that are more vs. less likely to be in their training data, and no experiment with a model trained on a known-closed corpus to establish a contamination-free baseline. The human expert results (88.6%) establish that the questions are answerable without memorization, but do not address whether models might be exploiting memorization shortcuts rather than reasoning.

Mitigation status. The paper partially addresses contamination through the question selection guidelines (Appendix H.7) and the manual curation process, but does not empirically quantify the residual risk. This is a common limitation in benchmark papers—comprehensive contamination analysis is expensive and for proprietary models like GPT-4V, impossible from the outside since training data is not disclosed. However, the paper could have included analyses that are feasible without access to training data: measuring question overlap with web search results, testing whether model confidence/calibration differs between likely-contaminated and likely-clean subsets, or designing a subset of questions that are deliberately novel (author-created rather than sourced) to serve as a contamination-free reference point. None of these analyses are performed.


6.6 Single Human Language and Cultural Context

The assumption or constraint. All MMMU questions are in English and are sourced from materials aligned with the U.S. and Western educational context—the 50 college student annotators are recruited from "university students... specializing in these majors" and collect questions from "major textbooks and online resources" (Section 3.2). The disciplines, subjects, and subfields reflect the structure of Western university curricula. The paper does not explicitly discuss the geographic, linguistic, or cultural scope of the benchmark, nor does it acknowledge this as a limitation.

The consequence. The benchmark implicitly equates "expert-level knowledge" with the knowledge canon of English-language, Western university education. This has two downstream effects:

First, coverage bias. Subjects and subfields that are prominent in other educational traditions but less emphasized in Western curricula may be underrepresented or absent. For example, Traditional Chinese Medicine, Ayurvedic medicine, or non-Western art history traditions are unlikely to appear in MMMU at all. A model that performs well on MMMU is being tested specifically on Western expert knowledge, not global expert knowledge—but the Expert AGI framing (Section 1) does not acknowledge this scope restriction.

Second, evaluation bias. Questions assume familiarity with cultural references, notation conventions, and problem framings that are standard in English-language education. A model trained on multilingual data might possess equivalent expertise but struggle with culturally specific problem framings or notational conventions. Conversely, a model trained predominantly on English-language data might benefit from cultural alignment with the benchmark, making cross-lingual and cross-cultural comparisons impossible.

What evidence exists in the paper. None. The paper does not report the geographic distribution of annotators, the language distribution of source materials, or any analysis of cultural representation in the question content. The subject taxonomy (Figure 7, Appendix Table 12) reflects standard Western academic disciplines. The paper's claim to "comprehensiveness" (Section 3.1) and "massive multi-discipline" coverage must be understood as comprehensive within the Western academic framework, but the paper does not explicitly bound the claim this way.

Mitigation status. The paper does not address this limitation. This is a significant omission given the Expert AGI framing: if the goal is to measure progress toward systems that can substitute for skilled adults globally, the benchmark should either be explicitly scoped to a particular educational tradition or systematically include cross-cultural coverage. The paper does neither. Compared to benchmarks like MMLU (which has received similar critiques for English-language and Western bias), MMMU could have distinguished itself by including questions from diverse educational systems or by explicitly acknowledging its cultural scope—but it does not. This limitation is particularly consequential because of the paper's ambition to inform public discourse about AGI and economic disruption (Section 1): a benchmark that measures only Western expert knowledge cannot support global claims about machine substitution for human labor.

7. Implications and Future Directions

How This Work Changes the Landscape

This work establishes a diagnostic framework for multimodal intelligence, not another leaderboard benchmark. MMMU's primary contribution to the research landscape is methodological rather than merely empirical: it creates a taxonomy of failure modes (perceptual, knowledge, reasoning) tied to concrete architectural bottlenecks, and it anchors evaluation difficulty to a fixed human-expert standard that does not drift as models improve. This is fundamentally different from benchmarks that aggregate scores across skill categories (MMBench, MM-Vet) or report coarse-grained difficulty bands without explaining why models fail where they do.

The conceptual shift is subtle but consequential. Prior multimodal benchmarks implicitly assumed that aggregate accuracy tracks general multimodal capability. MMMU demonstrates that aggregate accuracy on a sufficiently hard benchmark instead reflects the intersection of three separable capabilities—perception, knowledge, and reasoning—that degrade at different rates across different image types and difficulty levels. The evidence: models achieve 64–76% on photographs but 39–41% on geometric shapes (Figure 4, Table 13), and the performance gap between GPT-4V and open-source models narrows from ~40 points on Easy questions to ~3 points on Hard questions (Table 3). These are not subtle correlations—they are qualitative reversals that would be invisible without per-image-type and per-difficulty decomposition.

The "necessary but not sufficient" positioning creates a modular template for AGI benchmarking. Rather than competing for the title of "the AGI test"—a framing that has damaged the credibility of several prior benchmarks—MMMU connects to an external taxonomy (Morris et al.'s Levels of AGI), makes a specific, bounded claim (expert-level multimodal understanding is one necessary component of Expert AGI), and provides empirical calibration (human experts at 88.6%) against which progress can be measured. This modular framing makes it possible for future benchmarks to claim "necessary but not sufficient" status for other dimensions of general intelligence (e.g., common sense reasoning, open-ended creativity, social intelligence) without each benchmark implicitly claiming to be comprehensive.

The practical consequence of this reframing: the dialogue around benchmark scores becomes more honest. A model achieving 70% on MMMU is not "70% of the way to Expert AGI"—it has satisfied one necessary condition at 70% of the human-expert ceiling, with other necessary conditions (measured by other benchmarks) remaining unknown. This modular decomposition resists the pressure to over-interpret benchmark scores that has plagued prior AGI-adjacent benchmarks.

The error taxonomy reconciles contradictory intuitions about multimodal model failures. Before MMMU, the field lacked a shared vocabulary for diagnosing why models fail on hard multimodal problems. Different papers attributed failures to different causes—visual grounding (Liu et al., 2023), hallucination (Li et al., 2023), reasoning limitations (Lu et al., 2022), or domain knowledge gaps (Hendrycks et al., 2021)—but without a unified framework, it was impossible to tell whether these represented genuinely different failure modes or different descriptions of the same underlying problem. MMMU's tripartite taxonomy (perception: 35%, knowledge: 29%, reasoning: 26%) provides evidence that all three are simultaneously operative and roughly equally important, resolving the debate in favor of "all of the above" and redirecting research attention from identifying a single bottleneck to improving all three axes.

The negative result on OCR/captioning augmentation closes a plausible but incorrect shortcut. The finding that standard vision-to-text pipelines (MMOCR, LLaVA-1.5 captions) add essentially zero performance to text-only LLMs on MMMU—gains of 0.7–1.7 percentage points (Table 2)—establishes that expert-level multimodal questions contain visual information that cannot be adequately captured by current off-the-shelf image-to-text conversion. This result has direct practical implications: it suggests that research investment in better vision-to-text pipelines (as a cheap way to make text-only LLMs multimodal) will not scale to expert-level tasks, and that genuine multimodal architectures—models that jointly encode vision and language rather than converting vision to language first—are necessary for progress on MMMU-level challenges. This redirects resources away from captioning/OCR augmentation and toward multimodal pretraining and fusion.

The image-type performance breakdown reorients visual generalization research. MMMU's 30 image types, with performance reported per type across six models (Table 13, Appendix F), creates a concrete benchmark for visual generalization across representations. The finding that all models, including GPT-4V, perform substantially worse on specialized notation systems (sheet music, chemical structures, geometric shapes, mathematical notations) than on natural images (photographs, paintings, advertisements) provides a clear signal about where to invest: improving model performance on underrepresented image types requires either more diverse pretraining data or architectures that generalize better from common to rare visual formats—not simply more scale applied to the existing natural-image-heavy training distribution.

The benchmark makes incremental progress on research directions that become less attractive after MMMU. The finding that current few-shot learning mechanisms provide no benefit (OpenFlamingo flat, Otter degrading with more shots; Table 14, Appendix G) and that model performance converges near random on Hard questions regardless of scale or architecture (Table 3) suggests that scaling existing architectures and training paradigms alone will not close the gap to expert human performance. This makes "more of the same" approaches—larger models with the same multimodal fusion mechanisms, longer training on the same data distributions—less attractive as a sole strategy, and increases the appeal of research on novel multimodal reasoning architectures, domain-knowledge injection, and test-time reasoning strategies tailored to expert-level problems.


Follow-Up Research This Work Enables

1. Difficulty estimation from the question text alone, without requiring model sampling. The paper's difficulty categories are assigned by human annotators based on domain expertise (Section 3.2), but for the benchmark to enable adaptive evaluation—where testing resources are concentrated on questions that best differentiate model capabilities—it would be valuable to predict difficulty from the question text and image content before model inference. A concrete experiment: train a classifier on MMMU's training split (or on the validation set) that takes the question text, image type, subject, and subfield as input and predicts the human-assigned difficulty level (Easy/Medium/Hard). The output would be a per-question difficulty score that could be used to stratify evaluation, weight questions in aggregate metrics, or select maximally informative question subsets for efficient model comparison. MMMU makes this newly tractable because it provides 11.5K questions with human-annotated difficulty labels covering 30 subjects—enough data to train a non-trivial difficulty predictor. A strong follow-up would measure whether such a predictor generalizes across subjects (e.g., trained on Science and evaluated on Health & Medicine) and whether its difficulty rankings align with actual model performance (do models indeed find "predicted-hard" questions more difficult?).

2. Training curriculum studies that use MMMU's difficulty levels and image types to investigate how expert multimodal capabilities emerge. MMMU provides three orthogonal axes of question variation (subject, image type, difficulty) that could be used to construct training curricula for multimodal models. A concrete experiment: pretrain or fine-tune a multimodal model on MMMU training data (or MMMU-style synthetic data) in different orders—curriculum A starts with Easy questions on natural images (photographs, paintings) and gradually introduces Hard questions on specialized notation (chemical structures, sheet music); curriculum B does the reverse; curriculum C interleaves all difficulty levels and image types from the start. The hypothesis is that curriculum A might yield better generalization to novel subjects by first establishing basic multimodal reasoning skills before introducing complex visual notation. MMMU's 11.5K questions across 30 subjects and 30 image types provide sufficient diversity to test this hypothesis at scale. A strong follow-up would measure not just final MMMU accuracy but transfer to held-out subjects, providing evidence about whether structured curricula accelerate expert-level multimodal understanding—and whether the difficulty and image-type structure that MMMU makes explicit is actually the right structure for curriculum design.

3. Diagnostic benchmarking with the error taxonomy automated at scale. Section 5's error analysis on 150 GPT-4V responses manually categorizes errors as perceptual, knowledge, or reasoning. One could scale this diagnostic approach by training a classifier to automatically label model errors into the three MMMU categories, using the 150 human-annotated errors as training data. The classifier would take as input the question, the model's response, the ground-truth answer, and optionally the gold explanation (for the 17.62% of questions that have them), and predict the primary error category. The output would be a per-model, per-subject, per-difficulty error profile—e.g., "Model X on Hard Chemistry questions: 45% perceptual errors, 30% knowledge gaps, 25% reasoning errors." This would transform MMMU from a score-reporting benchmark into a diagnostic instrument that tells model developers not just how well their model performs, but why it fails. MMMU makes this newly tractable because it provides (a) the initial manually annotated error dataset, (b) sufficient question diversity for training a generalizable classifier, and (c) ground-truth explanations for a meaningful subset of questions that provide reasoning traces against which model outputs can be compared. A strong follow-up would validate the automated classifier against additional human annotations and demonstrate that models with different architectures (e.g., BLIP-2 vs. LLaVA vs. CogVLM) exhibit systematically different error profiles—which would confirm that the error taxonomy captures genuine architectural differences rather than random noise.

4. Cross-lingual and cross-cultural extensions to test whether expert multimodal knowledge is language- and culture-bound. MMMU's English-language, Western-curriculum framing (Section 6, Limitations) creates a natural follow-up: construct a parallel benchmark with the same disciplines, difficulty levels, and image types, but with questions sourced from non-English educational systems (e.g., Chinese Gaokao, Indian JEE, German Abitur, French Baccalauréat) and translated or adapted for cross-lingual evaluation. The concrete question: do models that perform well on English-language MMMU also perform well on the same subjects when the questions, notation conventions, and cultural references reflect different educational traditions? MMMU provides the blueprint—subject taxonomy, difficulty categories, image types, evaluation protocol, answer extraction pipeline—that can be replicated with different source materials. A strong follow-up would measure the performance gap between English-language MMMU and a non-English parallel version for the same models, quantifying whether "expert multimodal understanding" as measured by MMMU is genuinely general or is partly an artifact of language and cultural alignment between the benchmark and model training data. This would directly address one of MMMU's most significant unexamined limitations.

5. Combining MMMU's diagnostic dimensions with test-time compute scaling to probe whether expert reasoning can be amplified at inference time. The paper establishes that all models, including GPT-4V, perform poorly on Hard questions across all subjects—converging to near-random performance. An open question is whether additional test-time compute (chain-of-thought, self-consistency, or tree-search over reasoning paths) can move the needle on Hard MMMU questions, or whether these questions probe fundamental capability gaps that cannot be compensated for at inference time. This is analogous to the finding from the "compute-optimal test-time scaling" paper (attached reference example): test-time compute helps on problems within a model's capability range but provides zero benefit on problems where the model's pass@1 is near zero. MMMU provides the diagnostic infrastructure (per-difficulty breakdown, per-subject accuracy, image-type stratification) to test this hypothesis precisely. A concrete experiment: evaluate GPT-4V or a strong open-source model on MMMU with varying test-time compute budgets—majority voting at N={1, 4, 16, 64}, chain-of-thought prompting with and without self-consistency, and PRM-guided beam search if a process reward model can be trained on MMMU's 2,035 explained questions. The key measurement: does the performance gap between Easy and Hard questions narrow with additional compute, or does Hard-question performance remain flat regardless of budget? If it remains flat, this would provide strong evidence that MMMU's Hard questions indeed probe fundamental capability ceilings that require pretraining improvements rather than inference-time strategies—a finding with direct implications for research prioritization.

6. Stress-testing verifier robustness for multimodal reasoning chains using MMMU's expert-level content. The observation in Section 5 that GPT-4V commits reasoning errors (26% of failures) even when it correctly perceives visual information and recalls relevant knowledge motivates a specific line of verification research. Can an external verifier—trained to score the correctness of reasoning steps in multimodal expert solutions—reliably detect when a model's reasoning goes off-track? MMMU enables this research because (a) it provides 2,035 questions with gold explanations that can serve as positive examples of correct reasoning chains, (b) it provides 150 manually annotated GPT-4V errors with categorized failure modes that can serve as negative examples of specific reasoning failures, and (c) its diversity of subjects and image types ensures any trained verifier must generalize across expert domains rather than specializing in one subject's reasoning patterns. A concrete experiment: fine-tune a multimodal model (or a separate verifier head) to score intermediate reasoning steps for correctness, trained on MMMU's explained questions as positive examples and on synthetically generated reasoning errors (e.g., correct perception + incorrect knowledge application, correct knowledge + flawed calculation) as negative examples. Evaluate whether this verifier can detect reasoning errors in held-out subjects and at different difficulty levels, and whether verifier-guided rejection sampling (accepting only solutions whose reasoning steps pass the verifier's threshold) improves MMMU accuracy over standard best-of-N selection. This would test whether MMMU's error taxonomy can be operationalized as a training signal for more robust multimodal reasoning.


Practical Applications and Downstream Use Cases

1. Model selection for domain-specific deployment. Organizations deploying multimodal models in specialized domains (medical imaging analysis, engineering diagram interpretation, architectural plan review) currently lack principled criteria for choosing among available models. MMMU's per-discipline and per-image-type breakdown (Tables 5–10, Table 13) provides actionable guidance. For example: a radiology AI startup choosing between GPT-4V and a fine-tuned open-source model for chest X-ray interpretation would see from Table 2 that GPT-4V achieves 63.5% on Health & Medicine overall and 63.6% on Pathological Images specifically (Table 13), while the best open-source model at submission time (BLIP-2 FLAN-T5-XXL) achieves 33.7% and 35.6%, respectively—a ~28-point gap that likely justifies the proprietary model's cost for this use case. Conversely, for an art history education application where models need to analyze paintings and sculptures, the gap narrows: GPT-4V at 75.9% on Paintings vs. BLIP-2 FLAN-T5-XXL at 52.1%—meaningful but potentially not wide enough to justify the cost differential if the open-source model can be fine-tuned on domain-specific art data. MMMU's granular breakdowns make these deployment decisions empirically grounded rather than based on aggregate leaderboard rankings that obscure domain-specific capability differences.

2. Benchmark-driven data collection for multimodal pretraining. The finding that all models perform poorly on specialized image types—sheet music (GPT-4V: 38.8%), geometric shapes (40.2%), chemical structures (50.6%), technical blueprints (38.9%)—provides a directly actionable signal for pretraining data curation. Teams building next-generation multimodal models can use MMMU's image-type performance breakdown as a gap analysis: identify the image types where current models most underperform relative to human experts, and prioritize collecting or generating pretraining data in those categories. For instance, the ~40-point gap between GPT-4V's performance on photographs (64.2%) and sheet music (38.8%) strongly suggests that sheet music is dramatically underrepresented in web-scale pretraining data—a gap that could be addressed by targeted crawling of music notation repositories, IMSLP (International Music Score Library Project), or synthetic generation of sheet music images paired with their musical interpretations. MMMU provides the measurement instrument that makes this gap analysis possible; without it, data collection prioritization relies on intuition about which visual domains are underrepresented rather than empirical measurement of model failure.

3. Progress monitoring for Expert AGI as a societal early-warning signal. The paper's explicit connection to Morris et al.'s taxonomy and its framing of MMMU as measuring a necessary condition for Expert AGI creates a concrete use case for policymakers and industry observers monitoring AI progress. As model performance on MMMU improves—tracked via the public leaderboard—the gap between best-model accuracy and the human expert ceiling (88.6%) provides a quantifiable metric for how close current systems are to expert-level multimodal reasoning. When (and if) a model reaches within the human expert inter-quartile range (approximately 77–89%), that would signal that the "necessary condition" for Expert AGI measured by MMMU has been substantially satisfied—triggering closer attention to other necessary conditions (measured by other benchmarks) and to the real-world deployment implications the paper discusses (job displacement, economic disruption). The public leaderboard makes this monitoring transparent and contestable: different stakeholders can track different models, different disciplines, or different difficulty levels depending on their specific concerns. MMMU's large test set (10.5K questions) means that even substantial score improvements (e.g., 5–10 percentage points) represent genuine capability gains rather than test-set overfitting, making the leaderboard a credible signal rather than a gameable metric.

4. Diagnostic evaluation during iterative model development. For researchers building and improving multimodal models, MMMU's structure enables fine-grained development feedback that aggregate benchmarks cannot provide. A team training a new visual encoder can evaluate on MMMU's 30 image types separately and identify immediately whether their architectural change improved performance on the specific image types they were targeting (e.g., geometric shapes, chemical structures) without degrading performance on others. A team fine-tuning a model on scientific literature can check whether per-subject accuracy on Science subjects (Biology, Chemistry, Physics, Math, Geography) improves relative to a baseline, while monitoring that performance on non-science disciplines does not regress. The three-way error taxonomy (perceptual, knowledge, reasoning) provides a diagnostic vocabulary for interpreting development results: if a model's Chemistry accuracy improves but primarily by reducing perceptual errors (correctly reading chemical structures) rather than reducing reasoning errors (correctly applying reaction mechanisms), the development team knows their visual encoder improvements are working but their reasoning architecture needs separate attention. This granularity transforms MMMU from a final evaluation checkpoint (run once, report a score) into a development instrument (run frequently, diagnose specific improvements and regressions) that accelerates the model iteration cycle.