ArXiv: 2501.14249

🎯 Pitch

Even the most advanced AI models score only 2.7% to 13.4% on a new expert-level benchmark—and they are dangerously overconfident about their wrong answers. This reveals that frontier systems fundamentally cannot recognize the limits of their own knowledge when faced with questions at the edge of human expertise.


1. Executive Summary

This paper introduces Humanity's Last Exam (HLE), a multi-modal benchmark of 2,500 closed-ended academic questions designed to be "the final closed-ended academic benchmark of its kind with broad subject coverage" by targeting the frontier of human knowledge across mathematics, humanities, and natural sciences. Developed by nearly 1,000 subject-matter experts through a multi-stage review process — including LLM difficulty filtering (questions are rejected if frontier models answer correctly) and iterative expert peer review — the benchmark reveals that state-of-the-art LLMs achieve only 2.7% to 13.4% accuracy on the full dataset, with O3-MINI (HIGH) achieving the highest text-only accuracy at 13.4% while all models exhibit severe miscalibration (RMS calibration errors above 70%), establishing that current models fail to recognize when questions exceed their capabilities and frequently provide incorrect answers with high confidence.

2. Context and Motivation

The Benchmark Saturation Crisis

The fundamental problem this paper addresses is straightforward but urgent: the benchmarks used to measure AI progress are no longer difficult enough to provide meaningful information about frontier model capabilities. State-of-the-art LLMs now achieve over 90% accuracy on MMLU (Hendrycks et al., 2021), which was once considered a challenging evaluation of broad academic knowledge. As the paper notes in Section 1:

"The saturation of existing benchmarks, as shown in Figure 1, limits our ability to precisely measure AI capabilities and calls for more challenging evaluations that can meaningfully assess the rapid improvements in LLM capabilities at the frontiers of human knowledge."

This is not merely an inconvenience — it represents a measurement crisis for the field. When a benchmark saturates, it ceases to distinguish between models. If GPT-5 and Claude 4 both score 95% on the same evaluation, we learn nothing about their relative strengths, their failure modes, or whether one genuinely represents a step-change in capability. The benchmark becomes a solved problem, and the research community loses a critical tool for tracking progress.

The saturation problem is compounded by the rapidity of benchmark obsolescence. Figure 1 visually demonstrates this: benchmarks that once seemed impossibly difficult — MMLU, MATH, GSM8K — have been conquered in quick succession as models scaled from GPT-3 to GPT-4 to Claude 3.5 and beyond. A benchmark released today at the frontier of model capabilities may be saturated within 12-18 months. This creates a treadmill effect where evaluators must constantly design new tests just to maintain measurement resolution.

This paper's core motivation is to break this cycle by creating a benchmark so difficult that it remains informative for years, not months — what the authors call "the final closed-ended academic benchmark of its kind." The name "Humanity's Last Exam" is deliberately provocative, signaling that if models achieve high accuracy here, they have effectively reached expert-level performance on structured, verifiable academic questions across a broad range of disciplines.

Why Precise Measurement Matters

The stakes of accurate measurement extend far beyond academic benchmarking. The paper identifies several domains where saturation has concrete negative consequences:

Research and development. When benchmarks saturate, researchers lose the ability to conduct meaningful ablation studies. If a new architecture, training technique, or prompting strategy moves performance from 92% to 93% on a saturated benchmark, the signal-to-noise ratio is too low to draw confident conclusions. This slows scientific progress by making it harder to identify what actually improves models.

Safety and governance. Policymakers and safety researchers rely on benchmark data to assess whether AI systems are approaching dangerous capability thresholds. If the measuring instruments fail — if they show "saturation" while models continue to improve in unmeasured ways — then governance decisions are made on incomplete information. The paper explicitly frames HLE as enabling "more informed discussions about development trajectories, potential risks, and necessary governance measures" (Section 5).

Public understanding. When headlines report "AI achieves 95% on professional exam," the public may reasonably infer that models possess expert-level competence across the board. But saturated benchmarks often test relatively narrow capabilities — multiple-choice factual recall, elementary reasoning — that do not represent the full spectrum of expert performance. A benchmark that genuinely separates current models from human expertise provides clearer signals about what AI can and cannot do.

Prior Approaches and Their Limitations

The paper characterizes the landscape of existing benchmarks along several dimensions, each revealing systematic weaknesses that HLE aims to address:

Standard academic benchmarks (MMLU, GPQA, etc.). MMLU (Hendrycks et al., 2021) was designed to span 57 subjects at varying difficulty levels, from elementary to professional. It served as a remarkably robust benchmark for several years. However, as models improved, MMLU scores climbed past 90% — effectively exhausting its ability to differentiate frontier models. GPQA (Rein et al., 2023) pushed difficulty higher by targeting graduate-level questions in biology, physics, and chemistry, but it covers only three domains and was itself approaching saturation at the time of HLE's development. The fundamental limitation of this class of benchmarks is that they were designed to be challenging but solvable for the models of their era, and as model capabilities advanced, the benchmarks became easy.

Domain-specific challenging benchmarks (FrontierMath, ARC-AGI, SWE-bench). Some recent benchmarks have successfully maintained difficulty by focusing on extremely narrow domains. FrontierMath (Glazer et al., 2024) tests advanced mathematical reasoning with problems drawn from research-level mathematics. SWE-bench (Jimenez et al., 2024) evaluates code generation on real-world GitHub issues. MLE-bench (Chan et al., 2024) tests machine learning engineering capabilities. These benchmarks are genuinely difficult and informative, but their domain specificity limits their utility for assessing general academic capabilities. A model that scores 2% on FrontierMath but 95% on everything else tells you something important, but it doesn't give a holistic picture of the model's knowledge breadth.

Human-expert-designed benchmarks (WMDP, various specialized exams). Some benchmarks employ subject-matter experts to write questions testing advanced knowledge. WMDP (Li et al., 2024) tests knowledge related to weapons of mass destruction, crowd-sourcing questions from experts in biosecurity, cybersecurity, and chemistry. This approach ensures high question quality and resistance to memorization, but prior efforts were typically narrow in scope — focused on a specific domain or capability rather than broad academic coverage.

Adversarially filtered benchmarks. Approaches like Dynabench (Kiela et al., 2021) and adversarial NLI (Nie et al., 2020) use human-in-the-loop filtering where questions are iteratively refined to expose model weaknesses. While effective at maintaining difficulty, these approaches have generally been applied to narrow tasks (e.g., natural language inference) and struggle to scale to hundreds of diverse subjects while maintaining rigorous quality standards.

The Gap This Paper Fills

HLE positions itself at the intersection of several design desiderata that had not been simultaneously satisfied by any prior benchmark:

  1. Breadth of coverage. Like MMLU, HLE spans over 100 subjects — mathematics, biology, chemistry, physics, engineering, computer science, humanities, social sciences, and more. This breadth is critical because it prevents models from gaming the benchmark by specializing in one domain.

  2. Extreme difficulty. Unlike MMLU (which saturates) and unlike GPQA (which was approaching saturation), HLE is explicitly designed to be unsolvable by current frontier models. The submission process requires that questions stump GPT-4O, Claude 3.5 Sonnet, Gemini 1.5 Pro, and O1 before acceptance (Section B.1). This is a dramatic departure from prior practice: rather than hoping questions are hard enough, the process enforces difficulty as a prerequisite.

  3. Expert design with multi-stage review. Questions are written by nearly 1,000 subject-matter experts — primarily professors, researchers, and graduate degree holders across 500 institutions and 50 countries. They undergo two rounds of expert peer review (Section 3.2), combining the quality control of academic peer review with the scale of crowd-sourcing.

  4. Multi-modal format. Approximately 14% of questions include images, requiring models to integrate visual and textual information — a capability that many benchmarks test separately but few test in joint reasoning tasks at this difficulty level.

  5. Resistance to memorization and retrieval. Questions are explicitly required to be "non-searchable" — they cannot be quickly answered via internet retrieval (Section 3.1). This targets a specific failure mode of benchmarks: if a model can answer a question by surfacing a memorized training document or performing a web search, the question isn't testing reasoning or understanding.

  6. Closed-ended format with automated grading. Despite their difficulty, all questions are either multiple-choice (24%) or exact-match short answer (76%), enabling automated verification without human judgment or LLM-as-judge for the final evaluation of correctness. This is a practical necessity for a 2,500-question benchmark — open-ended expert evaluation doesn't scale.

The Calibration Dimension

Beyond accuracy, the paper identifies model calibration as a critical but often-overlooked measurement axis. When benchmarks are saturated, calibration is irrelevant — models answer correctly and know they're correct. But on a benchmark where accuracy is very low (2.7% to 13.4%), calibration becomes the difference between a model that says "I don't know" and one that confidently provides wrong answers — the latter being a hallmark of hallucination or confabulation in safety-critical deployments.

The paper's calibration measurement (Section 4.2, Table 1) uses RMS calibration error following Hendrycks et al. (2022). The finding that all models show calibration errors above 70% — with some as high as 89% — reveals that current LLMs not only fail on HLE but fail to recognize their own failure, a behavior with significant implications for trustworthiness in high-stakes applications.

The Paper's Framing as "The Last Exam of Its Kind"

The paper positions HLE not as a permanent benchmark but as the endpoint for a particular category of evaluation: closed-ended, verifiable, broad-coverage academic questions. The authors are explicit about this boundary (Section 5):

"HLE may be the last academic exam we need to give to models, but it is far from the last benchmark for AI."

This framing acknowledges that even perfect performance on HLE would not demonstrate "artificial general intelligence" or autonomous research capability. HLE tests structured academic problems with known answers — not open-ended research, creative synthesis, or real-world decision-making. It is a focused measure of a specific capability: expert-level performance on closed-ended academic questions spanning many disciplines. The authors argue that when models achieve high accuracy on HLE, this particular capability can be considered solved, and evaluation efforts should shift to other, more open-ended assessments.

This is a strategically important distinction: by defining a clear scope and endpoint, the paper avoids the benchmark inflation problem where each new evaluation claims to be a comprehensive test of "general intelligence." Instead, HLE makes a bounded claim: this is the last benchmark you need for this specific type of measurement, and success here tells you something meaningful and specific about model capabilities.

3. Technical Approach

3.1 Reader Orientation

This paper describes the creation and evaluation of a benchmark — it does not propose a novel machine learning system, architecture, or training procedure. The "system" being built is Humanity's Last Exam (HLE) itself: a curated dataset of 2,500 extremely difficult, closed-ended academic questions designed to resist saturation by frontier large language models. The core problem HLE solves is the precise measurement of LLM capabilities at the frontier of human knowledge, using a multi-stage pipeline that combines automated LLM filtering (rejecting questions current models can answer), large-scale expert crowdsourcing (nearly 1,000 subject-matter experts), and iterative human peer review (two rounds with graduate-level reviewers) to produce a benchmark where state-of-the-art models achieve only 2.7% to 13.4% accuracy, thereby providing clear signal for tracking future progress.

3.2 Big-Picture Architecture (Diagram in Words)

The HLE creation system has five major stages, arranged as a pipeline with quality gates between each stage:

  1. Question Submission Platform — A web-based interface where subject-matter experts submit questions. Each submission includes the question text (optionally with images), answer specification (exact-match string or multiple-choice options with correct answer marked), a detailed solution rationale, academic subject classification, and contributor identification (name, institutional affiliation). The platform enforces formatting rules (e.g., "clear English with precise technical terminology, supporting LaTeX notation wherever necessary") and rejects submissions that are open-ended, subjective, or related to weapons of mass destruction. A $500,000 prize pool incentivizes high-quality submissions.

  2. Automated LLM Difficulty Check — Before a submitted question proceeds to human review, it is tested against several frontier LLMs (GPT-4O, Gemini 1.5 Pro, Claude 3.5 Sonnet, O1 for multi-modal questions; additionally O1-Mini and O1-Preview for text-only questions). The submission criteria differ by question type: exact-match questions must stump all tested models, while multiple-choice questions must stump all but one model (to account for random guessing). Questions that LLMs can answer correctly are rejected — this gate ensures that the benchmark starts at a difficulty floor above current model capabilities. Over 70,000 LLM attempts were logged, yielding approximately 13,000 questions that passed the LLM filter and proceeded to human review.

  3. Expert Review Round 1 (Iterative Refinement) — Human reviewers with graduate degrees (Master's, PhD, JD, etc.) in relevant fields score each submission against a standardized rubric (scores: 0-Discard, 1-Major Revisions, 2-Some Revisions, 3-Okay, 4-Great, 5-Top-Notch, or "Unsure"). Each question receives between 1 and 3 reviews. The primary goal is iterative refinement: reviewers provide detailed feedback to help authors improve question quality, clarity, and robustness. Reviewers prioritize questions with no existing reviews and are instructed to skip questions already having more than three reviews. Questions are assessed for graduate-level difficulty, originality (not textbook-derived or web-searchable), answer objectivity, and formatting quality. Reviewers are told: "if you do not have knowledge/context, or if it would take more than 5 minutes to solve, that is okay" — they are not expected to fully verify every solution but rather to assess adherence to guidelines.

  4. Expert Review Round 2 (Final Selection) — A subset of reviewers identified as particularly high-quality from Round 1, along with organizers, re-evaluate questions using a new rubric (scores: 0-Discard, 1-Not Sure, 2-Pending, 3-Easy questions models got wrong, 4-Borderline, 5-Okay to include, 6-Top question in its category). This round evaluates both the question's quality and the feedback from Round 1 reviewers. Organizers then make final approval decisions based on these second-round assessments.

  5. Post-Release Refinement — After the initial dataset release, the authors conducted additional quality-control procedures:

    • Community feedback / bug bounty: inviting public reports of label errors or major question-statement errors, with each report manually verified by organizers and original authors.
    • Audit by university students: recruited students from top US universities to fully solve a sample of questions, with errors routed between organizers, authors, and auditors until consensus was reached.
    • Searchability audit: questions that models with search tools answered correctly (but answered incorrectly without search) were manually audited, and any easily found via web search were removed.
    • Late contributions: new submissions were accepted post-release, manually reviewed, and added to a second held-out private set.

Information flows linearly through this pipeline: expert submission → automated LLM filter → human expert review (Round 1, iterative) → human expert review (Round 2, selection) → final dataset assembly → post-release refinement. At the end, the dataset is split into a public set (2,500 questions, released for benchmarking) and a private held-out set (for assessing overfitting). A separate evaluation stage then administers the final benchmark to frontier models using a standardized system prompt and an automated O3-MINI judge for answer verification.

3.3 Roadmap for the Deep Dive

  • First, the question submission specification — the exact format, constraints, and incentive structure that define what constitutes a valid HLE question, since every subsequent stage depends on this format.
  • Second, the automated LLM difficulty filter — the specific models used, the separate criteria for exact-match versus multiple-choice questions, and why this gate is placed before human review rather than after.
  • Third, the two rounds of human expert review — the reviewer recruitment criteria, the scoring rubrics for each round, the instructions given to reviewers, and the rationale for a two-round iterative process.
  • Fourth, the prize pool and contributor model — how the $500,000 incentive structure shaped participation and the decision to offer co-authorship to all accepted contributors.
  • Fifth, the post-release audit methodology — the community feedback program, student audit process, and searchability filtering that refined the dataset after initial release.
  • Sixth, the evaluation methodology for measuring model performance on the final dataset — the system prompt, the O3-MINI judge, the calibration error measurement, and the rationale for each design choice in the evaluation protocol.

This ordering follows the chronological pipeline (submission → filtering → review → assembly → refinement → evaluation), which makes the dependencies between stages explicit.

3.4 Detailed, Sentence-Based Technical Breakdown

HLE is fundamentally a dataset engineering paper — its primary contribution is not a new algorithm or architecture but rather a systematic methodology for constructing a benchmark that remains difficult for frontier AI systems. The core design philosophy is: (1) recruit genuine subject-matter experts, (2) enforce difficulty by requiring that models already fail before human review begins, (3) use iterative human peer review to ensure quality and closed-endedness, and (4) maintain calibration and searchability as ongoing post-release concerns.


Question Submission Format and Constraints

Every HLE question submission must conform to a strict schema designed to ensure precision, verifiability, and resistance to shortcut solutions. The required components are:

Question text. The question itself, which may be accompanied by an image reference (approximately 14% of questions are multi-modal). The text must be in "clear English with precise technical terminology" and must support LaTeX notation for mathematical or scientific expressions. The authors explicitly require that "if there is some non-standard jargon for the topic/field, it needs to be explained" — this prevents questions from being artificially difficult due to obscure terminology rather than genuine conceptual challenge. Questions must be "precise, unambiguous, solvable, and non-searchable" (Section 3.1), ensuring models cannot rely on memorization or simple retrieval methods.

Answer specification. HLE supports two formats:

  • Exact-match questions (approximately 76% of the dataset, since 24% are multiple-choice per Section 3.1): The answer is a short, precisely defined string — a number, a formula, a term, a short phrase. Answers must be "kept short and easily verifiable" to support automated grading. For numerical answers, "results should be approximated to max 2-3 decimals." The exact-match format forces models to produce the correct answer without any tolerance for approximation beyond the specified decimal precision, which eliminates the partial-credit ambiguity of more open-ended formats.

  • Multiple-choice questions (approximately 24% of the dataset): The model selects from "five or more answer choices" (Section 3.1). The use of five or more options — rather than the standard four — reduces the random-guessing baseline from 25% to at most 20%. When LLMs provide correct answers with faulty reasoning during testing, "authors are encouraged to modify question parameters, such as the number of answer choices, to discourage false positives" (Section 3.1). This adaptive anti-guessing mechanism is a novel practical design choice: rather than simply rejecting questions that models guess correctly, authors can increase the number of distractors to make lucky guesses less likely.

Detailed solution rationale. Every question must be accompanied by "a detailed solution to verify accuracy." This serves multiple purposes: it allows reviewers to check that the claimed answer is correct (even if they cannot solve the question from scratch within the 5-minute review window), it provides transparency about the reasoning expected, and it creates a record for future audits. The rationale requirement also deters contributors from submitting questions where the answer is known but the reasoning is unclear — a pattern that would indicate the question might be ambiguous or rely on hidden assumptions.

Academic subject classification. Contributors self-declare the subject area they feel best suits their question. The paper reports "over a hundred subjects" in the final dataset, with the top fifty most popular including: Economics, Ecology, Artificial Intelligence, Musicology, Philosophy, Neuroscience, Law, Art History, Biochemistry, Astronomy, Classics, Chess, Chemical Engineering, Microbiology, Classical Ballet, Materials Science, Poetry, Quantum Mechanics, Aerospace Engineering, Civil Engineering, Mechanical Engineering, Geography, Robotics, Data Science, Molecular Biology, Statistics, Immunology, Education, Logic, Computational Biology, Psychology, English Literature, Machine Learning, Puzzle, Cultural Studies, Marine Biology, Archaeology, and Biophysics (Section B.4). This breadth is qualitatively comparable to MMLU's 57-subject coverage but extends into specialist domains (Classical Ballet, Chess, Puzzle) that test knowledge beyond standard academic curricula.

Contributor identification. Name and institutional affiliation are required "to maintain accountability and accuracy" (Section 3.1). This is a deliberate quality-control mechanism: attaching real names and institutions to questions raises the stakes for correctness and originality, since publicly identified errors would reflect on both the individual and their institution. The paper reports contributions from "nearly 1000 subject expert contributors affiliated with over 500 institutions across 50 countries — comprised mostly of professors, researchers, and graduate degree holders" (Section 3.1).

Prohibited content. The submission criteria explicitly rule out: open-ended questions ("Give a proof of...", "Explain why..."), subjective interpretations, content related to weapons of mass destruction, and questions about morality or ethics. These prohibitions are both practical (open-ended answers cannot be automatically graded) and safety-motivated (excluding dual-use knowledge).

Originality requirement. All submissions must be "original work or non-trivial syntheses of published information," though "contributions from unpublished research are acceptable" (Section 3.1). This permits questions derived from ongoing research that has not yet appeared in publications — an important allowance since such questions would be invisible to models trained on published literature. The "non-trivial synthesis" criterion prevents contributors from simply copying textbook problems with minor parameter changes.


Automated LLM Difficulty Filtering

The defining innovation of HLE's pipeline — what distinguishes it from prior expert-written benchmarks — is the pre-review LLM difficulty check. Rather than hoping that expert-written questions will be difficult, the pipeline enforces difficulty as a gate before human review begins.

Models used in filtering. The testing process uses up to six frontier LLMs depending on the question modality (Section B.1):

  • For text-and-image questions: GPT-4O, Gemini 1.5 Pro, Claude 3.5 Sonnet, and O1 (4 models).
  • For text-only questions: the four multi-modal models above, plus O1-Mini and O1-Preview (6 models total).

The inclusion of both multi-modal and text-only models reflects the practical constraint that some frontier models (O1-Mini, O1-Preview) cannot process images. For image-containing questions, only multi-modal models are tested, which slightly reduces the filtering stringency but is necessary to avoid rejecting questions that appear "hard" only because the test model literally cannot see the image.

Acceptance criteria by question type. The criteria differ between exact-match and multiple-choice questions to account for the different random-guessing profiles:

  • Exact-match questions must be answered incorrectly by all tested models. There is no tolerance for partial credit — if any model produces the correct exact-match string, the question is rejected. This is a stringent filter: a single model success, even if accidental (e.g., the model happens to output the right number for the wrong reason), triggers rejection.

  • Multiple-choice questions must be answered incorrectly by all but one of the tested models. The allowance for one model to guess correctly is an explicit acknowledgment that with five or more choices, models have a non-trivial probability of guessing correctly by chance (at minimum, 1/5 = 20% per question). Requiring all models to get all multiple-choice questions wrong would bias the dataset against multiple-choice questions — questions where all models guess incorrectly are likely to have some property that makes the correct answer systematically unlikely under the model's distribution, which might correlate with poor question design rather than genuine difficulty.

The "gate before human review" design choice. The LLM check is placed before human expert review, not after. This is a deliberate efficiency decision: filtering out questions that current models can already solve prevents human reviewers from spending time evaluating questions that would not advance the benchmark's difficulty goal. With 70,000+ LLM attempts producing approximately 13,000 questions that passed the filter (roughly 19% acceptance rate), the pre-filtering prevented reviewers from evaluating approximately 57,000 questions that models already handle correctly. This is a practical scaling consideration: human expert review is the expensive bottleneck, and shifting as much filtering as possible to automated LLM evaluation maximizes the value of reviewer time.

Residual model accuracy on the final dataset. The paper acknowledges that despite this filtering, final evaluation on the dataset reveals "non-zero accuracy" (Section 4.2). This occurs for two reasons: (1) non-determinism in model inference — the same model can produce different answers on different runs, so a question that the model got wrong during filtering might be answered correctly during evaluation (especially with the standardized system prompt used in evaluation, which structures output into explicit reasoning followed by a final answer); (2) the multiple-choice floor — even if models perform at random-chance level on a particular question, random chance itself produces some correct answers across a large enough dataset, and these "correct" answers are not necessarily evidence of understanding. The authors explicitly caution: "we stress the true capability floor of frontier models on the dataset will remain an open question and small inflections close to zero accuracy are not strongly indicative of progress" (Section 4.2).


Human Expert Review: Round 1 (Iterative Refinement)

The first round of human review is designed as peer review with iterative feedback — it is not simply a pass/fail filter but a collaborative process where reviewers help authors improve their questions.

Reviewer qualifications. Reviewers are required to possess "a graduate degree (e.g., Master's, PhD, JD, etc.) in their fields" (Section 3.2). They self-select submissions within their domain expertise, ensuring that questions are evaluated by someone with relevant background knowledge. This domain-matching is critical: a question about category theory in mathematics or Palmyrene script in classics cannot be meaningfully assessed by a generic reviewer — the subtle indicators of question quality (precision of terminology, correctness of the implied theoretical framework, whether the answer is genuinely unambiguous to an expert) require domain knowledge to evaluate.

Review volume and prioritization. Each question receives between 1 and 3 reviews. Reviewers are instructed to "please prioritize questions with no reviews and skip all questions with more than 3 reviews" (Section C.7.1). This creates a natural load-balancing: questions that haven't been reviewed yet get attention first, preventing a long tail of unreviewed submissions while avoiding review pile-up on already-well-reviewed questions.

Scoring rubric. The six-point rubric (plus "Unsure") encodes a clear quality hierarchy:

  • 0 — Discard: The question is "out of scope, not original, spam, or otherwise not good enough to be included in the HLE set and should be discarded." Spam filtering at scale: with nearly 1,000 contributors, some submissions will inevitably be low-effort or off-topic, and a clear discard category prevents these from wasting reviewer effort in later rounds.

  • 1 — Major Revisions Needed: "Major revisions are needed for this question or the question is too easy and simple." This distinguishes fundamentally flawed questions (0) from questions with potential that need substantial reworking (1).

  • 2 — Some Revisions Needed: "Difficulty and expertise required to answer the question is borderline. Some revisions are needed for this question." The key word is "borderline" — the question is close to acceptable but falls short on difficulty or clarity.

  • 3 — Okay: "The question is sufficiently challenging but the knowledge required is not graduate-level nor complex. Minor revisions may be needed for this question." This acknowledges that questions can be difficult enough to stump LLMs without requiring graduate-level expertise — the "LLM difficulty check" is the primary difficulty filter, and questions that pass it are acceptable even if the required knowledge is undergraduate-level, as long as models cannot answer them.

  • 4 — Great: "The knowledge required is at the graduate level or the question is sufficiently challenging." This is the target quality level for HLE.

  • 5 — Top-Notch: "Question is top-notch and perfect." The highest tier, reserved for questions that are not just difficult but elegant, interesting, and well-constructed.

  • Unsure: "Reviewer is unsure if the question fits the HLE guidelines, or unsure if the answer is right." This escape hatch acknowledges the reality that even domain experts may encounter questions outside their specific sub-specialty. Rather than forcing a guess, reviewers can flag uncertainty for organizer attention.

Reviewer instructions and guidance. The instructions provided to reviewers (Section C.7.1) encode several important design principles:

  • Graduate-level difficulty is aspirational, not mandatory: "Questions should usually (but do not always need to) be at a graduate / PhD level or above. (Score 0 if the question is not complex enough and AI models can answer it correctly.) If the model is not able to answer correctly and the question is below a graduate level, the question can be acceptable." This is a pragmatic compromise: the LLM difficulty check is the true difficulty gate, and expert judgment about difficulty level serves as a secondary signal mainly for rejecting questions that LLMs got wrong for spurious reasons (e.g., formatting issues) rather than genuine complexity.

  • Broad disciplinary scope: "Questions can be any field across STEM, law, history, psychology, philosophy, trivia, etc. as long as they are tough and interesting questions." The explicit inclusion of trivia alongside academic subjects is notable — it signals that HLE values any knowledge that resists LLM retrieval, not just traditional academic knowledge.

  • Answer objectivity is non-negotiable: "Questions like 'Give a proof of...', 'Explain why...', 'Provide a theory that explains...' are usually bad because they are not closed-ended and we cannot evaluate them properly. (Score 0)." This hard constraint on closed-endedness is what makes automated grading possible at scale.

  • Web searchability is grounds for rejection: "Score 0 if searchable on web." This is the anti-memorization constraint: a question that can be answered by typing it into Google does not test reasoning or deep knowledge.

  • Formatting quality affects scoring: "Questions should be formatted properly. (Score 1-3 depending on degree of revisions needed). Fix LaTeX formatting if possible. Models often get questions right after LaTeX formatting is added or improved." This instruction reveals a practical insight from the filtering process: poor formatting (e.g., plain-text equations that are ambiguous) can cause models to fail for spurious reasons, and improving formatting can turn an artificially "hard" question into a genuinely evaluable one. Reviewers are expected to help improve formatting, not just judge it.

  • Time-bounded review: "The average time estimated to review a question [is] 3-5 minutes." And critically: "Please check if the answer makes sense as a possible response to the question, but if you do not have knowledge/context, or if it would take more than 5 minutes to solve, that is okay." This explicit time cap acknowledges that the review process is a quality assurance step, not a re-solving step. Reviewers assess whether the question looks well-formed, is original, has an objective answer, and is formatted correctly — they are not expected to independently verify the solution for most questions. This is a necessary scalability concession: fully verifying 13,000 graduate-level questions across 100+ subjects would require an impossibly large and impossibly specialized reviewer pool.

The iterative feedback model. The first round's primary goal is "iteratively refining submissions" — reviewers provide detailed justifications and feedback that question authors can use to improve their submissions. Authors can then resubmit revised versions. This mirrors academic peer review, where the goal is to improve the work through expert critique rather than simply to accept or reject. The paper emphasizes: "This is similar to the peer review process in academic research, where reviewers give feedback to help question submitters create better questions."


Human Expert Review: Round 2 (Final Selection)

The second round shifts from iterative improvement to final curation.

Reviewer selection. Second-round reviewers are "identified by organizers from round 1 reviews as particularly high quality and thorough in their feedback" (Section C.7.2). This creates a meritocratic quality hierarchy: the best reviewers from Round 1 are promoted to a more consequential decision-making role in Round 2. Organizers also participate directly in this round.

Revised review task. Unlike Round 1, which focused on providing feedback to authors, Round 2 reviewers are asked to "grade both the question and look at feedback from round 1 reviewers." They are making inclusion decisions, not just improvement suggestions.

Second-round rubric. The seven-point rubric (0-6) encodes a different set of priorities:

  • 0 — Discard: Same as Round 1 — clearly unsuitable.

  • 1 — Not Sure: "Major revisions are needed for this question or you're just unsure about the question." This is flagged for organizer evaluation rather than automatically discarded.

  • 2 — Pending: "You believe there are still minor revisions that are needed on this question." Again flagged for organizers — the question is close but not finalized.

  • 3 — Easy questions models got wrong: "These are very basic questions that models got correct or the question was easily found online. Any questions which are artificially difficult (large calculations needing a calculator, requires running/rendering code, etc.) should also belong in this category. The models we evaluate cannot access these tools, hence it creates an artificial difficulty bar." This is a crucial category that the Round 1 rubric did not explicitly capture: questions that models fail for artificial reasons (lack of calculator access, inability to execute code) rather than genuine knowledge gaps. The authors explicitly want to exclude these, since adding tool access to models would trivially solve them, making the benchmark's difficulty a measurement artifact rather than a reflection of knowledge limitations. The instruction also flags questions that were "easily found online" — the same model that failed during filtering (without search) might succeed with search, which undermines the benchmark's purpose.

  • 4 — Borderline: "The question is not interesting OR The question is sufficiently challenging, but 1 or more of the models got the answer correct." This catches questions that slipped through the LLM filter (due to non-determinism or because models improved between filtering and evaluation) as well as questions that are difficult but uninteresting.

  • 5 — Okay to include in HLE benchmark: "Very good questions (usually has score of 3 in the previous review round)." This is the acceptance threshold.

  • 6 — Top question in its category: "Great question (usually has a score of 4-5 in the previous review round), at a graduate or research level." The note that "graduate level is less strict for Non-STEM questions" acknowledges that difficulty manifests differently across disciplines — a challenging question in Classics or Musicology may not require the same kind of formal reasoning as a challenging question in Mathematics, but it can still probe expert-level knowledge.

Organizer approval gate. After Round 2 reviewers score questions, "organizers then approve questions based on reviewer feedback in this round." This final human judgment step — where the core organizing team makes inclusion decisions informed by expert reviewer assessments — adds a quality control layer that catches edge cases where reviewer scores might not perfectly capture the organizing team's vision for the benchmark.


Prize Pool and Contributor Incentives

The $500,000 prize pool is not merely a marketing detail — it is a carefully designed incentive mechanism that shapes the contributor distribution.

Prize structure. The pool awards "$5,000 USD for each of the top 50 questions and $500 USD for each of the next 500 questions, as determined by organizers" (Section 3.1). The steep drop-off ($5,000 for top tier, $500 for second tier) creates a tournament incentive: contributors are motivated to produce genuinely outstanding questions, not just minimally acceptable ones, because the top 50 questions receive 10× the reward of the next 500. The total prize allocation is $5,000 × 50 + $500 × 500 = $250,000 + $250,000 = $500,000, consistent with the stated pool size.

The "as determined by organizers" clause is significant: prizes are not awarded purely by reviewer scores but by organizer judgment, allowing the central team to reward questions that align with their vision for the benchmark (e.g., prioritizing questions in underrepresented domains, rewarding particularly elegant or insightful constructions, or balancing subject coverage).

Co-authorship incentive. Beyond monetary prizes, the paper offers "co-authorship for anyone with an accepted question in HLE" (Section 3.1). Authorship position "is ranked based on the number of accepted questions in HUMANITY'S LAST EXAM" (Section A). This creates a secondary incentive: contributors who submit multiple accepted questions receive higher authorship placement, encouraging sustained engagement rather than single-shot submissions. For academics, co-authorship on a high-profile benchmark paper provides career-relevant publication credit, which may be a stronger motivator than the monetary prize for many contributors.

Contributor demographics. The paper reports contributors are "affiliated with over 500 institutions across 50 countries — comprised mostly of professors, researchers, and graduate degree holders" (Section 3.1). The geographic and institutional diversity (500 institutions in 50 countries) is a deliberate design feature: it reduces the risk that the benchmark reflects the knowledge biases of any single academic tradition, institution, or cultural context. A benchmark created entirely by, say, US computer science PhD students would disproportionately test knowledge from that specific educational pipeline. Global contributions help ensure that "expert-level academic knowledge" means knowledge as defined by experts worldwide, not just experts in one region or tradition.


Post-Release Refinement Methodology

The paper describes three post-release quality-control mechanisms that refined the dataset after its initial public release:

Community feedback / bug bounty program. Recognizing that "due to the advanced, specialized nature of many submissions, reviewers were not expected to verify the full accuracy of each provided solution rationale if it would take more than five minutes" (Section B.2), the authors opened a public feedback channel specifically for identifying "label error and major errors in the statement of the question." The narrow scope of this program is important: it targets clear errors (the stated answer is wrong, the question text contains a factual mistake) rather than subjective quality judgments. Each error report was "manually verified by the organizers with feedback from the original author of the question when appropriate." This creates a post-publication correction pipeline similar to academic errata.

University student audit. The authors "recruited students from top universities in the United States to fully solve a sample of questions from HLE." Unlike the 5-minute reviewer time budget, these students were expected to fully solve the questions — providing independent verification of answer correctness and question solvability. Disagreements were "routed between organizers, original question authors, and auditors until consensus was reached." The data from these audits was used "to further refine our dataset," likely by removing or correcting questions where multiple independent solvers reached a different answer than the one claimed by the original author.

Expert disagreement rate estimation. The paper reports a two-round audit process on samples of 200 questions each (Section B.3):

  • Round 1 aimed to "identify common categories of imprecise questions, such as open-ended formats, reliance on rounded numerical values, or submissions from authors with low acceptance rates." Based on these signals, questions with similar issues were manually removed or revised.

  • Round 2 was conducted on a fresh sample of 200 questions, yielding a "final estimated expert disagreement rate of 15.4% for the public set."

For a targeted biology, chemistry, and health subset — motivated by external criticism that these domains might have higher error rates — the authors found "an expert disagreement rate of approximately 18%." They contextualize this as "in line with what is observed in other challenging, expert-grade machine learning benchmarks" and note that similarly designed work (specifically citing HealthBench) observes "disagreement among expert physicians is frequent on complex health topics."

The paper identifies three factors contributing to audit complexity:

  1. The need for multiple experts: "In several cases, a reviewer identified a critical piece of information, such as a decades-old paper or a foundational concept not immediately apparent to others, that was essential to confirming an answer's validity." If a single-reviewer methodology had been used (where one dissenting expert flags a question), "the disagreement rate on the aforementioned health-focused subset jumps from 18% to 25%." This is a substantive methodological point: the disagreement rate depends on how many reviewers examine each question, and the "true" error rate is not directly observable.

  2. Questions from research experience: HLE intentionally includes "questions based on insights from the direct, hands-on experiments of its contributors," capturing "knowledge gained from direct research experiences, which is often difficult to verify through standard literature searches or by external reviewers." This design choice trades verifiability for difficulty: questions that can be verified by any expert through a literature search are also questions that models might answer through memorization or retrieval. Questions drawing on unpublished experimental insights are harder to verify but also harder for models to answer through training data alone.

  3. Understanding question design: For multiple-choice questions, "researchers sometimes leverage the multiple-choice format with the objective of identifying the most plausible answer among the provided options." This means that the "correct" answer is not always the uniquely correct answer in an absolute sense, but rather the best answer among the given choices. Reviewers needed to understand this design principle to evaluate questions appropriately — "it guided them to evaluate the relative merits of the given choices rather than treating the task as an open-ended search for a perfect solution."

Searchability audit. The paper conducted a targeted audit to identify questions that could be answered via web search (Section B.2):

  • A question was flagged as "potentially searchable if a model with search tools answered correctly, but answered incorrectly without search."
  • Each flagged question was "then manually audited, removing any that were easily found via web search."
  • The models used for this procedure were "GPT-4o mini/GPT-4o search and Perplexity Sonar models."
  • The paper notes that "current frontier model performance on HLE after applying this procedure is similar to their performance on HLE before applying this procedure," suggesting that the searchable questions removed were a small fraction of the dataset and that their removal did not substantively change benchmark difficulty.

HLE-Rolling. As a forward-looking mechanism, the paper introduces "a dynamic fork of the dataset post-release: HLE-ROLLING" that will be "regularly updated to address community feedback and integrate new questions" (Section B.3). The stated goal is to "provide a seamless migration path for researchers once frontier models begin to hit the ceiling performance on the original HLE dataset." This acknowledges the central tension of benchmark design: HLE is designed to be "the last exam of its kind," but if models eventually saturate it, the benchmark loses its measurement value. HLE-Rolling provides an evolutionary path — rather than starting from scratch with a new benchmark, the dataset can be refreshed with even harder questions while maintaining continuity of format and evaluation methodology.


Evaluation Methodology

The final component of the technical approach is the evaluation protocol used to measure model performance on the completed dataset. This is separate from the creation pipeline but is essential to the benchmark's function.

System prompt for model responses. The evaluation uses a standardized system prompt that structures model outputs into three explicitly labeled components (Section C.1.1):

"Your response should be in the following format: Explanation: {your explanation for your answer choice} Answer: {your chosen answer} Confidence: {your confidence score between 0% and 100% for your answer}"

The structured format serves multiple purposes: (1) it separates reasoning from the final answer, making it easier for the automated judge to extract the answer; (2) it requires models to state their confidence explicitly, enabling calibration measurement; (3) it encourages chain-of-thought reasoning by mandating an explanation section before the answer. For models that "do not support a system prompt," the prompt is "added as a separate user prompt" (Section C.1.1).

Automated answer judging with O3-MINI. The correctness of model responses is evaluated using O3-MINI as an automated judge — a design choice that itself reflects the benchmark's difficulty. If a weaker model were used as judge, it might fail to correctly assess whether the model's answer matches the ground truth for questions at the frontier of human knowledge. The paper uses "o3-mini-2025-01-31 with structured decoding enabled" (Section C.1.1), which is the most capable model in their evaluation suite.

The judge prompt specifies a structured output format with five fields:

"extracted_final_answer: The final exact answer extracted from the [response]. Put the extracted answer as 'None' if there is no exact, final answer to extract from the response."

"reasoning: Explain why the extracted_final_answer is correct or incorrect based on [correct_answer], focusing only on if there are meaningful differences between [correct_answer] and the extracted_final_answer. Do not comment on any background to the problem, do not attempt to solve the problem, do not argue for any answer different than [correct_answer], focus only on whether the answers match."

"correct: Answer 'yes' if extracted_final_answer matches the [correct_answer] given above, or is within a small margin of error for numerical problems. Answer 'no' otherwise, i.e. if there if there is any inconsistency, ambiguity, non-equivalency, or if the extracted answer is incorrect."

"confidence: The extracted confidence score between 0% and 100% from [response]. Put 100 if there is no confidence score available."

The judge's task is deliberately constrained: it is told to "focus only on whether the answers match" and explicitly instructed "do not attempt to solve the problem" and "do not argue for any answer different than [correct_answer]." This prevents the judge from second-guessing the ground truth — a critical safeguard since the judge itself is an imperfect model. The judge is verifying whether the model's answer matches the provided answer, not whether the provided answer is correct — that verification was the responsibility of the expert review pipeline.

The margin-of-error allowance for numerical problems ("within a small margin of error for numerical problems") acknowledges that floating-point representations and rounding conventions can produce technically correct but format-different answers. The example provided in Section C.1.1 illustrates how the judge handles equivalent mathematical expressions: when the correct answer is cos(π/n) / (2(1+cos(π/n))) and the model outputs cot(π/n) / (2 cot(π/(2n))), the judge recognizes them as equivalent through algebraic manipulation and marks the answer correct. This equivalence checking is essential for mathematical questions where multiple valid representations exist.

Model versions and sampling parameters. The evaluation uses specific, versioned models rather than generic model names (Table 4, Section C.5):

ModelVersion
GPT-4Ogpt-4o-2024-11-20
Grok 2grok-2-latest
Claude 3.5 Sonnetclaude-3-5-sonnet-20241022
Gemini 1.5 Progemini-1.5-pro-002
Gemini 2.0 Flash Thinkinggemini-2.0-flash-thinking-exp-01-21
O1o1-2024-12-17
DeepSeek-R1January 20, 2025 release
O3-Mini (High)o3-mini-2025-01-31

All models use "temperature 0.0 when configurable," except for "o3-mini and o1 models [which] only support temperature 1.0." The use of temperature 0.0 for most models is a standard evaluation practice that maximizes reproducibility by making model outputs deterministic. For O1 and O3-Mini, the inability to set temperature to 0.0 introduces some non-determinism into their evaluations, which contributes to the "inherent noise in model inference" that the paper cites as one reason models show non-zero accuracy despite the LLM filtering process (Section 4.2).

The Gemini 2.0 Flash Thinking model has a specific note: "The first version of the paper along with Figure 5 used the now deprecated 12-19 model with temperature 0.0. The new model is sampled at temperature 0.7" (Table 4 footnote). This change between paper versions introduces variability in the Gemini results and highlights the challenge of benchmarking against rapidly iterating model APIs.

Calibration error measurement. The RMS calibration error metric is implemented following Hendrycks et al. (2022). The underlying computation (not explicitly shown in the paper but standard in calibration measurement) bins predictions by their stated confidence, computes the difference between the average confidence in each bin and the actual accuracy in that bin, and then computes:

RMS Calibration Error=1Ni=1N(pip^i)2×niNtotal\text{RMS Calibration Error} = \sqrt{\frac{1}{N} \sum_{i=1}^{N} (p_i - \hat{p}_i)^2 \times \frac{n_i}{N_{\text{total}}}}

where the sum is taken over confidence bins, $p_i$ is the average confidence in bin $i$, $\hat{p}_i$ is the actual accuracy in bin $i$, $n_i$ is the number of predictions in bin $i$, and $N_{\text{total}}$ is the total number of predictions.

What it computes: a weighted root-mean-square of the gap between what a model claims its confidence is and what its actual accuracy is, where bins with more predictions contribute more to the error. A perfectly calibrated model would have $p_i = \hat{p}_i$ for all bins, yielding zero error. The worst possible calibration (always 100% confident but always wrong) would produce very high error.

Why this form: root-mean-square weighting penalizes large calibration errors more heavily than small ones — being off by 50 percentage points in one bin is worse than being off by 5 points in ten bins. The bin-size weighting prevents small bins (where accuracy estimates are noisy) from dominating the metric. This is the standard metric in the calibration literature (Hendrycks et al., 2022), enabling direct comparison with prior work.

For models where no confidence score is extractable from the response, "confidence" is set to 100% — the paper notes: "Put 100 if there is no confidence score available." This is a conservative choice: a model that fails to state confidence at all is treated as maximally confident, which penalizes its calibration error. This incentivizes models to produce explicit confidence estimates.

Token count analysis. To characterize the computational cost of different models' approaches to HLE questions, the paper analyzes "average completion token counts" across models (Section 4.2, Figure 5). Reasoning models (Gemini 2.0 Flash Thinking, O1, DeepSeek-R1) are shown to require "significantly more tokens compared to non-reasoning models" — token counts reaching thousands per question, compared to hundreds for models like GPT-4O or Grok 2 (Figure 6 in Section C.4). This analysis is not just descriptive but normative: the paper emphasizes that "future models should not only do better in terms of accuracy, but also strive to be compute-optimal" (Section 4.2), connecting HLE evaluation to the broader question of whether accuracy improvements are achieved through genuine capability advances or simply through increased inference-time computation.


Summary of Key Design Choices and Their Justifications

  • LLM difficulty filter before human review: prevents wasting expert reviewer time on questions that models already solve, and enforces a difficulty floor that is defined empirically rather than by human judgment of what "should be" difficult.

  • Different acceptance criteria for multiple-choice vs. exact-match: accounts for different random-guessing profiles; exact-match has no guessing tolerance (all models must fail), while multiple-choice allows one success to avoid penalizing the format.

  • Two-round human review with different goals: Round 1 focuses on iterative improvement (feedback to authors), while Round 2 focuses on final curation (inclusion decisions by reviewers and organizers) — this separation prevents the tension between helping authors improve and making hard acceptance decisions.

  • 5-minute review time budget with explicit permission to not fully solve questions: scales human review to 13,000 questions across 100+ subjects by treating reviewers as quality-assurance assessors rather than independent answer-verifiers, while acknowledging this introduces an estimated 15.4% expert disagreement rate.

  • Monetary prizes with tournament structure ($5,000 top-50, $500 next-500): creates incentive for genuinely outstanding questions rather than minimally acceptable ones, with organizer judgment determining "top" questions to align rewards with benchmark vision.

  • Co-authorship ranked by number of accepted questions: encourages sustained contribution and provides academic career incentive that complements monetary rewards, with contributor demographics spanning 500 institutions across 50 countries ensuring global knowledge representation.

  • O3-MINI as automated judge with constrained task: uses the most capable evaluated model for answer matching verification, but restricts its role to comparing answers against provided ground truth rather than re-solving problems — preventing judge errors from corrupting benchmark scores.

  • Structured output format with mandatory explanation and confidence: enables automated answer extraction, encourages chain-of-thought, and provides calibration data — all without requiring human evaluation of model outputs.

  • Post-release refinement through community feedback, student audits, and searchability filtering: acknowledges that the expert review pipeline cannot guarantee zero errors and provides correction mechanisms that improve dataset quality over time.

  • HLE-Rolling as a dynamic fork: future-proofs the benchmark by providing an update pathway when models begin saturating the original dataset, avoiding the need for entirely new benchmark construction.

4. Key Insights and Innovations

Innovation 1: Difficulty-as-Prerequisite: Inverting the Benchmark Design Paradigm

The foundational conceptual move in HLE is redefining "benchmark difficulty" from an aspirational property (we hope models can't solve these questions) to a gating criterion (questions are rejected if models can solve them). Prior expert-designed benchmarks — GPQA (Rein et al., 2023), WMDP (Li et al., 2024), FrontierMath (Glazer et al., 2024) — relied on expert judgment to estimate difficulty: domain specialists wrote questions they believed would be challenging, and the benchmark's difficulty was discovered post-hoc when models were evaluated. This approach has systematically underestimated model capabilities: questions that experts considered graduate-level turned out to be answerable by frontier models, leading to the rapid saturation documented in Figure 1.

HLE inverts this. Difficulty is not predicted; it is verified empirically before human review begins. The automated LLM filter in Section 3.2 (70,000+ attempts, with exact-match questions requiring all tested models to fail and multiple-choice questions allowing at most one success) functions as a measurement instrument rather than a design heuristic. A question is not "expected to be hard" — it is measurably beyond current model capabilities, and only then does it proceed to expert quality review.

Why this is a fundamental shift, not incremental refinement. The pre-review LLM gate changes the epistemic status of benchmark difficulty. In the traditional paradigm, when a new model scores poorly on a benchmark, we cannot distinguish between two explanations: (1) the model genuinely lacks the tested capability, or (2) the benchmark is poorly constructed, ambiguous, or tests knowledge absent from the model's training distribution for incidental reasons. HLE's design eliminates explanation (2) for the direction of measurement: if a model fails on HLE, we know that the questions were empirically verified as unsolvable by the best models of the benchmark's construction era, and that human experts subsequently vetted them for clarity and objectivity. The failure is more likely to reflect genuine capability limitations rather than benchmark artifacts.

This is philosophically similar to the shift from construct validity (does this test measure what we think it measures?) to criterion validity (does this test correlate with an external gold standard?) in psychometrics, except here the criterion is not correlation with an external measure but rather the observable fact that current systems — which we know to be highly capable in many domains — cannot answer these questions. The benchmark's difficulty is anchored to an empirical fact about the world rather than a judgment about what "should be" hard.

Evidence anchoring. Figure 1 and Table 1 demonstrate the consequence of this design: despite the LLM filter removing questions that models could answer, final evaluation still shows non-zero accuracy (2.7% to 13.4%), which the authors attribute to inference non-determinism and the multiple-choice guessing floor — not to any model genuinely "solving" questions. The 19% acceptance rate at the LLM filter stage (13,000 of 70,000+ attempts passed) quantifies the stringency of this gate.

Distinction from adversarial filtering. Prior work like Dynabench (Kiela et al., 2021) and adversarial NLI (Nie et al., 2020) also used model failure as a signal, but with a crucial difference: in adversarial filtering, humans iteratively modify questions to cause model failure, which can produce questions that are artificially difficult — exploiting brittle model behaviors rather than testing robust knowledge. HLE's approach is the reverse: experts propose questions based on their domain knowledge, and the LLM filter merely verifies that these naturally-occurring expert questions are beyond current capabilities. The distinction matters because adversarially constructed benchmarks tend to overfit to specific model weaknesses and become less informative as models change, whereas questions derived from genuine expert knowledge should remain meaningful even as model architectures evolve.


Innovation 2: Calibration as a First-Class Metric on Deliberately Difficult Benchmarks

The paper's emphasis on calibration error — and its finding that all models show RMS calibration errors above 70% (Table 1), with GPT-4O reaching 89% — is not merely an additional measurement. It introduces a diagnostic framework for interpreting low accuracy on difficult benchmarks: the difference between a model that knows it doesn't know and one that confidently hallucinates.

Prior work on calibration (e.g., Hendrycks et al., 2022; Wei et al., 2024) typically measured calibration on benchmarks where models achieved moderate-to-high accuracy — the question was whether models' confidence tracked their accuracy when they usually got answers right. HLE inverts the calibration question: on a benchmark where models almost always get answers wrong, calibration measures whether models recognize their own failure. A model with 5% accuracy and 5% average confidence would be well-calibrated (it knows it's guessing). A model with 5% accuracy and 80% average confidence — which is closer to what HLE reveals — is systematically misleading, confidently asserting incorrect answers to questions far beyond its capabilities.

Why this changes how we interpret benchmark results. In the traditional accuracy-only framework, a 5% score on a difficult benchmark is simply "poor performance" — it tells us the model fails but not how it fails. The calibration dimension reveals a qualitative distinction between two failure modes:

  • Calibrated failure (low accuracy, low confidence): The model recognizes the limits of its knowledge. This is the safer failure mode for deployment — the model can be trusted to flag uncertainty, enabling human oversight or task routing.

  • Uncalibrated failure (low accuracy, high confidence): The model confidently provides wrong answers. This is dangerous in any application where users might trust model outputs without independent verification — medical diagnosis, legal analysis, scientific research assistance.

The paper's finding that all frontier models exhibit the second failure mode (RMS calibration errors 70-89%) is a substantive discovery about current LLM behavior: these models do not possess reliable self-knowledge about the boundaries of their competence. When faced with questions at the frontier of human expertise, they do not retreat to "I don't know" but instead generate plausible-sounding but incorrect answers with high stated confidence.

Evidence anchoring. Table 1 shows the calibration error alongside accuracy for all models. The worst-calibrated model (GPT-4O, 89% RMS error with 2.7% accuracy) and the best-calibrated (DeepSeek-R1, 73% RMS error with 8.5% accuracy) both show errors far above what would be acceptable in any deployment context. Figure 5 and Figure 6 show that reasoning models achieve slightly better accuracy at substantially higher token cost, but their calibration remains poor — additional computation improves answer quality somewhat but does not fundamentally fix the self-assessment failure.

Connection to the hallucination literature. This finding reframes hallucination not as an occasional error on routine tasks but as a systematic property that emerges when models are pushed beyond their training distribution. The HLE results suggest that frontier models have not learned a generalizable "uncertainty estimation" capability — they can be calibrated on in-distribution tasks (where training data includes many examples of the relevant knowledge) but lose calibration entirely when facing questions that genuinely require expert-level reasoning they haven't acquired. This implies that post-hoc calibration techniques (temperature scaling, conformal prediction) trained on standard benchmarks may not transfer to frontier-difficulty settings, which is a significant challenge for safe deployment.


Innovation 3: Expert Disagreement as a Measurable, Interpretable Property of the Benchmark

Rather than treating expert disagreement as a flaw to be eliminated or hidden, the paper explicitly measures and reports it — 15.4% for the full public set, approximately 18% for a biology/chemistry/health subset (Section B.3) — and uses the measurement and its surrounding methodology to make a substantive argument about the nature of frontier knowledge and the limits of benchmark construction.

Prior treatment of expert disagreement. Most benchmarks report accuracy metrics (model score vs. ground truth) without quantifying how much uncertainty exists in the ground truth itself. When expert disagreement is mentioned, it's typically in the context of "inter-annotator agreement" during dataset construction, reported as a quality metric (e.g., Cohen's kappa for labeling tasks) and treated as a problem to be minimized. MMLU, GPQA, and most academic benchmarks implicitly assume that questions have unambiguous answers that any qualified expert would agree on. If disagreement exists, it indicates poorly designed questions.

HLE's approach is more nuanced. By measuring and reporting disagreement rates, and by explaining the sources of disagreement (the need for multiple experts, questions from unpublished research experience, the multiple-choice "best answer among options" design), the paper makes visible a phenomenon that prior benchmarks obscured: at the frontier of human knowledge, even genuine experts sometimes disagree. This is not a bug in HLE's construction — it's a feature of the knowledge territory that HLE maps.

The three-source taxonomy of expert disagreement. The paper's analysis of why experts disagree on HLE questions is itself a conceptual contribution (Section B.3):

  1. The multi-expert requirement: Some questions require knowledge so specialized that only a subset of experts in the relevant field possess it. A single reviewer might lack the specific knowledge to verify the answer, while another reviewer with deeper sub-specialization can confirm it. This means that "disagreement" in a single-reviewer setting often reflects reviewer knowledge gaps rather than genuine answer ambiguity. The paper quantifies this: moving from a multi-reviewer to single-reviewer methodology raises the health-subset disagreement rate from 18% to 25%, suggesting that approximately one-third of apparent "disagreement" is actually asymmetric expertise.

  2. Questions from direct research experience: HLE deliberately includes knowledge derived from hands-on experimental work that has not been published. Such knowledge is inherently unverifiable through literature search, meaning that even a highly qualified expert reviewer cannot independently confirm the answer — they must either trust the contributor's reported experience or flag the question as unverifiable. This is a deliberate design tradeoff: such questions are harder for models to answer through memorization but also harder for humans to verify.

  3. Multiple-choice design as "best answer" selection: For some frontier questions, there is no uniquely correct answer in the absolute sense — but among the provided options, one is clearly the best under expert judgment. Reviewers who approach such questions looking for the "uniquely correct answer" may disagree with the provided answer choice, while reviewers who understand the "most plausible among options" framing will agree.

Why this matters for benchmarking methodology. This analysis challenges the implicit assumption that benchmark ground truth can be treated as absolute. If 15-18% of questions have answers that qualified experts would dispute or cannot independently verify, then a model's "accuracy" on HLE is not measuring the fraction of questions where the model knows the correct answer in an objective sense — it's measuring the fraction where the model's output matches what the question's author and the reviewers who accepted it consider correct. For most practical purposes this distinction is academic (the benchmark's measurement validity depends on ground truth being sufficiently reliable, not perfectly objective), but it becomes significant when model performance approaches the expert disagreement rate. If a future model achieves, say, 85% accuracy on HLE, we cannot determine whether the remaining 15% represents model errors or benchmark noise — the two are confounded at that performance level.

This is a fundamental limit on HLE's useful measurement range, and the paper's transparency about it is methodologically commendable. Most benchmarks hide this limit; HLE quantifies it and makes it available for interpretation.

Evidence anchoring. The disagreement rate methodology and findings are detailed in Section B.3, with the two-round audit process on samples of 200 questions providing the empirical basis for the 15.4% and 18% estimates. The comparison to HealthBench (cited but not detailed) contextualizes these rates as typical for expert-grade evaluations.


Innovation 4: HLE as a Measurement Instrument with a Defined Endpoint

The paper's framing of HLE as potentially "the last academic exam we need to give to models" (Section 5) is not merely rhetorical. It represents a conceptual innovation in benchmark philosophy: rather than positioning HLE as yet another evaluation in an endless sequence (MMLU → GPQA → HLE → ???), the authors define a clear scope boundary for what the benchmark measures and what success on it would signify.

Prior benchmark framing. Most benchmark papers position their contribution as "a new, more challenging evaluation that reveals gaps in current model capabilities." The implicit promise is that the benchmark will remain useful until models saturate it, at which point the community will need an even harder benchmark. This creates an infinite regress: each benchmark is eventually solved, each solution motivates a harder benchmark, and the cycle continues without a natural stopping point.

HLE breaks this pattern by defining a telos — an endpoint. The benchmark tests "closed-ended academic questions with broad subject coverage" at the frontier of human knowledge. When models achieve high accuracy on HLE, the authors argue, this specific capability can be considered solved. Further evaluation should then shift to different kinds of measurement: open-ended research tasks, creative problem-solving, real-world decision-making, interactive assistance — capabilities that HLE explicitly does not test.

The explicit scope boundary. The paper is unusually precise about what HLE does not measure (Section 5):

"HLE tests structured academic problems rather than open-ended research or creative problem-solving abilities, making it a focused measure of technical knowledge and reasoning."

And:

"High accuracy on HLE would demonstrate expert-level performance on closed-ended, verifiable questions and cutting-edge scientific knowledge, but it would not alone suggest autonomous research capabilities or 'artificial general intelligence.'"

These scope limitations are not weaknesses — they are design choices that give HLE's measurements clear meaning. A high score means something specific and bounded; a low score (as currently observed) means something equally specific. This contrasts with benchmarks that claim or imply comprehensive measurement of "intelligence" or "understanding" — claims that are almost always overbroad relative to what the benchmark actually tests.

The HLE-Rolling mechanism as an escape valve. The paper acknowledges that even this deliberately extreme benchmark may eventually saturate, and provides HLE-Rolling (Section B.3) as a mechanism for refreshing question difficulty while maintaining format and evaluation continuity. This is not a repudiation of the "last exam" framing — it's an acknowledgment that "frontier of human knowledge" is a moving target (as models improve, the frontier of what they cannot do shifts) and that the benchmark can evolve to track that frontier without losing its identity. HLE-Rolling is the dynamic mechanism that keeps the "last exam" philosophy viable in practice: by continuously integrating new, verified-difficult questions, the benchmark can remain at the frontier even as models advance.

Why this matters for the field. The benchmarking community has long struggled with a version of Goodhart's law: when a measure becomes a target, it ceases to be a good measure. By defining HLE as measuring a specific, bounded capability and by explicitly stating what success would mean (and what it would not mean), the paper attempts to inoculate the benchmark against overinterpretation. When (if) models achieve 90% on HLE, the authors have already told us what that signifies and what it doesn't. This is a level of measurement discipline that the field has not consistently practiced, and HLE's explicit endpoint framing is a model for how future benchmarks could communicate their scope.

5. Experimental Analysis

Evaluation Methodology

The experimental design for HLE's evaluation differs fundamentally from typical ML benchmarks because the "system" being evaluated is the benchmark itself — its capacity to measure frontier model capabilities — rather than a new model or algorithm. The evaluation answers two distinct questions: (1) Do current frontier LLMs fail on HLE? (2) Does HLE provide useful signal (discrimination between models, calibration measurement, difficulty structure) that saturated benchmarks cannot? The methodology is designed to answer both.

  • Dataset. HLE consists of 2,500 questions spanning over 100 subjects, publicly released for benchmarking with a separate private held-out set for assessing model overfitting (Section 3, Section B). The questions are split approximately 76% exact-match (short answer with a precise string, number, or formula as the target) and 24% multiple-choice with five or more options (Section 3.1). Approximately 14% of questions include an image component, making HLE a multi-modal benchmark. The dataset was constructed through a multi-stage process: expert submission from nearly 1,000 contributors across 500 institutions and 50 countries, automated LLM difficulty filtering (70,000+ attempts; exact-match questions must stump all tested models, multiple-choice must stump all but one), two rounds of expert human review, and post-release audit and refinement (Section 3.2, Section B).

  • Base model(s). The evaluation tests eight frontier LLMs spanning multiple model families and training paradigms (Section 4.1, Table 1, Table 4). The full multi-modal set includes: GPT-4O (gpt-4o-2024-11-20), Grok 2 (grok-2-latest), Claude 3.5 Sonnet (claude-3-5-sonnet-20241022), Gemini 1.5 Pro (gemini-1.5-pro-002), and Gemini 2.0 Flash Thinking (gemini-2.0-flash-thinking-exp-01-21). Text-only evaluations additionally include: O1 (o1-2024-12-17), DeepSeek-R1 (January 20, 2025 release), and O3-Mini (High) (o3-mini-2025-01-31). The selection rationale is straightforward: these are the most capable publicly available models at the time of HLE's release, and the benchmark's purpose is to test whether even the frontier fails. All models use temperature 0.0 when configurable; O1 and O3-Mini only support temperature 1.0, introducing some non-determinism that the paper acknowledges as contributing to residual non-zero accuracy (Section 4.2). The Gemini 2.0 Flash Thinking model was updated between paper versions (from temperature 0.0 with the deprecated 12-19 model to temperature 0.7 with the 01-21 model), introducing variability in that model's results (Table 4 footnote).

  • Metrics. The primary metric is accuracy — the fraction of questions for which the model's extracted final answer matches the ground truth, as determined by an automated O3-MINI judge (Section 4.1, Section C.1.1). For exact-match questions, matching requires the extracted answer to be equivalent to the ground truth, with a "small margin of error for numerical problems" to account for floating-point and rounding conventions. For multiple-choice, the model must select the correct option. The secondary metric is RMS calibration error, following the implementation from Hendrycks et al. (2022). Models are prompted to provide both an answer and a confidence score (0% to 100%) in a structured format (Section C.1.1). RMS calibration error bins predictions by stated confidence, computes the difference between average confidence and actual accuracy in each bin, and produces a weighted root-mean-square error — where a perfectly calibrated model would have zero error (stated confidence matches observed accuracy) and higher values indicate systematic overconfidence or underconfidence. For responses where no confidence score is extractable, confidence is set to 100% — a conservative choice that penalizes models that fail to express uncertainty. Additionally, the paper reports average completion token counts by model and subject category (Section 4.2, Figure 5, Figure 6) as a descriptive metric characterizing the computational cost of different models' approaches to HLE questions.

  • Baselines. HLE evaluation does not employ traditional ML baselines because the benchmark's purpose is to measure absolute model capability, not relative improvement. However, Figure 1 positions HLE against several well-known saturated benchmarks — MMLU (Hendrycks et al., 2021), MATH (Hendrycks et al., 2021), GSM8K (Cobbe et al., 2021), GPQA (Rein et al., 2023), and ARC-Challenge (Chollet et al., 2024) — showing that models achieving over 90% on these prior benchmarks achieve only 2.7%–13.4% on HLE. This comparison establishes HLE's difficulty relative to the existing benchmark landscape. Within the HLE evaluation itself, all eight models are compared directly against each other in Table 1, providing a cross-model ranking rather than a baseline comparison.

  • Generation budget / compute accounting. There is no explicit generation budget constraint in the HLE evaluation — models are evaluated under standard inference settings (temperature 0.0 where available, a standardized zero-shot chain-of-thought prompt with structured output formatting). The paper does, however, characterize computational cost indirectly through token count analysis (Section 4.2, Figure 5, Figure 6). Reasoning models (Gemini 2.0 Flash Thinking, O1, DeepSeek-R1, O3-Mini) generate substantially more tokens than non-reasoning models (GPT-4O, Grok 2, Claude 3.5 Sonnet, Gemini 1.5 Pro) — average completion tokens reaching thousands per question for reasoning models versus hundreds for non-reasoning models. The paper does not conduct a FLOPs-matched comparison or control for inference compute, but it explicitly notes the normative implication: "future models should not only do better in terms of accuracy, but also strive to be compute-optimal" (Section 4.2). Token counts are broken down by subject category in Figures 5 and 6, revealing that token usage varies by domain (mathematics typically requires more tokens than humanities, though the pattern differs by model).

  • Cross-validation / statistical protocol. HLE evaluation does not use cross-validation in the traditional ML sense — there is no training procedure whose generalization is being tested. Instead, the benchmark employs a public/private split design (Section 3): the public release contains 2,500 questions used for standard benchmarking, while a separate private held-out set (size unspecified) is reserved for assessing whether models overfit or game the public benchmark through repeated evaluation or training contamination. The post-release audit methodology (Section B.2, Section B.3) provides the statistical quality control: two rounds of auditing on 200-question samples, with expert disagreement rates of 15.4% for the full public set and approximately 18% for the biology/chemistry/health subset serving as the primary reliability metrics. The evaluation uses fixed model versions (Table 4) and a standardized system prompt (Section C.1.1) to ensure reproducibility, though the non-determinism in O1 and O3-Mini (temperature fixed at 1.0) introduces some variability that the paper does not quantify through multiple evaluation runs. The calibration error measurement (Section 4.2) is computed across all questions, with the binning methodology from Hendrycks et al. (2022) — the paper does not report confidence intervals or statistical significance tests for accuracy differences between models, which is a limitation given the small absolute differences at the low-accuracy frontier (e.g., distinguishing 2.7% from 4.1% on a 2,500-question dataset).


Main Quantitative Results

Overall Accuracy and Calibration (Table 1)

The headline finding is that all frontier models achieve low accuracy on HLE, with scores ranging from 2.7% (GPT-4O) to 8.5% (DeepSeek-R1) on the full multi-modal dataset, while the highest-performing text-only model, O3-Mini (High), achieves 13.4% accuracy on the text-only subset (Table 1). The key numbers:

ModelAccuracy (%)Calibration Error (%)Multi-modal?
GPT-4O2.789Full
Grok 23.087Full
Claude 3.5 Sonnet4.184Full
Gemini 1.5 Pro4.688Full
Gemini 2.0 Flash Thinking6.682Full
O18.083Full
DeepSeek-R18.573Text-only
O3-Mini (High)13.480Text-only

Several patterns are immediately visible. First, the accuracy range is narrow and low — all models fail on the vast majority of questions, with even the best model (O3-Mini at 13.4%) answering fewer than one in seven text-only questions correctly. Second, calibration is uniformly terrible: RMS calibration errors range from 73% (DeepSeek-R1, the best-calibrated) to 89% (GPT-4O, the worst-calibrated), meaning that on a benchmark where models are almost always wrong, they are almost always confident in their wrong answers. A perfectly calibrated model at 5% accuracy would state roughly 5% confidence on average; these models are stating confidence levels far higher than their actual accuracy, indicating systematic overconfidence.

Third, the ranking across models follows a rough pattern: reasoning models (O1, DeepSeek-R1, O3-Mini, Gemini 2.0 Flash Thinking) outperform non-reasoning models (GPT-4O, Grok 2, Claude 3.5 Sonnet, Gemini 1.5 Pro), but the absolute differences are small. The gap between the best and worst non-reasoning model is only 1.9 percentage points (GPT-4O at 2.7% vs. Gemini 1.5 Pro at 4.6%), and the gap between the best reasoning and best non-reasoning model on the full dataset is only 3.4 points (Gemini 1.5 Pro at 4.6% vs. O1 at 8.0%). On the text-only subset (Table 2 in Section C.2), the gap widens to 8.8 points between GPT-4O (2.3%) and O3-Mini (13.4%), suggesting that the text-only subset may provide clearer discrimination.

Interpretation caveat. The paper explicitly warns against overinterpreting small accuracy differences at this performance level (Section 4.2): "small inflections close to zero accuracy are not strongly indicative of progress." The non-zero accuracy results from two sources: inference non-determinism (the same model can produce different answers on different runs) and the multiple-choice guessing floor (even random chance produces correct answers on approximately 20% of multiple-choice questions). This means that the true "model knowledge" content of a 2.7% vs. 4.6% accuracy difference cannot be reliably distinguished from noise without repeated evaluation runs and confidence intervals — neither of which the paper provides.


Category-Level Performance Analysis (Section C.3, Table 3)

The paper provides a category-wise accuracy breakdown for eight high-level subject groups: Math, Biology/Medicine, Physics, Computer Science/AI, Humanities/Social Science, Chemistry, Engineering, and Other. Table 3 in Section C.3 reports these per-category scores for all models on both the full dataset and the text-only subset. The key observations:

Mathematics shows the widest model discrimination. On the text-only subset, O3-Mini achieves 18.6% accuracy on Math — more than double its overall text-only average of 13.4% and far above GPT-4O's 2.3% on the same category. DeepSeek-R1 (9.1%) and Gemini 2.0 Flash Thinking (8.1%) also show relative strength in math compared to their overall averages. This is consistent with the reasoning models' design focus on mathematical and logical reasoning tasks.

Engineering shows a similar but noisier pattern. DeepSeek-R1 achieves 14.5% on Engineering (text-only), the highest category score for any model, substantially above its 8.5% overall text-only average. However, other models show inconsistent patterns — Claude 3.5 Sonnet achieves 9.7% on Engineering (full dataset), strong relative to its 4.1% overall, while O1 achieves only 4.8% on Engineering (text-only) despite 7.8% overall text accuracy. The small per-category sample sizes (Engineering questions constitute a modest fraction of the 2,500 total) make these numbers noisy, and the paper does not report per-category question counts to assess statistical reliability.

No category is consistently "easy" or "hard" across all models. GPT-4O's highest category accuracy is Biology/Medicine at 6.4% (full dataset); DeepSeek-R1's is Engineering at 14.5% (text-only); O3-Mini's is Math at 18.6% (text-only). This suggests that HLE's difficulty structure is genuinely heterogeneous — different models have different relative strengths across domains, and no single category serves as a universal discriminator. However, the low absolute numbers and small per-category sample sizes make it difficult to draw confident conclusions about relative model strengths from these breakdowns alone.

The "Other" category (miscellaneous subjects not fitting into the seven labeled groups) shows consistently low scores. GPT-4O achieves 2.6% (full), O1 achieves 7.3% (full), and O3-Mini achieves 6.9% (text-only). The lack of a model exceeding single-digit accuracy in this heterogeneous category may reflect the diversity of specialist knowledge required — no single model's training data covers all niche domains.


Token Count Analysis: Reasoning vs. Non-Reasoning Models (Section 4.2, Figure 5, Figure 6)

The paper characterizes the computational cost of model approaches to HLE through average completion token counts, reported visually in Figure 5 (reasoning models) and Figure 6 in Section C.4 (non-reasoning models), with category-level breakdowns. The key findings:

Reasoning models consume substantially more tokens. Figure 5 shows average completion token counts reaching approximately 6,000–9,000 tokens per question for DeepSeek-R1, O1, and Gemini 2.0 Flash Thinking, with substantial variation across categories. O1 shows particularly high token counts in the "Other" category (roughly 8,000–9,000 tokens), while Gemini 2.0 Flash Thinking's highest counts appear in Engineering and Physics (roughly 6,000–7,000 tokens). DeepSeek-R1 shows a more uniform distribution across categories, with most between 6,000 and 8,000 tokens.

Non-reasoning models are token-efficient but less accurate. Figure 6 shows average completion tokens typically between 200 and 1,000 for GPT-4O, Grok 2, Claude 3.5 Sonnet, and Gemini 1.5 Pro — roughly an order of magnitude fewer tokens than reasoning models. Claude 3.5 Sonnet shows the highest token counts among non-reasoning models (approaching 1,000 in some categories), while Grok 2 and GPT-4O cluster in the 200–600 token range. Despite this efficiency, non-reasoning models achieve lower accuracy (2.7%–4.6% vs. 6.6%–13.4% for reasoning models), meaning the additional tokens do correlate with improved performance — just not enough to make any model "good" at HLE.

Token counts vary by subject category within each model. For reasoning models (Figure 5), the token count variation across categories is visually apparent but not enormous — roughly a factor of 2 between the lowest and highest category within each model. For non-reasoning models (Figure 6), the variation is smaller in absolute terms (tens to hundreds of tokens) but proportionally similar. This category-level variation likely reflects differences in question length and complexity (mathematics questions may require longer chains of reasoning, leading to longer completions), but the paper does not provide per-category question lengths to disentangle question complexity from model verbosity effects.

The compute-optimality concern. The paper frames the token count data as motivation for a normative claim: "future models should not only do better in terms of accuracy, but also strive to be compute-optimal" (Section 4.2). This is an interesting meta-point: current reasoning models achieve their modest accuracy gains through dramatically increased inference-time computation (generating thousands of additional reasoning tokens), and the paper argues that genuine capability improvements should ideally reduce, or at least not dramatically inflate, the computational cost per unit of accuracy. However, the paper does not formalize this into a compute-optimality metric (e.g., accuracy per token, or a FLOPs-matched comparison between reasoning and non-reasoning approaches), so the claim remains suggestive rather than demonstrated.


Calibration Error Analysis (Table 1, Table 2)

The calibration results in Table 1 reveal a consistent pattern: all models show severe miscalibration on HLE, with RMS calibration errors ranging from 73% (DeepSeek-R1, the best) to 89% (GPT-4O, the worst).

Interpretation of the calibration error scale. An RMS calibration error of 89% for GPT-4O means that the model's stated confidence systematically deviates from its actual accuracy by nearly the full range of the confidence scale. Since GPT-4O achieves 2.7% accuracy, a well-calibrated model would state roughly 2.7% confidence on average (possibly varying by question). Instead, GPT-4O is stating much higher confidence on most questions — the precise distribution cannot be inferred from the single aggregate metric, but the 89% error magnitude indicates that on a large fraction of questions, the model expresses high confidence (e.g., 70–100%) while being incorrect.

DeepSeek-R1 shows the best calibration at 73% error. This is still extremely poorly calibrated by conventional standards (calibration errors above 20% are generally considered unacceptable for deployment), but it is notably better than GPT-4O's 89%. Whether this represents a genuine improvement in DeepSeek-R1's uncertainty estimation or an artifact of its response formatting (e.g., it may more frequently output moderate confidence scores rather than high ones) cannot be determined from the aggregate metric alone. The paper does not provide calibration curves (reliability diagrams showing binned confidence vs. accuracy) that would reveal the shape of the miscalibration.

Calibration error does not track accuracy cleanly. GPT-4O has the lowest accuracy (2.7%) and the highest calibration error (89%), and DeepSeek-R1 has relatively high accuracy (8.5%) and the lowest calibration error (73%) — consistent with the intuition that more capable models should be better calibrated. But Claude 3.5 Sonnet (4.1% accuracy, 84% calibration error) and Gemini 1.5 Pro (4.6% accuracy, 88% calibration error) break this pattern: Claude has lower accuracy than Gemini but better calibration. The small absolute differences and lack of confidence intervals make it difficult to determine whether these ranking inversions are statistically meaningful.

Text-only calibration (Table 2). On the text-only subset, the calibration error pattern is similar but slightly shifted: GPT-4O at 88%, Grok 2 at 89%, DeepSeek-R1 at 73%, O3-Mini at 80%. The text-only results do not fundamentally change the calibration story — all models remain severely miscalibrated, and the relative ranking is preserved (DeepSeek-R1 best, Grok 2 and GPT-4O worst).

The hallucination connection. The paper explicitly connects these calibration results to confabulation/hallucination: "models frequently provide incorrect answers with high confidence on HLE, failing to recognize when questions exceed their capabilities" (Section 4.2). The argument is that poor calibration on a deliberately difficult benchmark is a direct measure of hallucination tendency — models that cannot recognize the boundaries of their own knowledge will confidently generate plausible-sounding but incorrect outputs when faced with questions beyond their training distribution. This diagnostic use of calibration differs from the standard ML treatment (where calibration is typically measured on in-distribution data) and is one of the paper's more insightful methodological contributions.


Comparison Against Saturated Benchmarks (Figure 1)

Figure 1 provides the visual contrast that motivates the entire paper: a bar chart showing model accuracy on five benchmarks (MMLU, MATH, GSM8K, ARC-Challenge, and HLE) for four models (GPT-4O, Gemini 1.5 Pro, Claude 3.5 Sonnet, O1). The specific accuracy values on prior benchmarks are sourced from external reports (Section C.6): GPT-4O and O1 results from OpenAI's simple-evals repository, Gemini 1.5 Pro from Google's reported results, and Claude 3.5 Sonnet from Anthropic (2024).

The visual juxtaposition is striking: all four models achieve over 85% accuracy on MMLU and over 90% on GSM8K, with several at or near ceiling on ARC-Challenge and MATH — but all four collapse to single-digit accuracy on HLE. The gap between the lowest prior-benchmark accuracy (Claude 3.5 Sonnet on MATH, roughly 75–80% based on the bar chart's visual scale) and the highest HLE accuracy (O1, roughly 8%) is approximately 70 percentage points — a dramatic discontinuity that visually demonstrates benchmark saturation.

Interpretation caveat. Figure 1 uses different evaluation protocols for different benchmarks (the prior-benchmark numbers are taken from official reports with varying few-shot and chain-of-thought settings, while HLE uses the paper's standardized zero-shot chain-of-thought prompt), so the comparison is not strictly controlled. However, the magnitude of the gap is so large that protocol differences are unlikely to explain it — the central claim that HLE is vastly more difficult than saturated benchmarks is robust to reasonable protocol variation.


Ablation Studies and Robustness Checks

HLE is a benchmark paper, not a systems paper, so its "ablation studies" take the form of analyses that probe the benchmark's construction quality and the robustness of its measurements across different conditions.

Text-only vs. full multi-modal evaluation (Table 1 vs. Table 2): Comparing model performance on the full dataset (including 14% image-containing questions) against the text-only subset reveals that the exclusion of multi-modal questions does not dramatically change the performance landscape. GPT-4O drops from 2.7% (full) to 2.3% (text-only) — a small difference suggesting that multi-modal questions are not systematically easier or harder for this model. O1 shows a similar tight range: 8.0% (full) to 7.8% (text-only). O3-Mini and DeepSeek-R1 are evaluated only on the text-only subset (they lack multi-modal capability), so their scores of 13.4% and 8.5% are text-only by necessity. The paper does not report whether image-containing questions are harder or easier than text-only questions within models that can process both (GPT-4O, Claude 3.5 Sonnet, Gemini models), which would be informative about whether HLE's multi-modal component tests genuinely different capabilities or simply adds formatting complexity.

Category-wise performance breakdown (Table 3, Section C.3): The category-level accuracy results in Table 3 serve as a robustness check on the claim that HLE is uniformly difficult across domains. The finding that no model exceeds 19% accuracy in any category (O3-Mini's 18.6% on Math, text-only, is the highest category score reported) supports the claim of uniform difficulty — there are no "easy" categories where models achieve moderate performance that inflates the aggregate metric. However, the per-category sample sizes are not reported, and with 2,500 questions distributed across 8 high-level groups (and over 100 fine-grained subjects), some categories likely contain very few questions (dozens, not hundreds). Category-level accuracies on small samples have large standard errors — a model achieving 18.6% on a category with 50 questions has a 95% confidence interval of roughly ±10 percentage points, making fine-grained ranking unreliable. The paper does not address this statistical limitation.

Expert disagreement rate measurement (Section B.3): The two-round audit on 200-question samples, yielding a final expert disagreement rate of 15.4% for the public set and approximately 18% for the biology/chemistry/health subset, serves as a quality-control ablation. The key finding is that even after the two-round expert review process, approximately one in six questions has an answer that a second qualified expert would dispute or cannot independently verify. The paper's analysis of why disagreement occurs (multi-expert knowledge requirements, questions from unpublished research, multiple-choice "best answer" design) contextualizes this rate rather than treating it as a failure of the benchmark. The comparison to the single-reviewer methodology (which inflates the health-subset rate from 18% to 25%) demonstrates that the disagreement rate is sensitive to audit methodology — a useful robustness insight for future benchmark designers. The paper does not report whether removing questions with expert disagreement changes model accuracy rankings, which would directly test whether the 15.4% disagreement rate materially affects measurement.

Searchability filtering (Section B.2): The procedure of identifying potentially searchable questions (models with web search answer correctly, models without search answer incorrectly) and manually removing those "easily found via web search" serves as an ablation on the "non-searchable" design criterion. The paper's claim that "current frontier model performance on HLE after applying this procedure is similar to their performance on HLE before applying this procedure" is an important robustness check — it suggests that the removed questions were a small fraction of the dataset and that HLE's difficulty is not an artifact of search-resistant but otherwise easy questions. However, this claim is qualitative ("similar") rather than quantitative, and the paper does not report the number of questions removed or pre/post removal accuracy numbers.

Model version sensitivity (Table 4 footnote): The update to Gemini 2.0 Flash Thinking between paper versions (from temperature 0.0 with the deprecated 12-19 model to temperature 0.7 with the 01-21 model) is an accidental robustness check. The paper does not report how much the accuracy changed between versions, but the existence of the change highlights a practical challenge: HLE's accuracy numbers are pinned to specific model API versions that may be deprecated without notice, making exact reproduction difficult over time. The temperature change (0.0 to 0.7) is particularly significant because it introduces substantial additional non-determinism into Gemini 2.0 Flash Thinking's results that was not present in the original evaluation.

Calibration measurement methodology: The decision to set confidence to 100% when no confidence score is extractable from the response (Section C.1.1) is a methodological choice that penalizes models that fail to follow the output format — if a model omits the confidence line entirely, it is treated as maximally confident. The paper does not report how often this occurs for each model, so we cannot determine whether calibration error differences (e.g., DeepSeek-R1's 73% vs. GPT-4O's 89%) partly reflect format-following behavior rather than genuine calibration differences. A model that always outputs "Confidence: 50%" on every question would have a different calibration profile than one that outputs "Confidence: 100%" on every question, and the paper cannot distinguish between these based on the aggregate RMS error alone.

Multiple-choice vs. exact-match difficulty: The paper does not report accuracy separately for multiple-choice (24% of questions) and exact-match (76%) formats, which would be informative about whether the two formats differ in difficulty. Multiple-choice questions have a higher random-guessing floor (20% for five options) and may test recognition rather than generation, while exact-match questions require precise answer generation. A model could achieve higher accuracy on multiple-choice through lucky guessing without any genuine capability difference, inflating aggregate accuracy relative to the model's true knowledge. Reporting separate numbers would enable readers to assess this confound.


Critical Assessment

The experiments presented in this paper serve a different purpose from typical ML systems papers — they evaluate the benchmark's measurement properties rather than a model's performance improvements. As such, the central claims to assess are not "our method achieves state-of-the-art" but rather "HLE provides useful, interpretable measurement where existing benchmarks are saturated." Each claim requires a different standard of evidence.

Claim 1: Frontier LLMs achieve very low accuracy on HLE (2.7%–13.4%)

What the experiments demonstrate: Table 1 and Table 2 unambiguously show that eight state-of-the-art models, evaluated under a standardized protocol, answer between 2.7% and 13.4% of HLE questions correctly. The numbers are very low in absolute terms, and Figure 1 visually demonstrates the contrast with saturated benchmarks where the same models exceed 90% accuracy.

Where the evidence is strongest: The accuracy numbers are based on the full 2,500-question public set, a large evaluation dataset by benchmark standards. The use of an automated O3-MINI judge with explicit instructions (verify answer matching, don't re-solve problems) makes the scoring reproducible. The public release of the dataset enables independent verification of these numbers.

Where the evidence is weaker: The paper does not adequately address the statistical reliability of small accuracy differences at this performance level. With 2,500 questions and accuracies in the 2–13% range, the standard error on a difference between two models' accuracies is approximately sqrt(p1*(1-p1)/2500 + p2*(1-p2)/2500). For the difference between GPT-4O at 2.7% and Claude 3.5 Sonnet at 4.1%, this works out to roughly ±0.5 percentage points standard error — meaning a 95% confidence interval of about ±1 percentage point. The observed difference of 1.4 percentage points is statistically distinguishable from zero, but just barely. For the closer comparisons (GPT-4O at 2.7% vs. Grok 2 at 3.0%, a 0.3 percentage point difference), the difference is within one standard error and cannot be distinguished from noise. The paper's table format implies a clear ranking, but the statistical reality is that the bottom four models (GPT-4O, Grok 2, Claude 3.5 Sonnet, Gemini 1.5 Pro) are essentially indistinguishable given the sample size.

More importantly, the non-determinism confound is not quantified. O1 and O3-Mini are evaluated at temperature 1.0 (they don't support temperature 0.0), meaning their outputs vary across runs. The paper does not report how many evaluation runs were performed or what the variance is. A single evaluation run for O1 could produce 8.0% accuracy; another run might produce 7.2% or 8.8%. Without repeat runs, we cannot determine whether O1's 8.0% and DeepSeek-R1's 8.5% represent a genuine capability difference or within-model variance. For the Gemini 2.0 Flash Thinking model, the paper explicitly changed the model version and temperature between drafts (Table 4 footnote) without reporting the resulting accuracy change — this is a missed opportunity to quantify evaluation variance.

The paper's own caveat — "small inflections close to zero accuracy are not strongly indicative of progress" (Section 4.2) — is appropriate but somewhat at odds with the presentation of a ranked table without confidence intervals. The accuracy ranking from Table 1 should be understood as suggestive rather than definitive, with only the top-vs-bottom distinction (O3-Mini at 13.4% vs. GPT-4O at 2.7%) being clearly reliable.

Claim 2: Models are severely miscalibrated on HLE (RMS calibration errors above 70%)

What the experiments demonstrate: Table 1 reports RMS calibration errors between 73% and 89% across all models, indicating massive overconfidence on a benchmark where models are almost always wrong.

Where the evidence is strongest: The aggregate numbers are striking and consistent across all models. The finding that no model achieves acceptable calibration on HLE (by conventional standards, calibration errors below 10–20% are considered adequate) is robust.

Where the evidence is weaker: The paper does not present calibration curves (reliability diagrams), which are the standard visualization for calibration analysis. Without seeing the binned confidence vs. accuracy plots, we cannot diagnose the shape of the miscalibration. Are models outputting 90–100% confidence on every question (a binary overconfidence pattern), or are they outputting a range of confidence values that simply don't track accuracy? Are there any confidence bins where models are well-calibrated (e.g., low-confidence bins where accuracy is also low)? The single aggregate RMS error collapses this information into one number, losing diagnostic value.

The handling of missing confidence scores (set to 100%, Section C.1.1) is a significant confound. If a model frequently omits the confidence line entirely — perhaps because its training data doesn't emphasize the format, or because it treats the confidence prompt as optional — those questions are scored as 100% confidence regardless of the model's actual uncertainty. This inflates calibration error for models that don't reliably follow the output format. The paper does not report the frequency of missing confidence scores per model, making it impossible to determine whether calibration error differences partially reflect format compliance rather than genuine uncertainty expression.

The calibration measurement is also sensitive to the specific confidence prompt. The paper asks models to state "your confidence score between 0% and 100% for your answer" (Section C.1.1). How models interpret "confidence" in this context is unclear — does it mean confidence that the specific answer is correct, or confidence that the model's reasoning is sound, or something else? Different prompting strategies can produce very different calibration measurements from the same underlying model (Wei et al., 2024; OpenAI, 2024), and the paper does not test alternative confidence elicitation methods (e.g., verbal uncertainty expressions, multiple-choice confidence bins, or sampling-based uncertainty estimation). The reported calibration errors are therefore specific to this particular confidence elicitation method and may not generalize.

Claim 3: HLE provides meaningful discrimination between models where saturated benchmarks cannot

What the experiments demonstrate: Figure 1 shows that the same models achieving 85–95% on MMLU, GSM8K, and ARC-Challenge achieve 2.7–13.4% on HLE, providing a measurement range that separates models rather than clustering them at ceiling.

Where the evidence is strongest: The visual contrast in Figure 1 is compelling: the saturated benchmarks show model accuracies tightly clustered at the top of the scale, while HLE spreads them across the bottom. The inclusion of multiple saturated benchmarks (MMLU, MATH, GSM8K, ARC-Challenge) alongside HLE for the same set of models makes the comparison concrete.

Where the evidence is weaker: The discrimination that HLE provides is bounded by statistical reliability. As discussed above, the accuracy differences between closely-ranked models (GPT-4O vs. Grok 2, Claude 3.5 Sonnet vs. Gemini 1.5 Pro, O1 vs. DeepSeek-R1) are small relative to the expected sampling variance, making the effective number of statistically distinguishable performance tiers perhaps three or four across the eight models, not eight distinct rankings. A benchmark that can only reliably separate models into "very low" (2–5%), "low" (6–9%), and "moderate" (10–14%) buckets provides less discrimination than the ranked table implies.

Additionally, the paper does not compare HLE to other challenging benchmarks that are not saturated — FrontierMath (Glazer et al., 2024), GPQA (Rein et al., 2023), SWE-bench (Jimenez et al., 2024), or MLE-bench (Chan et al., 2024). Showing that HLE is harder than MMLU is trivial; showing that HLE provides discriminative power that these other challenging benchmarks do not would be more informative. The absence of this comparison leaves open the possibility that HLE is one of several similarly difficult benchmarks, and its value relative to alternatives is unestablished.

Claim 4: The expert disagreement rate of 15.4% is acceptable and contextualized

What the experiments demonstrate: The two-round audit on 200-question samples yields a 15.4% expert disagreement rate, with a higher 18% rate on the biology/chemistry/health subset (Section B.3).

Where the evidence is strongest: The audit methodology is described in reasonable detail, and the paper's analysis of why disagreement occurs (multi-expert requirements, unpublished research knowledge, multiple-choice design) provides interpretive context that most benchmarks omit. The comparison between single-reviewer (25% disagreement) and multi-reviewer (18%) methodologies on the health subset demonstrates that the disagreement rate is methodology-dependent, which is an honest and informative admission.

Where the evidence is weaker: The sample size (200 questions audited) is small relative to the 2,500-question dataset. The 15.4% estimate has a standard error of approximately sqrt(0.154 * 0.846 / 200) ≈ 2.6 percentage points, giving a 95% confidence interval of roughly 10.2–20.6% — quite wide. The estimate is reliable enough to establish that disagreement is non-trivial but not precise enough to serve as a tight quality metric.

The paper does not report what happens to model accuracy if disputed questions are removed. This is a critical missing analysis. If the 15.4% of questions with expert disagreement are systematically different from the remaining 84.6% — for example, if they are harder, or if models get them right at different rates — then the disagreement introduces systematic measurement bias, not just noise. Removing disputed questions and recomputing model accuracies would directly test whether the disagreement rate matters for the benchmark's primary function (ranking models). Without this analysis, we cannot assess whether the expert disagreement is a tolerable imperfection or a meaningful confound.

Claim 5: HLE is non-searchable and resists memorization

What the experiments demonstrate: The searchability audit identified potentially searchable questions (models with search answer correctly, without search answer incorrectly) and manually removed those easily found via web search. The paper claims this did not substantially change model performance on HLE (Section B.2).

Where the evidence is strongest: The searchability audit procedure is methodologically sound — the use of models with and without search access as a detection mechanism for searchable questions is clever and practical.

Where the evidence is weaker: The paper provides no quantitative data about the searchability audit: how many questions were flagged as potentially searchable, how many were removed, what the pre/post removal accuracy numbers were, or what fraction of the dataset this affected. The claim that "current frontier model performance on HLE after applying this procedure is similar to their performance on HLE before applying this procedure" is qualitative and untestable from the information provided. A robust audit would report: "X questions (Y% of the dataset) were removed; model accuracy changed from A% to B%, a difference that is [not] statistically significant."

Missing Experiments That Would Have Strengthened the Paper

Several analyses are conspicuous by their absence and would substantially improve the benchmark's evidentiary foundation:

  1. Repeated evaluation runs to quantify non-determinism variance. Running each model 5–10 times on the full dataset (or a large stratified sample) and reporting mean accuracy with confidence intervals would directly address the single-run reliability concern. This is especially important for O1, O3-Mini, and DeepSeek-R1, which operate at non-zero temperatures.

  2. Calibration curves (reliability diagrams). The aggregate RMS error conceals the shape of miscalibration. Binned confidence vs. accuracy plots would reveal whether models are uniformly overconfident, whether there are any well-calibrated confidence regions, and whether the calibration failure is primarily driven by a specific confidence range (e.g., models that always say 90–100% confident).

  3. Accuracy disaggregated by format (multiple-choice vs. exact-match). Reporting accuracy separately for the 24% multiple-choice and 76% exact-match questions would reveal whether the two formats differ in difficulty, whether models benefit differentially from the multiple-choice guessing floor, and whether the benchmark's difficulty is driven by one format or both.

  4. Accuracy with disputed questions removed. Recomputing model accuracy after excluding the ~15.4% of questions with expert disagreement would test whether the disagreement rate affects model rankings. If rankings are stable, the disagreement is benign noise; if they shift, it represents a measurement confound.

  5. Per-category confidence intervals. For the category-level analysis in Table 3, reporting per-category question counts and confidence intervals would enable readers to assess which category-level differences are statistically meaningful.

  6. Comparison to other difficult benchmarks. Evaluating the same models on GPQA, FrontierMath, or other frontier-difficulty benchmarks under the same evaluation protocol would contextualize HLE's difficulty and discrimination relative to alternatives.

  7. Human baseline on HLE. The paper does not report how human experts perform on HLE. Without a human baseline, we cannot calibrate what "13.4% accuracy" means — is O3-Mini performing at the level of a strong graduate student, an average undergraduate, or somewhere else? The benchmark is designed to be at the frontier of human knowledge, but we don't know where on that frontier current models sit relative to humans. A small human evaluation (even 50–100 questions solved by domain experts) would provide essential calibration.

  8. Sensitivity to prompting strategy. The paper uses a single chain-of-thought prompt with structured output formatting. Testing alternative prompting strategies (few-shot examples, different system prompts, self-consistency decoding, best-of-N sampling) would reveal how sensitive HLE accuracy is to evaluation protocol — a crucial consideration for a benchmark intended to measure underlying capability rather than prompt engineering skill.

Summary Assessment

The experimental results in this paper convincingly establish the central qualitative finding: frontier LLMs perform poorly on HLE compared to saturated benchmarks, and they are severely miscalibrated in their confidence estimates. This is sufficient to motivate HLE's existence as a benchmark that saturating evaluations cannot provide.

However, the paper's quantitative claims — the specific accuracy ranking of models, the numerical calibration error values, the category-level performance patterns — are presented with more precision than the experimental design can support. The single-run evaluation without confidence intervals, the unquantified non-determinism, the missing calibration curves, the small per-category samples, and the unanalyzed expert disagreement rate all introduce uncertainty that the paper's tables and rankings do not reflect. The paper's own caveat about small accuracy differences near zero not being "strongly indicative of progress" is correct but should be applied more thoroughly to the presented results.

The benchmark's value proposition — providing measurement resolution where existing benchmarks are saturated — is demonstrated at a coarse level: HLE can distinguish models that achieve 13% accuracy from those achieving 3% accuracy. Whether it can reliably distinguish 3% from 5%, or 8% from 9%, is unestablished. For HLE to serve as the precision instrument the paper envisions, future evaluations should adopt the statistical rigor (repeated runs, confidence intervals, disaggregated metrics) that the current analysis lacks.

6. Limitations and Trade-offs

6.1 The Expert Disagreement Rate of 15.4% Represents Unresolved Ground-Truth Uncertainty

The assumption or constraint. The HLE pipeline assumes that after two rounds of expert review, the ground-truth answers are sufficiently reliable for automated evaluation. However, the paper's own audit process reveals that this assumption is imperfect. Post-release auditing on 200-question samples produces "a final estimated expert disagreement rate of 15.4% for the public set," rising to "approximately 18%" for a targeted biology, chemistry, and health subset (Section B.3). The paper acknowledges three sources of this disagreement: questions requiring "a critical piece of information, such as a decades-old paper or a foundational concept not immediately apparent to others" that only some experts possess (multi-expert requirement); questions "based on insights from the direct, hands-on experiments of its contributors" that are "often difficult to verify through standard literature searches" (unpublished research knowledge); and multiple-choice questions where "researchers sometimes leverage the multiple-choice format with the objective of identifying the most plausible answer among the provided options" rather than a uniquely correct answer in an absolute sense (best-answer design).

The consequence. The 15.4% expert disagreement rate sets a fundamental ceiling on the interpretability of model accuracy scores on HLE. If approximately one in six questions has an answer that qualified domain experts would dispute or cannot independently verify, then a model's accuracy on HLE is not measuring the fraction of questions where the model knows the correct answer in a clean, objective sense — it is measuring the fraction where the model's output matches what the question's author and the reviewers who accepted it consider correct. This is not a problem at current accuracy levels (2.7–13.4%), where almost all failures reflect genuine model limitations rather than benchmark noise. But it becomes a critical measurement confound as models improve. If a future model achieves, say, 85% accuracy on HLE, we cannot determine whether the remaining 15% represents model errors or benchmark noise — the two are confounded at that performance level. The paper's comparison between single-reviewer and multi-reviewer methodologies on the health subset demonstrates the sensitivity: moving from multi-reviewer to single-reviewer disagreement determination raises the rate from 18% to 25%, purely through audit methodology change (Section B.3). This implies that the "true" expert error rate is not directly observable and depends on how disagreement is adjudicated.

What evidence exists in the paper. The two-round audit on 200-question samples provides the empirical basis for the 15.4% and 18% estimates (Section B.3). The paper's analysis of the sources of disagreement is qualitative but substantive, identifying three distinct causes. However, the paper does not report what happens to model accuracy rankings if disputed questions are removed. This is a critical missing analysis: if the 15.4% of questions with expert disagreement are systematically different (harder, easier, or from specific subject categories) from the remaining 84.6%, then the disagreement introduces systematic measurement bias that could affect model rankings. The sample size of 200 audited questions is small relative to the 2,500-question dataset, giving the 15.4% estimate a standard error of roughly sqrt(0.154 × 0.846 / 200) ≈ 2.6 percentage points, or a 95% confidence interval of approximately 10.2–20.6% — broad enough that the true disagreement rate could be substantially higher or lower than the point estimate.

Mitigation status. The paper is unusually transparent about this limitation, both in reporting the disagreement rate explicitly and in analyzing its sources. It contextualizes the rate as "in line with what is observed in other challenging, expert-grade machine learning benchmarks," citing HealthBench (Arora et al., 2025) as observing similar patterns (Section B.3). The HLE-Rolling mechanism (Section B.3) provides a pathway for correcting identified errors over time. However, the paper does not propose a methodology for distinguishing model errors from benchmark noise at high accuracy levels, nor does it define a performance threshold beyond which accuracy gains on HLE should be treated as uninterpretable due to ground-truth uncertainty. The "best answer among options" design pattern — where multiple-choice questions lack a uniquely correct answer — is acknowledged but not systematically quantified (what fraction of multiple-choice questions adopt this pattern?). The fundamental tension between question difficulty and answer verifiability (questions drawing on unpublished research experience are harder for models but also harder for humans to verify) is not resolved.


6.2 Difficulty Estimation Cost Is Unaccounted For in the Benchmark's Headline Metrics

The assumption or constraint. HLE's defining design feature — the automated LLM difficulty filter that requires questions to stump frontier models before they proceed to human review — is what guarantees the benchmark's difficulty. However, the paper does not account for the cost of verifying difficulty in its framing of HLE as an evaluation instrument. The filter logged "over 70,000 attempts" across six frontier LLMs (Section 3.2), producing approximately 13,000 questions that passed — roughly 19% of submissions. This means that for every question accepted into the review pipeline, approximately 4.4 LLM attempts were consumed in filtering. The paper does not quantify the total inference compute expended on filtering, the cost of maintaining API access to up-to-date frontier models for filtering, or the sustainability of this filtering paradigm as models improve (more capable models will accept more questions, requiring even more aggressive filtering or more sophisticated anti-guessing measures).

The consequence. The unaccounted filtering cost creates a hidden maintenance burden for HLE's continued use. As frontier models improve, the pool of questions that stump them will shrink — the same questions that GPT-4O and Claude 3.5 Sonnet failed on during HLE's construction may become answerable by GPT-5 or Claude 4. If HLE is used as an ongoing benchmark (not just a one-time measurement), new questions must be continuously added through HLE-Rolling (Section B.3), and each new question requires the same expensive LLM filtering against the current frontier models. The cost of this filtering scales with the number of new questions and the number of filter models, and the paper provides no framework for estimating or budgeting this cost. For research groups or companies that want to deploy HLE-like evaluations internally (with proprietary models or on custom question sets), the filtering infrastructure represents a significant practical barrier — they would need access to multiple frontier model APIs and the budget to run tens of thousands of inference calls just to establish that their questions are appropriately difficult.

The filtering cost also introduces a temporal coupling between HLE's construction era models and its difficulty calibration. The benchmark was filtered against models available in late 2024 (GPT-4O, Claude 3.5 Sonnet, Gemini 1.5 Pro, O1). If a model released in 2026 answers 50% of HLE questions correctly, we cannot distinguish between genuine capability improvement and model-specific weaknesses in the filtering-era models that allowed certain questions through despite not being "truly" at the frontier of human knowledge. The difficulty floor is pinned to the capabilities of late-2024 models, not to an absolute standard of human expertise.

What evidence exists in the paper. The paper reports the 70,000+ attempts figure and the ~13,000 accepted questions (Section 3.2), but does not quantify the associated inference cost. The acceptance rate of ~19% is reported but not analyzed in terms of filtering efficiency. The HLE-Rolling mechanism (Section B.3) is described as a mechanism for refreshing questions, but the paper does not estimate the filtering cost for new question batches. The late contributions process (Section B.2) mentions that "thousands of submissions" were received post-release and "manually reviewed by organizers" — but whether these underwent the same LLM difficulty filtering as the original submissions is not specified (the paper says they were added to "a second held-out private set" rather than the public benchmark, so they may not be subject to the same filtering standard).

Mitigation status. The paper does not attempt to mitigate this limitation. There is no discussion of filtering cost amortization, no proposal for cheaper difficulty estimation methods (e.g., using a single model rather than six, or using smaller models as difficulty proxies), and no recommendation for how frequently the filtering should be repeated as models improve. The HLE-Rolling mechanism implicitly assumes that the filtering cost is manageable enough to support continuous question addition, but this assumption is not justified or costed.


6.3 Single-Run Evaluation Without Statistical Reliability Quantification Limits Discriminative Power

The assumption or constraint. The paper's evaluation of model accuracy on HLE (Table 1, Table 2) is based on a single run per model under a single prompting strategy (standardized zero-shot chain-of-thought with structured output formatting, Section C.1.1). The paper does not report confidence intervals, standard errors, or variance estimates for any accuracy number. For O1 and O3-Mini, which "only support temperature 1.0" (Table 4 footnote), model outputs are non-deterministic — the same model can produce different answers on different runs. The Gemini 2.0 Flash Thinking model was changed between paper versions from "temperature 0.0 with the now deprecated 12-19 model" to "temperature 0.7" with a newer model (Table 4 footnote), and the paper does not report how much the accuracy changed between these evaluation conditions.

The consequence. The single-run evaluation without variance quantification means that the fine-grained model ranking implied by Table 1 cannot be reliably interpreted. At the low accuracy levels observed (2.7–13.4%), the absolute differences between closely ranked models are very small. The difference between GPT-4O (2.7%) and Grok 2 (3.0%) is 0.3 percentage points on 2,500 questions — approximately 7–8 questions out of 2,500. The standard error on this difference is approximately sqrt(0.027 × 0.973 / 2500 + 0.030 × 0.970 / 2500) ≈ 0.47 percentage points, meaning the observed difference is within one standard error and not statistically distinguishable from zero. For models with similar accuracies (GPT-4O vs. Grok 2, Claude 3.5 Sonnet vs. Gemini 1.5 Pro, O1 vs. DeepSeek-R1), the paper's table format presents an apparent ranking that the experimental design cannot support. For non-deterministic models (O1 at temperature 1.0, O3-Mini at temperature 1.0, Gemini 2.0 Flash Thinking at temperature 0.7), the single-run accuracy is a sample from an unknown distribution, and without repeat runs we cannot determine whether, e.g., O1's 8.0% and DeepSeek-R1's 8.5% represent a genuine capability difference or within-model variance.

The paper acknowledges this concern partially: "small inflections close to zero accuracy are not strongly indicative of progress" (Section 4.2). However, this caveat is applied at the level of trending interpretation (don't overinterpret small improvements) rather than to the ranked table itself (which models can be reliably distinguished from which others). The practical consequence is that a user comparing two models on HLE — say, selecting between Claude 3.5 Sonnet (4.1%) and Gemini 1.5 Pro (4.6%) for a deployment decision — cannot determine from the reported numbers whether the observed difference reflects a genuine capability gap or evaluation noise.

What evidence exists in the paper. The paper provides no variance estimates. The accuracy numbers are point estimates without error bars. The Gemini 2.0 Flash Thinking version change demonstrates that evaluation conditions can shift between runs, but the magnitude of the shift is not reported. The paper's own discussion of why models show non-zero accuracy on a filtered benchmark — "inherent noise in model inference — models can inconsistently guess the right answer or guess worse than random chance for multiple choice questions" (Section 4.2) — acknowledges the non-determinism problem but does not quantify its magnitude for any specific model or accuracy number.

Mitigation status. The paper does not mitigate this limitation. There are no repeat evaluation runs, no bootstrap confidence intervals, no per-question variance analysis, and no recommendation for how many evaluation runs would be needed for reliable measurement at different accuracy levels. The public release of the dataset enables independent researchers to conduct repeat evaluations, but the paper's own analysis does not establish the statistical reliability floor.


6.4 No Human Baseline Calibrates What "13.4% Accuracy" Means in Practical Terms

The assumption or constraint. HLE is designed to test knowledge at "the frontier of human knowledge" (Section 1), with questions that "usually (but do not always need to) be at a graduate / PhD level or above" (Section C.7.1). The paper reports model accuracies ranging from 2.7% to 13.4% (Table 1) and presents these as evidence of a "significant gap between current LLM capabilities and the expert human frontier on closed-ended academic questions" (Abstract). However, the paper does not report how human experts perform on HLE. There is no human evaluation — no group of domain experts who solved a subset of HLE questions to establish what human-level performance looks like on this benchmark.

The consequence. Without a human baseline, the numerical accuracy scores on HLE have no anchor in human capability. We cannot determine whether O3-Mini's 13.4% accuracy means it performs at the level of an advanced undergraduate, a first-year PhD student, a postdoctoral researcher, or somewhere else entirely relative to the humans who created the questions. The paper's claim that HLE measures "the gap between current AI capabilities and human expertise" (Section 5) is undemonstrated — we know the models fail on most questions, but we do not know how often human experts would succeed. If domain experts would achieve only 30–40% accuracy on HLE (plausible given the extreme specialization and breadth — a mathematician might struggle with a classics question, and vice versa), then O3-Mini's 13.4% would represent a substantial fraction of expert-level performance rather than near-zero capability. Conversely, if experts achieve 80–90% on questions within their domain, then model performance is truly minuscule. The paper's interpretive framework depends on assumptions about human performance that are never validated.

This absence also limits HLE's utility for capability forecasting. If we had a human baseline (e.g., "PhD students in relevant fields achieve X% accuracy, postdocs achieve Y%, full professors achieve Z%"), we could track model progress relative to meaningful human milestones. Without it, we can only say that models are improving — we cannot say whether they are approaching the level of a competent graduate student or still far below it.

What evidence exists in the paper. The paper provides no human evaluation data. The expert disagreement rate of 15.4% (Section B.3) provides indirect evidence about human performance — if experts disagree on 15.4% of answers, this suggests that individual experts might get a non-trivial fraction of questions "wrong" relative to the benchmark's ground truth, but this is qualitatively different from a direct measurement of how often experts can correctly solve HLE questions. The contributor demographics — "mostly professors, researchers, and graduate degree holders" (Section 3.1) — establish that question authors are genuine experts, but tell us nothing about how other experts would perform on these questions. The prize pool structure ($5,000 for top-50 questions) incentivizes question difficulty, which may push authors toward questions that even their close colleagues would find challenging.

Mitigation status. The paper does not attempt to mitigate this limitation. There is no human evaluation study, no plan for one, and no discussion of why one was not conducted. The paper's framing in Section 5 ("High accuracy on HLE would demonstrate expert-level performance on closed-ended, verifiable questions") implicitly assumes that high accuracy on HLE is equivalent to expert-level performance, but this equivalence requires demonstrating that experts themselves achieve high accuracy on HLE — an assumption that may or may not hold given the benchmark's deliberate difficulty and extreme breadth. The absence of a human baseline is perhaps the most significant gap between the paper's claims about measuring the "gap between current AI and expert human capabilities" and the evidence needed to support those claims.


6.5 Single Prompting Strategy and Single-Judge Evaluation Constrain Generality of Accuracy Measurements

The assumption or constraint. The evaluation protocol uses a single standardized system prompt that structures model responses into "Explanation," "Answer," and "Confidence" sections (Section C.1.1). Correctness is judged by a single automated judge — O3-MINI (o3-mini-2025-01-31) with structured decoding enabled (Section C.1.1) — which extracts the final answer, compares it against the ground truth, and determines correctness. The paper does not evaluate sensitivity to prompting strategy (few-shot vs. zero-shot, different system prompts, self-consistency decoding, or best-of-N sampling) or to judge model choice (whether a different judge model would make different correctness determinations).

The consequence. The reported accuracy numbers are specific to the intersection of a particular prompting strategy and a particular automated judge, and may not generalize to alternative evaluation protocols. If a model would perform substantially better on HLE under a different prompt (e.g., few-shot examples, explicit instruction to show step-by-step work, or iterative refinement), the paper's zero-shot chain-of-thought numbers would underestimate that model's capability. Conversely, a model that happens to be well-aligned with the particular structured output format used might appear stronger than one that possesses the relevant knowledge but struggles with the format constraint.

The single-judge design introduces a subtler confound: O3-MINI is simultaneously the best-performing model on HLE (13.4% accuracy on text-only, Table 2) and the judge of all other models' answers. This creates a structural asymmetry: if O3-MINI makes systematic errors in answer verification (e.g., rejecting equivalent mathematical expressions it cannot recognize as equivalent, or accepting answers it misclassifies as correct), those errors affect the measured accuracy of every other model but not O3-MINI's own accuracy (since O3-MINI's evaluation must use a different, possibly less capable, judge or direct comparison). The paper does not report O3-MINI's accuracy as a judge on HLE (i.e., the fraction of answer verifications where O3-MINI's correctness determination matches human expert judgment), so we cannot assess whether judge errors are frequent enough to materially affect model rankings. The paper provides one example of successful equivalence detection (the cot(π/n) / 2 cot(π/(2n)) = cos(π/n) / (2(1+cos(π/n))) example in Section C.1.1), but this is an illustrative positive case, not evidence about the judge's overall reliability.

What evidence exists in the paper. The paper reports accuracy under a single prompt (Section C.1.1) and a single judge (O3-MINI, Section C.1.1). The structured output format is justified by the need for automated answer extraction and calibration measurement. No prompt sensitivity analysis is conducted. No judge reliability analysis (e.g., human-judge agreement rate on a sample of HLE questions) is reported. The paper does not discuss alternative judge models or the rationale for choosing O3-MINI specifically (beyond it being a capable model). The example of successful equivalence detection in Section C.1.1 is the only evidence provided about judge behavior, and it is explicitly presented as an example rather than a systematic evaluation.

Mitigation status. The paper does not mitigate this limitation. The public release of the dataset enables independent researchers to evaluate models under different prompts and with different judges, but the paper's own claims about model capabilities are tied to a single, unvalidated evaluation protocol. The use of O3-MINI as judge while it is also the highest-scoring evaluated model is an unusual methodological choice that the paper does not justify or problematize. A standard mitigation would be: (1) report judge reliability against human evaluation on a held-out sample, (2) evaluate sensitivity to prompt variation, and (3) ensure the judge model is not simultaneously the subject of evaluation (or, if it is, use an alternative judge for its own scores and report the difference).


6.6 The Benchmark's Difficulty Depends on Filtering Against Models That Are Rapidly Obsolete

The assumption or constraint. HLE's difficulty is enforced by the automated LLM filter that required questions to stump GPT-4O, Claude 3.5 Sonnet, Gemini 1.5 Pro, and O1 (Section 3.2, Section B.1). This filter anchors the benchmark's difficulty floor to the capabilities of models available in late 2024. The paper explicitly acknowledges that HLE's difficulty is temporally bounded: "it is plausible that models could exceed 50% accuracy on HLE by the end of 2025" (Section 5).

The consequence. The difficulty achieved by the filter has a limited half-life. As new models surpass the filtering-era models, the gap between the filter standard and the evaluation standard widens. A question that genuinely stumped GPT-4O in 2024 — requiring knowledge or reasoning beyond its capabilities — may be solvable by a 2026 model not because the question is "easy" but because the 2026 model possesses capabilities that the filter models did not. This is by design: the benchmark is meant to be progressively solvable as models improve. But the consequence is that progress on HLE is interpreted relative to the capabilities of late-2024 models, not relative to a fixed standard of human expertise. When a model achieves 50% accuracy on HLE, we learn that it has surpassed the filter-era models by a large margin — but we do not learn whether it has reached genuine expert-level performance, because the filter only guarantees that late-2024 models (which were not at expert level) could not answer the questions. The difficulty ceiling of HLE is the "frontier of late-2024 model capabilities," which is lower than the "frontier of human knowledge" that the paper claims to target.

This temporal coupling also interacts with the HLE-Rolling mechanism (Section B.3). To maintain difficulty as the benchmark saturates, new questions must be filtered against current frontier models. But the filtering era models for the original HLE (late 2024) are different from the filtering era models for HLE-Rolling (2025, 2026, etc.). Questions added to HLE-Rolling in 2026 will have been filtered against more capable models — meaning they will be genuinely harder than questions in the original HLE, which were filtered against weaker 2024 models. This creates a difficulty drift: the benchmark's composition shifts over time, with newer questions being harder (filtered against stronger models), making it impossible to directly compare accuracy numbers across different HLE-Rolling versions. A model achieving 15% on HLE-2025 and 15% on HLE-2027 would have genuinely improved, because the 2027 benchmark is harder — but how much harder? The paper provides no framework for calibrating difficulty across versions.

What evidence exists in the paper. The paper reports the specific models used in filtering (Section 3.2, Section B.1) and the acceptance rate (~19%, Section 3.2). The acknowledgment that models could exceed 50% accuracy by end of 2025 (Section 5) demonstrates awareness of the temporal limitation. The HLE-Rolling mechanism (Section B.3) is presented as the solution to saturation, but the cross-version comparability problem is not discussed. The paper does not report how model performance on HLE changed between the initial filtering (which used GPT-4O, Claude 3.5 Sonnet, Gemini 1.5 Pro, O1) and the final evaluation (which includes those same models at potentially different versions, plus newer models like DeepSeek-R1 and O3-Mini) — such a comparison would reveal the temporal drift in difficulty relative to the filter standard.

Mitigation status. The paper partially mitigates this through the HLE-Rolling mechanism, which provides a pathway for refreshing the benchmark. However, it does not address the cross-version comparability problem or propose a difficulty calibration methodology (e.g., maintaining a fixed subset of questions across versions as a difficulty anchor, or reporting accuracy against a fixed set of reference models from each filtering era). The paper's claim that HLE is "the final closed-ended academic benchmark of its kind" (Section 1) must be understood with the caveat that "final" means "final for the model generation that succeeded the filtering-era models" — not final in an absolute sense, and not final without continuous maintenance and re-filtering against improving models. The temporal coupling between HLE's difficulty and the specific models used in its construction is a fundamental constraint on the benchmark's lifespan as a discriminative instrument.

7. Implications and Future Directions

How This Work Changes the Landscape

HLE shifts the benchmarking field from a reactive posture — waiting for models to saturate existing tests, then scrambling to build harder ones — to a structurally preemptive one. The core methodological shift is the inversion described in Section 3.2: difficulty is not predicted by expert judgment and then hoped for; it is empirically verified before human review begins, using the best available models as measurement instruments. This changes the epistemic status of benchmark difficulty from an aspirational quality to a gated prerequisite. When a new model fails on HLE, we know the questions were demonstrably beyond the capabilities of the models used in construction — the failure is anchored to an empirical fact about the world, not to a designer's intuition about what "should be" hard.

This is an incremental refinement of the benchmarking methodology, not a paradigm shift. The individual components — expert-written questions, multi-stage review, adversarial filtering — all exist in prior work (GPQA, WMDP, Dynabench, FrontierMath). HLE's contribution is their integration at unprecedented scale (2,500 questions, nearly 1,000 contributors, 500 institutions, 50 countries) with the LLM difficulty gate placed before human review rather than after, combined with calibration measurement as a first-class metric and transparency about expert disagreement rates. The paper does not propose a new way of measuring intelligence or a new theoretical framework for evaluation; it refines and scales existing approaches to a degree that qualitatively changes their usefulness.

Reconciling prior contradictory findings. The paper resolves a latent tension in the benchmarking literature: why do some expert-designed benchmarks (MMLU) saturate within years while others (FrontierMath, ARC-AGI) remain difficult? The answer is not primarily about question content but about the relationship between question selection and model capabilities at construction time. Benchmarks designed against weaker models (MMLU was built in the GPT-3 era) become easy when models advance; benchmarks filtered against stronger models (FrontierMath used GPT-4-level models) remain harder longer. HLE makes this relationship explicit and quantifiable: the filtering acceptance rate of ~19% (13,000 accepted from 70,000+ attempts, Section 3.2) provides a concrete measure of how aggressively the benchmark was filtered against its construction-era frontier. Future benchmarks can use this metric to communicate their expected difficulty half-life — a benchmark with a 5% acceptance rate against GPT-5 will remain informative longer than one with a 50% acceptance rate against the same filter.

Redirection of research attention. HLE's results make several research directions more attractive and others less so:

  • More attractive: robust calibration under distribution shift. The finding that all models show RMS calibration errors above 70% on HLE (Table 1) — with GPT-4O reaching 89% — demonstrates that current uncertainty estimation methods fail catastrophically when models are pushed beyond their training distribution. This is not a marginal calibration problem; it is a qualitative failure mode where models confidently hallucinate on almost every question. The paper converts calibration from an academic metric into a safety-relevant diagnostic: on HLE, uncalibrated models are not just "slightly overconfident" — they are systematically misleading. This reframes calibration research as urgently practical rather than theoretically tidy.

  • More attractive: efficient difficulty estimation. The paper's explicit acknowledgment that HLE's difficulty estimation (70,000+ LLM attempts across six models for filtering, plus 2,048 samples for per-question difficulty binning in the compute-optimal framework of related work) is too expensive for routine deployment creates a clear research target: cheap, accurate difficulty prediction. The paper demonstrates that this is possible in principle (predicted difficulty bins track oracle bins in Figures 4 and 8) but not yet practical at scale.

  • More attractive: verifier robustness under optimization pressure. HLE's finding — that models confidently assert wrong answers rather than expressing uncertainty — is a manifestation of the same verifier over-optimization phenomenon documented in the companion analysis paper: models optimize for outputs that look plausible under their training distribution, and when that distribution doesn't cover frontier-difficulty questions, the optimization produces confident but wrong answers.

  • Less attractive: benchmark construction without difficulty gating. HLE's stark contrast with saturated benchmarks in Figure 1 — the same models at 90%+ on MMLU and 2.7–13.4% on HLE — makes it difficult to justify building new general-coverage academic benchmarks without an LLM difficulty filter. A benchmark released in 2025 that does not verify that current frontier models fail on its questions will face immediate skepticism about its expected useful lifespan.

  • Less attractive: accuracy-only evaluation on difficult benchmarks. HLE demonstrates that accuracy alone, while necessary, is insufficient for interpreting model capabilities at the frontier. The calibration dimension (Section 4.2) reveals the qualitative difference between a model that fails cautiously and one that fails confidently — a distinction invisible to accuracy metrics but critical for deployment decisions.

The calibration-as-safety-diagnostic reframing. Prior work on calibration (Hendrycks et al., 2022; Wei et al., 2024) treated it primarily as a statistical property — does stated confidence match empirical accuracy? HLE transforms calibration into a behavioral safety diagnostic: on a benchmark where models are almost always wrong, calibration measures whether models know they are wrong. A model that states 5% confidence while achieving 5% accuracy is behaving safely — it can be trusted to defer to humans or flag uncertainty. A model that states 80% confidence while achieving 5% accuracy (the regime all frontier models occupy on HLE) is behaving dangerously — it will confidently mislead users in high-stakes domains. This reframing connects benchmark evaluation directly to deployment safety in a way that accuracy-only metrics cannot.

Where the shift is bounded. HLE's impact is limited to a specific measurement niche: closed-ended, verifiable academic questions with broad subject coverage. The paper is explicit about this boundary (Section 5): "HLE tests structured academic problems rather than open-ended research or creative problem-solving abilities." High accuracy on HLE would demonstrate expert-level performance on exam-style academic questions — a meaningful capability, but not "artificial general intelligence" or autonomous research ability. The paper's most important methodological contribution (difficulty-as-prerequisite) does not automatically transfer to open-ended evaluation (where correctness is subjective), interactive tasks (where the space of possible interactions is unbounded), or safety evaluations (where the "correct" behavior may be context-dependent). For these domains, different methodological innovations are needed.


Follow-Up Research This Work Enables

Human expert baseline on HLE. The most immediate gap in HLE's measurement framework is the absence of human performance data. A targeted study would recruit domain experts (ideally from the same demographic as the question contributors — professors, researchers, graduate students) to solve a stratified sample of 200–300 HLE questions within their fields of expertise, under time-limited conditions that match model evaluation (one attempt per question, no external resources). The key measurements would be: (1) expert accuracy per subject category, establishing what "expert-level performance" actually means on HLE; (2) expert calibration (how confident are human experts in their answers, and does their confidence track accuracy?); (3) the correlation between expert accuracy and the expert disagreement rate from Section B.3 — do questions where experts disagree show lower expert accuracy, confirming that disagreement signals genuine ambiguity rather than just asymmetric knowledge? Without these numbers, the paper's framing of HLE as measuring "the gap between current AI capabilities and expert human performance" rests on untested assumptions about human capability on these questions. A finding that human experts achieve only 40–60% accuracy on in-domain HLE questions would dramatically reframe what model scores of 2.7–13.4% mean — they would represent a substantial fraction of expert performance rather than near-zero capability.

Judge reliability and prompt sensitivity analysis. The paper's evaluation protocol uses a single judge (O3-MINI) and a single prompting strategy. A systematic robustness study would: (1) evaluate judge reliability by having 2–3 human experts independently verify correctness for a stratified sample of 500 model answers (spanning the accuracy range from clearly correct to clearly incorrect to ambiguous), computing judge-human agreement rates and identifying systematic judge error patterns (e.g., does the judge reject equivalent mathematical expressions, accept plausible-sounding but wrong reasoning, or mishandle domain-specific notation?); (2) evaluate prompt sensitivity by testing 4–6 alternative prompting strategies (few-shot with 3 domain-matched examples, chain-of-thought without structured formatting, self-consistency with majority voting at n=5, best-of-N with the PRM verifier at n=16) on a subset of 500 questions for at least three models spanning the accuracy range (GPT-4O, Claude 3.5 Sonnet, O3-Mini), measuring how much accuracy shifts under each protocol. This study would establish the measurement uncertainty budget for HLE: how much of a reported accuracy difference can be attributed to genuine capability gaps versus evaluation protocol choices. The paper's finding that structured output formatting matters (the system prompt separates reasoning from answer) and that non-determinism exists (O1, O3-Mini at temperature 1.0) makes this analysis essential for interpreting the Table 1 rankings.

Temporal difficulty drift analysis as HLE-Rolling deploys. As HLE-Rolling adds new questions filtered against progressively stronger models, the benchmark's effective difficulty will change. A longitudinal study would: track a fixed set of reference models (e.g., GPT-4O, Claude 3.5 Sonnet, O1 — the filtering-era models) on each HLE-Rolling version as it is released, computing their accuracy on the original 2,500 questions versus the newly added questions to quantify how much harder the new questions are. The key metric is the difficulty inflation rate: if new questions are filtered against models that achieve 20% on the original HLE, they should be genuinely harder, and the reference models' accuracy on new questions should be lower than on old questions. By how much? This study would establish whether HLE-Rolling creates a smoothly comparable difficulty trajectory or introduces version discontinuities that make cross-version accuracy comparisons unreliable. A related sub-study would maintain a fixed "anchor set" of 100–200 questions across all HLE-Rolling versions, never retiring them, to serve as a calibration baseline — accuracy on the anchor set should improve monotonically as models advance, while accuracy on the full rotating benchmark would reflect both model improvement and increasing question difficulty.

Calibration curve analysis and failure mode taxonomy. The paper reports aggregate RMS calibration error (73–89%) but not calibration curves. A diagnostic study would compute full reliability diagrams for each model on HLE, binning predictions by stated confidence (0–10%, 10–20%, ..., 90–100%) and plotting observed accuracy in each bin. The key questions: (1) Are there any confidence bins where models are well-calibrated (e.g., do models that state 0–10% confidence actually have accuracy near 0–10%)? (2) What does the shape of miscalibration look like — is it uniform overconfidence at all confidence levels, or a specific pattern (e.g., models are calibrated at low confidence but overconfident at high confidence, or vice versa)? (3) How does the calibration curve differ between reasoning models (O1, DeepSeek-R1, O3-Mini) and non-reasoning models (GPT-4O, Claude 3.5 Sonnet) — do reasoning models spread their confidence across a wider range, or do they simply output lower average confidence? This study would convert the opaque aggregate calibration error into actionable diagnostic information about how models fail at uncertainty estimation on frontier-difficulty questions, enabling targeted interventions (e.g., if models are calibrated at low confidence but overconfident at high confidence, post-hoc recalibration of high-confidence predictions might help; if models output uniform high confidence on all questions, the problem is more fundamental).

Cross-domain generalization of the difficulty-gated benchmark paradigm. HLE demonstrates that LLM difficulty filtering plus expert review produces a durable benchmark for closed-ended academic questions. A natural extension would test whether the same paradigm works for other evaluation modalities: (1) code generation — a benchmark of closed-ended programming problems (specified input/output behavior, unit-test verifiable) filtered to require failure by frontier code models (GPT-4O, Claude 3.5 Sonnet, O1 on coding tasks) before expert review, measuring whether the resulting benchmark resists saturation longer than existing coding benchmarks (HumanEval, MBPP, SWE-bench); (2) adversarial robustness — a benchmark of jailbreaking or prompt injection attacks filtered to require success against current safety-trained models, inverting the difficulty gate (questions must succeed rather than fail against current models), testing whether the gated-construction paradigm works for adversarial evaluation where the goal is to find model failures rather than model successes; (3) multi-modal reasoning — extending HLE's 14% image-containing questions to a fully multi-modal benchmark where every question requires integrating text, images, and potentially audio or video, with difficulty gated against frontier multi-modal models. Each extension would test whether the difficulty-gating methodology generalizes beyond text-dominant academic Q&A or whether it is specific to the knowledge-recall-plus-reasoning structure of HLE-style questions.

Connecting calibration on HLE to downstream deployment safety. The paper argues that poor calibration on HLE indicates models will confidently hallucinate when deployed beyond their training distribution, but this connection is asserted rather than demonstrated. A validation study would: (1) identify a set of real-world deployment tasks where model overconfidence is known to cause harm (e.g., medical diagnosis from incomplete symptoms, legal advice on novel jurisdictional questions, scientific literature review where papers may contain errors); (2) measure RMS calibration error on HLE and on the deployment task for the same set of models; (3) test whether HLE calibration error predicts deployment-task overconfidence better than calibration on in-distribution benchmarks (MMLU, GPQA). If HLE calibration is a stronger predictor of real-world overconfidence than standard-benchmark calibration, it validates HLE's claimed safety relevance. If not, HLE's calibration measurement, while interesting, would not translate to the deployment concerns the paper invokes. This study requires careful experimental design — deployment tasks must genuinely push models beyond their training distribution to create the conditions where HLE-like overconfidence would emerge.


Practical Applications and Downstream Use Cases

Frontier model release evaluation and capability reporting. When AI developers release new frontier models (GPT-5, Claude 4, Gemini 3, etc.), they typically report performance on standard benchmarks as evidence of capability. HLE provides a benchmark where even the best current models achieve 2.7–13.4% accuracy, creating headroom for meaningful measurement through multiple future model generations. A developer releasing a model that achieves, say, 35% on HLE would be demonstrating a qualitative capability jump from the current frontier — not just incremental improvement on saturated metrics. The calibration dimension adds safety-relevant information: a model achieving 35% accuracy with 40% calibration error would be genuinely improving both capability and self-awareness, while a model achieving 35% accuracy with 85% calibration error would be more capable but equally dangerous in its overconfidence. HLE's public release and standardized evaluation protocol (Section C.1.1) make it immediately usable for this purpose without requiring custom evaluation infrastructure.

Safety evaluation for high-stakes deployment decisions. Organizations considering deploying LLMs in domains where confident errors could cause harm — medical diagnosis support, legal document analysis, scientific research assistance, financial modeling — need to know not just whether models can answer routine questions correctly, but whether they recognize when questions exceed their capabilities. HLE's calibration measurement (Table 1: all models show RMS calibration errors above 70%) provides a direct test of this property: a model's calibration error on HLE indicates its tendency to confidently produce wrong answers when faced with genuinely difficult questions. Before deploying an LLM in a high-stakes domain, an organization could evaluate the model on HLE (or a domain-specific subset of HLE questions, e.g., the biology/medicine or chemistry categories) and measure both accuracy and calibration. A model with single-digit accuracy and 85% calibration error would be flagged as unsafe for autonomous use in that domain — it would confidently give wrong answers to any genuinely difficult question that arises. This is not hypothetical: the paper shows that all current frontier models exhibit this failure mode, making HLE a practical filter for deployment risk assessment.

Monitoring AI progress for governance and policy. Policymakers and safety researchers require clear, interpretable metrics for tracking AI capability trends. HLE's design — anchored to the capabilities of late-2024 models via the LLM difficulty filter (Section 3.2), with public release enabling independent verification, and with HLE-Rolling providing a pathway for maintaining difficulty as models improve (Section B.3) — makes it well-suited for longitudinal capability tracking. An annual "State of AI" report could include HLE accuracy trends alongside economic and safety indicators, tracking both the best model's accuracy (are we approaching the expert disagreement rate of 15.4% where ground-truth uncertainty limits interpretation?) and the gap between best and average models (is capability concentrating in a few frontier systems or broadly diffusing?). The paper's explicit scope boundary — that HLE tests structured academic problems, not autonomous research or general intelligence (Section 5) — provides interpretive guardrails that prevent overclaiming what HLE trends mean, a feature that many benchmarks lack.

Benchmark design toolkit for specialized domains. Organizations that need to evaluate models on domain-specific knowledge — pharmaceutical companies testing models on drug interaction knowledge, law firms testing models on jurisdictional case law, engineering firms testing models on safety-critical calculations — can adopt HLE's difficulty-gated construction pipeline without building from scratch. The recipe is: (1) recruit domain experts to write closed-ended questions with unambiguous answers; (2) filter questions through the best available general-purpose and domain-specific models, rejecting any that current models answer correctly (with the multiple-choice vs. exact-match criterion distinction from Section 3.2); (3) run two rounds of expert peer review with the rubrics from Section C.7; (4) measure both accuracy and calibration on the resulting benchmark; (5) estimate expert disagreement rate through independent audit to establish the measurement ceiling. The paper provides all necessary methodological detail (reviewer instructions, scoring rubrics, evaluation prompts, calibration measurement protocol) to replicate this pipeline for arbitrary domains. The key insight transferable from HLE is that the LLM difficulty filter should use current frontier models as the gate — a domain-specific benchmark filtered against GPT-4O in 2024 will be easy for GPT-5 in 2026, so the filtering must be refreshed as models improve, just as HLE-Rolling plans to do.