Humanity's Last Exam
HLE-Diamond Logo

Introducing HLE-Diamond

Center for AI Safety&Scale AI
Hugging FaceDatasetload_dataset("cais/hle-diamond")

Center for AI Safety and Scale AI

We are releasing HLE-Diamond, a refined subset of Humanity’s Last Exam (HLE) question collection, following a year-long process of cleaning and refinement with input from research communities.

HLE-Diamond consists of 1,000 questions.

Main results. We compare the performance of current models on HLE-Diamond without tools.

All models are evaluated with reasoning high.

HLE-Diamond
  • GPT-6Astra
    59.9%
  • ClaudeOpus 5.5
    54.6%
  • GPT-6.1Sol
    53.2%
  • ClaudeFable 5.1
    50.7%
  • Gemini3.8 Flash
    33.3%
  • ClaudeSonnet 5.5
    33.1%
  • GPT-6Sol
    32.8%
  • MuseSpark 1.3
    24.6%
  • Grok4.7
    22.8%

Token Usage

  • GPT-6Astra
    3k
  • GPT-6.1Sol
    4k
  • GPT-6Sol
    5k
  • ClaudeSonnet 5.5
    9k
  • ClaudeOpus 5.5
    10k
  • MuseSpark 1.3
    10k
  • ClaudeFable 5.1
    16k
  • Gemini3.8 Flash
    21k
  • Grok4.7
    43k

Dataset. HLE-Diamond consists of 500 reasoning and 500 knowledge questions, measuring reasoning and expert knowledge, respectively.

Reasoning and Knowledge Partitions
ReasoningKnowledge
  • GPT-6Astra
    74.2%
    45.6%
  • ClaudeOpus 5.5
    62.4%
    46.8%
  • GPT-6.1Sol
    67.4%
    39.0%
  • ClaudeFable 5.1
    60.8%
    40.6%
  • Gemini3.8 Flash
    36.6%
    30.0%
  • ClaudeSonnet 5.5
    42.4%
    23.8%
  • GPT-6Sol
    42.2%
    23.4%
  • MuseSpark 1.3
    30.0%
    19.2%
  • Grok4.7
    31.2%
    14.4%

All models are evaluated with reasoning high.

Evaluation with tools. HLE-Diamond questions are designed to be answerable in a closed-book setting, testing both reasoning and expert knowledge. Since HLE is also used to evaluate the capabilities of agentic systems, we outline our recommended settings for evaluating HLE-Diamond with tools here.

HLE-Diamond with tools
Without toolsWith tools (web+code)
  • GPT-6Astra
    59.9%
    82.9%
  • ClaudeOpus 5.5
    54.6%
    73.9%
  • ClaudeFable 5.1
    50.7%
    72.3%
  • Gemini3.8 Flash
    33.3%
    60.6%
  • GPT-6Sol
    32.8%
    65.0%
  • MuseSpark 1.3
    24.6%
    55.4%

All models are evaluated with reasoning high.

Acknowledgement. We extend our deepest gratitude to all participating question contributors and experts involved in creating and refining the dataset, and to the researchers whose inputs across HLE-Rolling shaped HLE-Diamond.

For any inquiries, please contact agibenchmark@safe.ai.

Citation

@article{phan2025lastexam,
      title = {A benchmark of expert-level academic questions to assess {AI} capabilities},
      author = {{Center for AI Safety} and {Scale AI} and {HLE Contributors Consortium}},
      journal = {Nature},
      volume = {649},
      pages = {1139--1146},
      year = {2026},
      doi = {10.1038/s41586-025-09962-4},
      eprint = {2501.14249},
      archivePrefix = {arXiv},
      primaryClass = {cs.LG},
      url = {https://arxiv.org/abs/2501.14249}
}