We are releasing HLE-Diamond, a refined subset of Humanity’s Last Exam (HLE) question collection, following a year-long process of cleaning and refinement with input from research communities.
HLE-Diamond consists of 1,000 questions.
Main results. We compare the performance of current models on HLE-Diamond without tools.
All models are evaluated with reasoning high.
GPT-6Astra
59.9%ClaudeOpus 5.5
54.6%GPT-6.1Sol
53.2%ClaudeFable 5.1
50.7%Gemini3.8 Flash
33.3%ClaudeSonnet 5.5
33.1%GPT-6Sol
32.8%MuseSpark 1.3
24.6%Grok4.7
22.8%
Token Usage
GPT-6Astra
3kGPT-6.1Sol
4kGPT-6Sol
5kClaudeSonnet 5.5
9kClaudeOpus 5.5
10kMuseSpark 1.3
10kClaudeFable 5.1
16kGemini3.8 Flash
21kGrok4.7
43k
Dataset. HLE-Diamond consists of 500 reasoning and 500 knowledge questions, measuring reasoning and expert knowledge, respectively.
GPT-6Astra
74.2%45.6%ClaudeOpus 5.5
62.4%46.8%GPT-6.1Sol
67.4%39.0%ClaudeFable 5.1
60.8%40.6%Gemini3.8 Flash
36.6%30.0%ClaudeSonnet 5.5
42.4%23.8%GPT-6Sol
42.2%23.4%MuseSpark 1.3
30.0%19.2%Grok4.7
31.2%14.4%
All models are evaluated with reasoning high.
Evaluation with tools. HLE-Diamond questions are designed to be answerable in a closed-book setting, testing both reasoning and expert knowledge. Since HLE is also used to evaluate the capabilities of agentic systems, we outline our recommended settings for evaluating HLE-Diamond with tools here.
GPT-6Astra
59.9%82.9%ClaudeOpus 5.5
54.6%73.9%ClaudeFable 5.1
50.7%72.3%Gemini3.8 Flash
33.3%60.6%GPT-6Sol
32.8%65.0%MuseSpark 1.3
24.6%55.4%
All models are evaluated with reasoning high.
Acknowledgement. We extend our deepest gratitude to all participating question contributors and experts involved in creating and refining the dataset, and to the researchers whose inputs across HLE-Rolling shaped HLE-Diamond.
For any inquiries, please contact agibenchmark@safe.ai.
Citation
@article{phan2025lastexam,
title = {A benchmark of expert-level academic questions to assess {AI} capabilities},
author = {{Center for AI Safety} and {Scale AI} and {HLE Contributors Consortium}},
journal = {Nature},
volume = {649},
pages = {1139--1146},
year = {2026},
doi = {10.1038/s41586-025-09962-4},
eprint = {2501.14249},
archivePrefix = {arXiv},
primaryClass = {cs.LG},
url = {https://arxiv.org/abs/2501.14249}
}