Linear Digressions

← Linear Digressions17 Aug · 23 min

Better Know a Benchmark: Humanity's Last Exam

Better Know a Benchmark: Humanity's Last Exam17 Aug23 min

Humanity's Last Exam was designed with a bold premise: questions that human experts can answer, but AI models can't. Originally dubbed "Humanity's Last Stand," this benchmark is a massive academic collaboration — hundreds of contributors, thousands of fiendishly hard questions spanning a wild range of domains. In this Better Know a Benchmark installment, we unpack what HLE is actually testing, how it was built, and what it means when a model finally starts cracking it.