Outsmarting LLM: Humans vs Their Toughest Challenge Yet
ai-safety benchmarks
| Source: Dev.to | Original article
Humans vs HLE quiz app challenges users against advanced models.
A new quiz app, Humans vs HLE, challenges users to compete against state-of-the-art language models on Humanity's Last Exam, a benchmark consisting of 2,500 questions across various subjects. This app allows individuals to test their knowledge against frontier models, with uncheatable server-side grading and a leaderboard powered by Durable Objects.
The Humanity's Last Exam benchmark was created by the Center for AI Safety and Scale AI, and is considered one of the more challenging tests for language models. Previous results have shown that leading models struggle with this exam, scoring under 30% on average. This highlights the significant gap between current language model capabilities and human expertise.
As language models continue to advance, benchmarks like Humanity's Last Exam will play a crucial role in measuring their capabilities. With the release of the Humans vs HLE quiz app, users can now experience the challenge of competing against these models firsthand. It will be interesting to watch how users perform compared to the models, and whether this app can help identify areas where language models need improvement.
Sources
Back to AIPULSEN