Open 180B Model Leads 10 Launches Official Hugging Face Leaderboards with Zero‑Token Judge
huggingface open-source
| Source: Mastodon | Original article
The open‑source Darwin‑180B‑RSI model tops ten official Hugging Face leaderboards—the most for any organization—with its Zero‑Token Judge evaluating actions in just 0.06 seconds.
Darwin‑180B‑RSI, an open‑weight language model released by Korean startup VIDRAFT, now sits at the top of ten official Hugging Face leaderboards – the most first‑place slots held by any organization. The model’s dominance spans mathematical, scientific, general‑knowledge and multimodal reasoning tracks, where it has posted perfect scores on the AIME 2026 and HMMT 2026 contests and high marks on GPQA Diamond (94.44 %), MMLU‑Pro (88.12 %) and MMMU‑Pro (79.48 %).
The surge is powered by a “zero‑token judge” (ZTC) that evaluates model actions in just 0.06 seconds, enabling rapid, fine‑grained benchmarking across the Hugging Face Open LLM leaderboard. VIDRAFT also unveiled POCKET‑Darwin‑180B, a compressed build of the same 180‑billion‑parameter mixture that can run without a GPU, underscoring a shift from the rack‑scale hardware traditionally required for frontier models.
Why it matters is twofold. First, the results prove that open‑source, community‑driven models can match or exceed the performance of proprietary offerings on a breadth of tasks, challenging the monopoly of commercial giants in high‑end AI. Second, the ability to run a 180 B model on modest hardware lowers the barrier to entry for research labs and developers, accelerating democratization of advanced language‑model capabilities.
Looking ahead, the community will watch how POCKET‑Darwin‑180B is adopted in real‑world applications and whether its GPU‑free footprint spurs further compression breakthroughs. Updates to the Hugging Face leaderboards will reveal if the model can sustain its lead as new benchmarks emerge. Additionally, the ZTC evaluation framework may become a standard tool for rapid model assessment, influencing how future open‑source LLMs are compared and refined.
Sources
Back to AIPULSEN