Open-Source LLM and Leaderboard 2026 Collaboration
benchmarks deepseek llama open-source qwen
| Source: Mastodon | Original article
Open-source LLMs are being benchmarked for performance. DeepSeek-V2.5 scores 76.3% on MATH-500.
The Open-Source LLM Leaderboard 2026 has been updated, providing a comprehensive comparison of open-source and open-weight Large Language Models (LLMs). According to the leaderboard, DeepSeek-V2.5, released in December 2024, has achieved a score of 76.3% on the MATH-500 benchmark. This independently measured score offers a reliable benchmark for evaluating the model's performance.
This update matters because it provides developers and users with a transparent and trustworthy comparison of open-source LLMs. The leaderboard includes models such as Llama, DeepSeek, Qwen, and Kimi, allowing users to evaluate their performance, pricing, speed, and context windows. As the field of AI continues to evolve, such benchmarks are essential for identifying the most effective and efficient models.
As the landscape of open-source LLMs continues to shift, it will be interesting to watch how these models perform in various tasks, such as coding, math, and chat benchmarks. The leaderboard will likely be updated regularly, reflecting new releases and improvements to existing models. Users can track these developments and compare the latest models on the leaderboard, available at olud.ai/leaderboard.html.
Sources
Back to AIPULSEN