Open-Source LLM and Leaderboard 2026 Collaboration
benchmarks claude deepseek llama open-source qwen reasoning
| Source: Mastodon | Original article
Open-source LLMs narrow gap with proprietary models, with top open-source option being 2x cheaper.
The Open-Source LLM Leaderboard 2026 has been updated, comparing the performance of open-source and proprietary large language models. According to the leaderboard, Kimi K3 is currently the best open-weight model with a score of 57.1, while Claude Opus 5 leads the proprietary models with a score of 60.7. Notably, Kimi K3 is approximately 2x cheaper per 1M output tokens than Claude Opus 5, despite a 3.6-point gap in their scores.
This leaderboard matters because it provides a comprehensive comparison of open-source and proprietary LLMs, helping developers and users make informed decisions about which models to use. The rankings are based on a range of tasks, including reasoning, coding, math, and multilingual tasks, giving a well-rounded view of each model's capabilities.
As the LLM landscape continues to evolve, it will be interesting to watch how these rankings change over time. With multiple sources tracking the performance of open-source LLMs, including BenchLM.ai and WhatLLM.org, users can expect to see ongoing updates and new models emerging. The Open-Source LLM Leaderboard 2026 can be found at olud.ai/leaderboard.html, providing a valuable resource for those looking to explore the latest developments in open-source LLMs.
Sources
Back to AIPULSEN