Comparison of Top Media Models: Open-Source and Proprietary Systems
open-source qwen speech
| Source: Mastodon | Original article
Open-source and proprietary media models compete in a leaderboard. Top models ranked by human preference include downloadable TTS options.
The Media Model Leaderboard has sparked interest in the AI community by pitting open-source models against their proprietary counterparts. As of the latest update, Kokoro 82M v1.0 leads the pack among downloadable TTS models, yet it trails behind Qwen-Audio-3.0-TTS-Plus by 173 ELO in blind human preference tests. This ranking system, hosted on olud.ai, allows for a direct comparison of open-source and proprietary models across various media formats, including image, video, and speech models.
The significance of this leaderboard lies in its use of human-preference rankings, where real people compare outputs from different models without knowing which model produced which output. This approach provides a more nuanced understanding of model performance, moving beyond traditional benchmarking methods. The fact that open-source models are being compared directly to proprietary ones highlights the growing competitiveness of the open-source community in the AI landscape.
As the AI landscape continues to evolve, it will be interesting to watch how open-source models fare against their proprietary rivals. With resources like the Media Model Leaderboard and the Open LLM Leaderboard, the community has access to detailed comparisons and rankings. The ongoing updates to these leaderboards will likely influence the development and adoption of AI models, potentially shifting the balance between open-source and proprietary solutions.
Sources
Back to AIPULSEN