Darwin-180B-RSI: VIDRAFT's recursive self‑improvement model Topped 7 the leaderboards without legal training data
huggingface training
| Source: Mastodon | Original article
VIDRAFT has pushed an open‑source model to the top of several Hugging Face benchmarks. Its new 180‑billion‑parameter mixture‑of‑experts vision‑language model, Darwin‑180B‑RSI, claimed first place on seven official leaderboards, including the legal‑focused LEXam and LEXam‑hard suites, despite being trained without any legal‑domain data.
The model builds on the Qwen/Qwen3.8‑Flash‑Next “parent” model, preserving its 512 routed experts, router and vision encoder. VIDRAFT’s Darwin framework then diagnoses the parent’s weak pathways and strengthens them by altering a mere 0.02 % of the parameters. The key innovation is a recursive self‑improvement (RSI) loop that refines answers against verifiable references, feeding the corrected outputs back into the model. This lightweight, targeted tuning allowed the 180 B‑parameter system to outpace much larger models—some with 685 B parameters—on the same tasks.
The achievement matters for several reasons. First, it shows that strategic, data‑efficient refinement can rival brute‑force scaling, a claim that could reshape how researchers allocate compute and data resources. Second, by publishing the weights on Hugging Face, VIDRAFT gives the broader community access to a state‑of‑the‑art legal‑reasoning model without the licensing constraints that often accompany proprietary systems. Finally, the result underscores the growing relevance of RSI techniques, a niche that has recently attracted attention in open‑source decision‑model projects such as Clef.
Going forward, the community will watch how Darwin‑180B‑RSI performs on broader multimodal tasks and whether its RSI pipeline can be generalized to other domains. Scrutiny of the self‑reported leaderboard scores is also likely, given recent discussions about evaluation protocols on Hugging Face. If the model’s gains hold up under independent testing, we may see a wave of similarly lean, self‑improving models challenging the dominance of ever‑larger, data‑hungry architectures.
Sources
Back to AIPULSEN