VA-Judger Uses Human Preference Feedback for Joint Video‑Audio Generation
benchmarks reinforcement-learning
| Source: HF Papers | Original article
VA-Judger introduces a human‑feedback reward model for reinforcement‑learning post‑training of joint video‑audio generators, overcoming the limits of metric‑based reward signals.
Researchers have unveiled **VA‑Judger**, the first reward model designed specifically for joint video‑audio generation. The system learns from human‑preference feedback, using a chain‑of‑thought “omni‑reward” architecture that evaluates a text prompt alongside two complete clips to judge overall quality, synchronization and fidelity. To train the model, the team compiled the VAPref‑10K resource, containing roughly 9 000 prompts and 10.3 000 fine‑grained paired comparisons, providing a richer signal than the separate audio, visual and sync metrics traditionally used.
VA‑Judger tackles a long‑standing gap in reinforcement‑learning pipelines for multimodal generation. Existing approaches typically combine isolated quality metrics, which can miss holistic judgments that humans make when assessing video‑audio pairs. By first learning from pairs with clear quality gaps, then distilling preference explanations for near‑identical samples through rejection sampling verified against human annotations, the model delivers dimension‑wise reinforcement signals that better reflect user expectations.
The authors also introduced **VA‑Judger‑Bench**, a benchmark that pits in‑domain and out‑of‑domain models against each other to test how well reward models align with human preferences. Early results suggest that VA‑Judger’s explanations improve the reliability of preference discrimination, potentially leading to more coherent and engaging generated content.
The development could accelerate the refinement of generative systems for entertainment, education and virtual‑reality applications, where synchronized audio‑visual output is critical. Observers will watch for adoption of the VA‑Judger framework in open‑source projects and commercial pipelines, as well as follow‑up studies that expand the dataset or apply the chain‑of‑thought reward approach to other multimodal tasks.
Sources
Back to AIPULSEN