CorrGRPO unveils correlation‑normalized GRPO for multi‑reward learning
reasoning
| Source: HF Papers | Original article
Researchers introduce CorrGRPO, a correlation‑normalized extension of Group Relative Policy Optimization that improves multi‑reward learning by normalizing total rewards within prompt groups.
A new variant of Group Relative Policy Optimization (GRPO) has been released that tackles a long‑standing bias in multi‑reward reinforcement learning. Dubbed CorrGRPO, the method replaces the raw covariance terms in GRPO’s normalizer with Pearson correlation coefficients, preventing large‑scale reward components from overwhelming smaller ones. The change leaves the centered total reward untouched while re‑balancing the influence of each reward according to its correlation with the others.
The improvement matters because modern reasoning language models are often trained on several objectives at once—code correctness, formatting, tool‑call precision, safety, and more. Standard GRPO aggregates these rewards and normalizes by the within‑group standard deviation, a process that can be skewed when one reward dominates the variance. CorrGRPO’s correlation‑based scaling lets advantage estimates reflect the true interplay among objectives, leading to more stable and effective learning.
Early results on the LeetCodeDataset using the Qwen2.5‑Coder family show the approach delivering 2.09‑4.21 Pass@1 points over vanilla GRPO across model sizes from 0.5 B to 7 B parameters. The gains suggest that CorrGRPO can unlock higher performance for multi‑reward tasks without additional data or model changes, a promising development for both academic research and industrial deployments of code‑generation and other reasoning systems.
The community will now watch for broader adoption of CorrGRPO in other multi‑objective domains, such as tool‑use agents and safety‑constrained language models. Follow‑up studies are likely to explore how the correlation‑normalized scheme interacts with curriculum learning strategies and agent‑prior guided policies that have recently been proposed. If the early gains hold across diverse benchmarks, CorrGRPO could become the new default for multi‑reward reinforcement learning in large‑scale language models.
Sources
Back to AIPULSEN