Study Shows Multi-Agent Code Judge Can Be Grounded with Label‑Free Metrics, Yet It Refuses to Guess
agents reasoning
| Source: ArXiv | Original article
A new arXiv paper examines how language models judge each other's code, revealing that judges often give confident verdicts without evidence, and proposes label‑free metrics to assess grounding.
A new arXiv pre‑print — 2609.30328v1 — examines how reliably multi‑agent systems can judge the correctness of generated code. The authors focus on MARCH, a published framework that splits a verdict into smaller, checkable claims and enlists several language‑model helpers to verify each piece. Running MARCH unmodified across 80 condition‑by‑cell measurements on two established code‑judging benchmarks, they find the system declares both candidate solutions “equally good” in 78 % to 95 % of comparisons. That behaviour translates into a meagre 4.4 % accuracy, far below the 43.7 % accuracy achieved when the same model is asked directly to evaluate the code. The gap persists regardless of problem difficulty or the size of the judging model.
The paper’s key contribution is a label‑free method that detects when a judge lacks sufficient evidence. Rather than issuing a confident but unfounded verdict, the system can now decline to guess. The authors argue that current multi‑agent code reviewers often present reasoning that looks grounded even when the underlying evidence is missing, a risk that grows as enterprises adopt AI‑driven code review pipelines.
Why it matters is twofold. First, AI‑based code checking is already being rolled out in continuous‑integration environments, where an erroneous “pass” can let bugs or security flaws slip into production. Second, the ability to recognize and signal uncertainty aligns with broader concerns about AI agents acting without proper safeguards—a theme echoed in recent work on agent safety platforms and liability for rogue behavior.
What to watch next includes whether tool vendors incorporate the label‑free abstention mechanism into commercial code‑review products, and how larger models respond to the same tests. Researchers are likely to extend the evaluation to real‑world development workflows and to explore complementary techniques for grounding multi‑agent judgments, a step that could make AI‑assisted programming both more trustworthy and more transparent.
Sources
Back to AIPULSEN