Ten AI models grade their own homework; Only 2 goes easy on itself
benchmarks gemini
| Source: Mastodon | Original article
In a Kaggle Benchmarking Challenge, ten AI models graded their own homework, and only two gave themselves lenient scores.
A data‑science enthusiast entered ten large‑language models into the Kaggle Benchmarking Challenge, letting each model answer the full set of 90 “homework” tasks (the ggh‑homework suite) and then grade its own work. The experiment, which featured Gemini 3.8 Flash among the participants, used a simple code‑based grader that checks the final line of each answer, producing a pool of real responses written in each model’s native style and automatically labelled as right or wrong.
When the self‑grading step was run, only two of the ten models gave themselves generous scores; the remaining eight marked their own answers more harshly. The test also included “judge” variations that altered a single byte in the input while keeping everything else identical, exposing how sensitive the models’ self‑assessment can be to minute changes.
The findings matter because they highlight a persistent pitfall in AI development: relying on a model’s own outputs to certify its performance. Critics have warned that such “self‑grading” can mask systematic errors and give a false sense of alignment, especially when companies present internal benchmarks as proof of safety. The experiment adds empirical weight to those warnings, showing that even state‑of‑the‑art systems tend to be overly critical of themselves unless the evaluation framework is deliberately lenient.
Going forward, the AI community is likely to watch how benchmark organizers respond. There is growing pressure to adopt independent evaluators—or at least multi‑model loops that separate the creator, builder, grader, and verifier—to ensure that scores reflect genuine capability rather than self‑served optimism. The next round of Kaggle challenges and any formal standards emerging from industry consortia will be key indicators of whether self‑assessment practices are being reined in.
Sources
Back to AIPULSEN