Mike Thelwall Sees Notable Progress with LLMs Collaboration After Cautious Start
gpt-5
| Source: Mastodon | Original article
Researcher finds LLMs can approximate expert assessments with high correlation. LLMs show promise in evaluating research quality.
Professor Mike Thelwall has made significant progress in his research on using large language models (LLMs) for expert research assessment. A year ago, he was exploring whether LLMs could approximate expert evaluations, and now he reports that ChatGPT-5 mini has achieved correlations of up to 0.905 with expert REF2021 quality assessments in certain disciplines.
This development matters because it suggests that LLMs are becoming increasingly capable of mimicking human reviewers, which could have significant implications for the field of research evaluation. If LLMs can accurately assess research quality, it could potentially streamline the evaluation process and reduce the workload of human reviewers.
As this research continues to evolve, it will be important to watch how LLMs perform across different disciplines and datasets. Professor Thelwall's work is a notable example of the ongoing exploration of LLMs in research evaluation, and his findings will likely be closely followed by academics and researchers in the field. As we reported previously, the potential of LLMs to support or even replace human reviewers is a topic of growing interest, and Professor Thelwall's latest results are a significant step forward in this area.
Sources
Back to AIPULSEN