Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Won-Chan Lee

Comparing Human and AI Rater Effects Using the Many-Facet Rasch Model

May 28, 2025

Hong Jiao, Dan Song, Won-Chan Lee

Figure 1 for Comparing Human and AI Rater Effects Using the Many-Facet Rasch Model

Figure 2 for Comparing Human and AI Rater Effects Using the Many-Facet Rasch Model

Figure 3 for Comparing Human and AI Rater Effects Using the Many-Facet Rasch Model

Figure 4 for Comparing Human and AI Rater Effects Using the Many-Facet Rasch Model

Abstract:Large language models (LLMs) have been widely explored for automated scoring in low-stakes assessment to facilitate learning and instruction. Empirical evidence related to which LLM produces the most reliable scores and induces least rater effects needs to be collected before the use of LLMs for automated scoring in practice. This study compared ten LLMs (ChatGPT 3.5, ChatGPT 4, ChatGPT 4o, OpenAI o1, Claude 3.5 Sonnet, Gemini 1.5, Gemini 1.5 Pro, Gemini 2.0, as well as DeepSeek V3, and DeepSeek R1) with human expert raters in scoring two types of writing tasks. The accuracy of the holistic and analytic scores from LLMs compared with human raters was evaluated in terms of Quadratic Weighted Kappa. Intra-rater consistency across prompts was compared in terms of Cronbach Alpha. Rater effects of LLMs were evaluated and compared with human raters using the Many-Facet Rasch model. The results in general supported the use of ChatGPT 4o, Gemini 1.5 Pro, and Claude 3.5 Sonnet with high scoring accuracy, better rater reliability, and less rater effects.

Via

Access Paper or Ask Questions

Investigating AI Rater Effects of Large Language Models: GPT, Claude, Gemini, and DeepSeek

May 24, 2025

Hong Jiao, Dan Song, Won-Chan Lee

Figure 1 for Investigating AI Rater Effects of Large Language Models: GPT, Claude, Gemini, and DeepSeek

Figure 2 for Investigating AI Rater Effects of Large Language Models: GPT, Claude, Gemini, and DeepSeek

Figure 3 for Investigating AI Rater Effects of Large Language Models: GPT, Claude, Gemini, and DeepSeek

Figure 4 for Investigating AI Rater Effects of Large Language Models: GPT, Claude, Gemini, and DeepSeek

Via

Access Paper or Ask Questions