[New Pub] Estimating LLM Grading Ability and Response Difficulty in Automatic Short Answer Grading via Item Response Theory

New Pub
On July 4, 2026, Longwei Cong virtually presented the paper “Estimating LLM Grading Ability and Response Difficulty in Automatic Short Answer Grading via Item Response Theory” at the ACL SIGEDU 21st Workshop on Innovative Use of NLP for Building Educational Applications (BEA 2026). The work investigates automatic short answer grading with large language models from a psychometric perspective. Instead of relying only on aggregate evaluation metrics such as accuracy or macro-F1, the study applies Item Response Theory to jointly estimate the grading ability of individual LLMs and the difficulty of student responses. The results show that LLM graders differ not only in their overall performance, but also in their robustness when evaluating more difficult responses. In addition, highly difficult responses were found to induce systematic grading errors, including a tendency…
Read More

[New Pub] Confidence Estimation in Automatic Short Answer Grading with LLMs

New Pub
  On June 30, Longwei Cong presented the work “Confidence Estimation in Automatic Short Answer Grading with LLMs” at the 27th International Conference on Artificial Intelligence in Education (AIED 2026) in Seoul. The work investigates how confidence estimates can make automatic short-answer grading with large language models more reliable in practice. In addition to commonly used model-based confidence signals, it also considers uncertainty originating from the data itself. In particular, aleatoric uncertainty is estimated from the semantic heterogeneity of student responses and combined with several model-side confidence signals. The study evaluates these uncertainty estimates with a focus on practical use: whether they can support selective grading and help identify responses for which automated grading is less reliable and human review would be preferable. The paper is available via the following…
Read More