On July 4, 2026, Longwei Cong virtually presented the paper “Estimating LLM Grading Ability and Response Difficulty in Automatic Short Answer Grading via Item Response Theory” at the ACL SIGEDU 21st Workshop on Innovative Use of NLP for Building Educational Applications (BEA 2026).

The work investigates automatic short answer grading with large language models from a psychometric perspective. Instead of relying only on aggregate evaluation metrics such as accuracy or macro-F1, the study applies Item Response Theory to jointly estimate the grading ability of individual LLMs and the difficulty of student responses.

The results show that LLM graders differ not only in their overall performance, but also in their robustness when evaluating more difficult responses. In addition, highly difficult responses were found to induce systematic grading errors, including a tendency of models to assign the partially_correct_incomplete category.

The study contributes to a more fine-grained understanding of the strengths and limitations of LLM-based grading systems and demonstrates how psychometric methods can complement conventional model evaluation in educational NLP.

The paper is available via the following link: Estimating LLM Grading Ability and Response Difficulty in Automatic Short Answer Grading via Item Response Theory – ACL Anthology