On June 30, Longwei Cong presented the work “Confidence Estimation in Automatic Short Answer Grading with LLMs” at the 27th International Conference on Artificial Intelligence in Education (AIED 2026) in Seoul.
The work investigates how confidence estimates can make automatic short-answer grading with large language models more reliable in practice. In addition to commonly used model-based confidence signals, it also considers uncertainty originating from the data itself. In particular, aleatoric uncertainty is estimated from the semantic heterogeneity of student responses and combined with several model-side confidence signals.
The study evaluates these uncertainty estimates with a focus on practical use: whether they can support selective grading and help identify responses for which automated grading is less reliable and human review would be preferable.
The paper is available via the following link: Confidence Estimation in Automatic Short Answer Grading with LLMs | Springer Nature Link
