[New Pub] Estimating LLM Grading Ability and Response Difficulty in Automatic Short Answer Grading via Item Response Theory

New Pub
On July 4, 2026, Longwei Cong virtually presented the paper “Estimating LLM Grading Ability and Response Difficulty in Automatic Short Answer Grading via Item Response Theory” at the ACL SIGEDU 21st Workshop on Innovative Use of NLP for Building Educational Applications (BEA 2026). The work investigates automatic short answer grading with large language models from a psychometric perspective. Instead of relying only on aggregate evaluation metrics such as accuracy or macro-F1, the study applies Item Response Theory to jointly estimate the grading ability of individual LLMs and the difficulty of student responses. The results show that LLM graders differ not only in their overall performance, but also in their robustness when evaluating more difficult responses. In addition, highly difficult responses were found to induce systematic grading errors, including a tendency…
Read More

[New Pub] Confidence Estimation in Automatic Short Answer Grading with LLMs

New Pub
  On June 30, Longwei Cong presented the work “Confidence Estimation in Automatic Short Answer Grading with LLMs” at the 27th International Conference on Artificial Intelligence in Education (AIED 2026) in Seoul. The work investigates how confidence estimates can make automatic short-answer grading with large language models more reliable in practice. In addition to commonly used model-based confidence signals, it also considers uncertainty originating from the data itself. In particular, aleatoric uncertainty is estimated from the semantic heterogeneity of student responses and combined with several model-side confidence signals. The study evaluates these uncertainty estimates with a focus on practical use: whether they can support selective grading and help identify responses for which automated grading is less reliable and human review would be preferable. The paper is available via the following…
Read More
New Pub: The Dependency Dilemma of Machine Learning Aids

New Pub: The Dependency Dilemma of Machine Learning Aids

Artificial Intelligence, Journal, New Pub
AI and machine-learning decision aids can be a great support for employees and help them improve their work performance. An overreliance on algorithmic recommendations, on the other hand, may result in a long-term reduction of the employees’ ability to develop and maintain their own decision-making skills. This could be especially problematic when such AI and machine-learning decision aids are not available. Instead of overly relying on such AI-based recommendations, organizations should train their employees to continue developing their critical thinking and decision-making skills. Hendrik Drachsler and his research partners address this problematic dependency on AI tools and the conflict of interests in the workplace in their newly published article “The Dependency Dilemma: How Machine Learning Decision Aids can Undermine Skill Growth”. Using a controlled experiment, the authors found that participants…
Read More
New Pub: Recommendations for Higher Education in the Age of Generative AI

New Pub: Recommendations for Higher Education in the Age of Generative AI

Artificial Intelligence, Higher Education, Publication, Report
Generative AI cannot be treated as just another digital tool that has come along. As it becomes more and more embedded in higher education, universities face the challenge of responsibly navigating the many challenges and opportunities that generative AI brings with it. One central guiding principal for institutions and stakeholders engaged with generative AI is intellectual sovereignty, which the German Science and Humanities Council (Wissenschaftsrat, WR) highlights in its newly published position paper, which was presented in a digital press release on 06.07.2026. Intellectual sovereignty refers to the ability to think independently, exercise critical judgment and maintain autonomy in the creation and evaluation of knowledge. Rather than relying uncritically on AI-generated outputs, this concept encourages students, educators and institutions to actively question, assess and contextualize information. By placing intellectual sovereignty…
Read More

[New Pub] Report on the BEA 2026 Shared Task on Rubric-based Short Answer Scoring for German

New Pub
At BEA 2026, we organized the first shared task on rubric-based short answer scoring for German, introducing a new benchmark dataset spanning multiple STEM domains and designed to evaluate both in-domain performance and generalization to previously unseen questions. The goal was to better understand how modern NLP systems can interpret and apply textual scoring rubrics—an ability that closely mirrors how human assessors evaluate student responses. The shared task attracted multiple teams, who explored a wide range of approaches, including fine-tuned large language models, retrieval-augmented prompting, hybrid symbolic–neural systems, and ensemble methods. Across all four evaluation tracks, systems that explicitly incorporated rubric semantics consistently achieved the strongest performance, while scoring previously unseen questions remained the biggest challenge. We hope this benchmark provides a foundation for future research on more robust, interpretable,…
Read More

[New Pub] Rubrics as Semantic Subspaces: A New Perspective on AI Assessment

New Pub
How can AI not only score student answers accurately, but also explain why it assigned a particular score? In our latest paper, we introduce AGRAA (Aspect-Grounded Rubric–Answer Alignment), a new framework for rubric-based assessment that represents rubric criteria as semantic subspaces. Instead of treating scoring as a standard classification problem, AGRAA measures how strongly a student's response aligns with the latent semantic aspects defined by each rubric level. We evaluate the approach on both short-answer and essay scoring benchmarks, where it consistently achieves performance highly competitive with strong transformer-based baselines. More importantly, the model naturally produces interpretable explanations by showing which rubric-defined aspects contributed to each scoring decision. We believe this work is a step toward automated assessment systems that are not only accurate, but also transparent, trustworthy, and aligned…
Read More
Congratulations, Dr. Karademir! Turning Data into Pedagogical Action

Congratulations, Dr. Karademir! Turning Data into Pedagogical Action

Alumni, Learning Analytics, PhD defense
On 1 July 2026, Onur Karademir successfully defended his doctoral dissertation at the Goethe-University Frankfurt — and we could not be prouder. Onur joined the EduTec team in February 2021, but his story started even earlier: as a student frustrated by 90-minute maths lectures, he channelled that experience into founding StudyCore, an EdTech startup delivering digital teaching tools to universities and schools. That hands-on instinct — building things that actually reach learners — defined his entire PhD trajectory. His dissertation, Turning Data into Pedagogical Action: Designing the Learning Analytics Cockpit, a Teacher Dashboard for Feedback, tackles one of the most pressing challenges in educational technology: making learning analytics useful in real classrooms. Over three interconnected studies — a proof of concept, a co-design process with secondary school teachers and a…
Read More