Today, Hendrik Drachsler gave a keynote at the VHB annual Conference in Bamberg, speaking to an audience of over 33 AI project leaders from higher education institutions in Germany. The topic was “From Experiment to Infrastructure: Scaling, Evaluation, and Governance of AI Systems in Higher Education,” reflecting his thinking on various grassroots projects and how to scale them across the whole university.
The gap between performance and learning
Hendrik opened with a provocation: Is AI the new calculator? Both technologies automate cognitive processes, both faced early scepticism, and both promised efficiency gains. But the analogy breaks down quickly. A calculator takes over arithmetic, not problem-solving. AI takes over the formulation, analysis, and argument. That is a qualitative difference, and it demands a different institutional / governance response.
Mixed results on learning
The research picture reinforces this. Studies show that AI tutoring can improve learning outcomes when carefully designed, but the same tools can also have the opposite effect, leading to lower learning gains.
Overreliance is the empirical default
The finding Hendrik returned to throughout the talk: students accept AI-generated feedback largely uncritically, especially under time pressure or uncertainty. This is not an isolated phenomenon; it seems to be the empirical default. But Human oversight, as demanded by the EU AI Act, only works if people have enough domain knowledge to judge whether the AI is right or wrong. That expertise cannot be assumed; it is also not an ethical position; it is a competence that needs to be trained.
What this means for institutions
At Goethe University Frankfurt, Hendrik and his team at studiumdigitale are building their response around four pillars:
- AI literacy for all staff and students,
- sovereign and privacy-compliant infrastructure,
- a multi-level quality framework that explicitly measures overreliance,
- and governance structures that distribute responsibility across faculties.These four pillars only work as a system. A well-resourced infrastructure without evaluation is blind. Evaluation without governance cannot scale.
The most important design principle Hendrik proposed: domain expertise before AI use. Students cannot meaningfully supervise a system they do not yet understand well enough to critique.
The core argument
AI will not follow the calculator’s adoption pattern. It will trigger a co-evolutionary process in which technology, teaching, assessment formats, and competency goals change together. The greatest challenge in rolling out AI at scale is not the technology. It is the question of how we preserve human judgement — at scale.
Scaling AI requires scaling human judgment alongside it. That is the institutional task before us.
