Dagstuhl Seminar 25272
Challenges of Human Oversight: Achieving Human Control of AI-Based Systems
( Jun 29 – Jul 04, 2025 )
Permalink
Organizers
- Raimund Dachselt (TU Dresden, DE)
- Markus Langer (Universität Freiburg, DE)
- Q. Vera Liao (Microsoft - Montréal, CA)
- Tim Miller (University of Queensland - Brisbane, AU)
- Nava Tintarev (Maastricht University, NL)
Contact
- Marsha Kleinbauer (for scientific matters)
- Susanne Bach-Bernhard (for administrative matters)
Press/News
Schedule
What is effective human oversight of AI systems? The Dagstuhl Seminar 25272 “Challenges of Human Oversight: Achieving Human Control of AI-Based Systems” brought together interdisciplinary experts from artificial intelligence, human-computer interaction, human factors and psychology, philosophy and ethics, as well as law to explore conceptual, technical, legal, and practical dimensions of human oversight of AI. Across the seminar, participants provided perspective talks from the different disciplines and engaged in working groups and use-case specific discussions in order to establish a science of human oversight of AI systems. The main outcome of this seminar is a general framework that outlines the architecture, processes, and sociotechnical design dimensions of human oversight of AI systems. In the following, we present some of the key insights of this seminar in more detail.
Conceptual Foundations of Human Oversight. Human oversight is defined as a human activity to monitor and intervene in AI-supported tasks (typically at runtime) with the aim of sufficiently mitigating risks. Mitigating risks means detecting errors, system malfunctions, or inadequate outputs. Effectiveness depends on epistemic access, causal power, self-control, and fitting intentions of the human oversight personnel. In order to optimize the human oversight effectiveness, it requires designing the sociotechnical dimensions of human oversight: the human factors, technical design, and contextual considerations. Human oversight can operate at multiple layers and across distributed roles within human oversight teams and is inherently interdependent with other risk mitigation measures.
Human Factors, Technical Design, and Contextual Considerations. Human factors cover situation awareness, decision making, cognitive biases, workload, motivation, training, and collaboration between oversight personnel. Technical design must support human detection of system errors and failures, for example via visualization, technical support tools, adaptive automation, handover design, and personalization. Contextual considerations to support human oversight include time and resource requirements for effective human oversight as well as the clarity of human oversight roles and duties.
Legal, and Normative Considerations. Human oversight effectiveness requires normative judgment beyond legal mandates. Human oversight objects include individual AI systems, highly-autonomous agentic AI in high-risk domains, as well as AI systems operated by users such as patients using mental health chatbots who themselves are not considered human oversight personnel. The seminar highlighted the importance to consider the relation between human oversight, other risk management measures, and technical standards.
Evaluating Human Oversight Effectiveness. Evaluation of human oversight implementation is crucial given that human oversight effectiveness can only be achieved iteratively. Metrics include effectiveness of monitoring and interventions (e.g., detecting erroneous AI outputs, overriding these outputs), alignment with human oversight protocols, and long-term performance outcomes. Mixed-methods approaches (quantitative and qualitative) and comparative studies of different human oversight design options (e.g., varying support interfaces) were discussed as possible options to evaluate human oversight effectiveness.
Human Oversight Effectiveness as an Iterative and Multi-Layered Challenge. Continuous updating of human oversight design is essential, integrating empirical feedback and ensuring institutional support for high-quality and sustainable human oversight of AI. Furthermore, we saw that effective oversight required identifying information and workflows across regulatory, technical, and interface layers.
Conclusion. The seminar demonstrated that human oversight of AI is a multifaceted, interdisciplinary challenge, involving conceptual clarity, human factors, technical design, contextual considerations, evaluation frameworks, as well as legal and ethical considerations. The outputs of this seminar provide a foundation for theoretical modeling, empirical research, practical design guidance, and normative reflection, establishing a roadmap for advancing effective human oversight in AI systems contributing to the safe implementation of AI in highrisk contexts. Next steps include joint publications (e.g., a framework for human oversight of AI), developing technical support tools for effective human oversight, and community building through workshops at key human-computer interaction and AI conferences.
Raimund Dachselt, Markus Langer, Q. Vera Liao, Tim Miller, and Nava Tintarev
This Dagstuhl Seminar will investigate challenges and approaches to designing Artificial Intelligence (AI) systems that ensure meaningful human oversight and control of their operation.
The past decade has seen a substantial leap forward in AI research and technology, spanning from novel neural network-based deep learning approaches to applications of generative AI based on large language models. Rapidly, these technologies are being applied to numerous fields. Especially the deployment of AI-based systems in high-risk contexts entails threats to safety and fundamental human rights. To alleviate such risks, emerging ethical guidelines and legislation (such as the European AI Act) around the globe are calling for Human Oversight of AI-based systems in high-risk contexts.
However, the conditions for effective human oversight are ill-understood and so is their technological basis. Among the urgent questions that developers and deployers of AI-based systems have to deal with once such AI regulations are in force include: How to design interfaces and the communication between humans and systems to enable overseers to effectively control AI-based systems? Can explainability and visualization approaches promote an understanding of system capacities and limitations? How do we ensure that people override system outputs in situations where this reduces risks and does not introduce new risks?
This urgently calls for an interdisciplinary discussion among researchers in artificial intelligence, computer system design and verification, human-computer interaction, psychology, ethics, and law. This seminar involves AI systems researchers who need to work on ensuring system interpretability and advancing explainability approaches in a way that serves the needs of human overseers. This involves formal methods researchers who will help to ensure the accuracy and reliability of explanations stemming from approaches in Explainable Artificial Intelligence (XAI) research. This involves language processing and visualization researchers to effectively communicate information to human overseers enabling them, for instance, to effectively grasp current system states and risky situations. And this critically involves human-computer interaction researchers , for designing human-system decision workflows and oversight support tools that are tailored to the tasks of human overseers.
Apart from computer science experts, the seminar needs the perspective of psychology to understand people’s needs in ensuring effective human oversight, and in developing evaluation methods to empirically test whether human oversight is truly effective. We also need the perspectives of law and normative ethics to interpret the foundational assumptions behind emerging regulation and its connections to other legal frameworks. And eventually, we need all perspectives to respond to the intertwined questions: How to assess whether we have achieved effective human oversight? What are the requirements for “overseeability-by-design”? What if we find that there are limits to human oversight? How can and perhaps even should we get involved in shaping policy-making to take into account the limits and conditions of effective human oversight?
Among the major topics to be discussed during the seminar are: the scope and requirements of human oversight, AI systems built for human oversight, effective and intuitive user interfaces for human oversight, the evaluation of human oversight approaches, decision-making and associated challenges, and how to interface to law, norms, and society.
Raimund Dachselt, Markus Langer, Q. Vera Liao, Timothy Miller, and Nava Tintarev
- Kevin Baum (DFKI - Saarbrücken, DE) [dblp]
- Raimund Dachselt (TU Dresden, DE) [dblp]
- Virginia Dignum (University of Umeå, SE) [dblp]
- Anna Maria Feit (Universität des Saarlandes - Saarbrücken, DE) [dblp]
- Ujwal Gadiraju (TU Delft, NL) [dblp]
- Susanne Gaube (University College London, GB) [dblp]
- Holger Hermanns (Universität des Saarlandes - Saarbrücken, DE) [dblp]
- Oana Inel (Universität Zürich, CH) [dblp]
- Harmanpreet Kaur (University of Minnesota - Minneapolis, US) [dblp]
- Mark T. Keane (University College Dublin, IE) [dblp]
- Richard Landers (University of Minnesota - Minneapolis, US) [dblp]
- Markus Langer (Universität Freiburg, DE) [dblp]
- Anne Lauber-Rönsberg (TU Dresden, DE) [dblp]
- Johann Laux (University of Oxford, GB)
- Q. Vera Liao (Microsoft - Montréal, CA) [dblp]
- Brian Lim (National University of Singapore, SG) [dblp]
- Philip Meinel (TU Dresden, DE) [dblp]
- Tim Miller (University of Queensland - Brisbane, AU) [dblp]
- Linda Onnasch (TU Berlin, DE) [dblp]
- Carola Plesch (BSI - Bonn, DE)
- Tim Schrills (Universität Lübeck, DE) [dblp]
- Marija Slavkovik (University of Bergen, NO) [dblp]
- Liz Sonenberg (University of Melbourne, AU) [dblp]
- Sarah Sterz (Universität des Saarlandes - Saarbrücken, DE) [dblp]
- Chenhao Tan (University of Chicago, US) [dblp]
- Nava Tintarev (Maastricht University, NL) [dblp]
- Silja Voeneky (Universität Freiburg, DE) [dblp]
- Ziang Xiao (Johns Hopkins University - Baltimore, US) [dblp]
- Hanwei Zhang (Universität des Saarlandes - Saarbrücken, DE) [dblp]
Classification
- Artificial Intelligence
- Computers and Society
- Human-Computer Interaction
Keywords
- artifical intelligence
- explainable AI
- norms and regulations
- human oversight
- safety

Creative Commons BY 4.0
