Institut für Business Analytics

Design and Evaluation of Explanations for
User-Centric Human-AI Interactions

Philipp Noah Schröppel · 14 September 2026
Universität Ulm

Agenda

1
Overview
2
Deep Dive
3
Discussion
2

The capital committed to AI

ANNUAL AI INFRASTRUCTURE CAPEX · USD BN
0 400 800 1200 1600 $765 $1,011 $1,220 $1,392 $1,579 $1,636 2026 2027 2028 2029 2030 2031
Compute
Data centres
Power
Source: G. Lee & L. Greenbaum (2026), Tracking Trillions: The Assumptions Shaping the Scale of the AI Build-Out, Goldman Sachs Global Institute. Baseline anchored to NVIDIA data-centre revenue estimates.
3

The capital committed to AI

ANNUAL AI INFRASTRUCTURE CAPEX · USD BN
0 400 800 1200 1600 494 1,127 232 436 $765 $1,011 $1,220 $1,392 $1,579 $1,636 German federal budget 2026 · $635 bn 2026 2027 2028 2029 2030 2031
Compute
Data centres
Power
Source: G. Lee & L. Greenbaum (2026), Tracking Trillions: The Assumptions Shaping the Scale of the AI Build-Out, Goldman Sachs Global Institute. Baseline anchored to NVIDIA data-centre revenue estimates. German federal budget: Bundesministerium der Finanzen, Bundeshaushalt 2026, €545 bn, converted at €1 = $1.165 (10 Sep 2026).
4

Where that electricity is consumed

DATA-CENTRE ELECTRICITY CONSUMPTION · TWh PER YEAR
0 200 400 600 800 1000 415 TWh ≈ 945 TWh 425 280 107 133 2024 2030 base case
United States
China
Europe
Rest of world
Source: IEA (2025), Data centre electricity consumption by region, Base Case, 2020-2030, IEA, Paris. https://www.iea.org/data-and-statistics/charts/data-centre-electricity-consumption-by-region-base-case-2020-2030, Licence: CC BY 4.0. Regional values derived from the published 2024 shares and 2024–2030 growth figures in Energy and AI.
5

Where that electricity is consumed

DATA-CENTRE ELECTRICITY CONSUMPTION · TWh PER YEAR
0 200 400 600 800 1000 415 TWh ≈ 945 TWh 425 280 107 133 Germany · 509 TWh 2024 2030 base case
United States
China
Europe
Rest of world
Source: IEA (2025), Data centre electricity consumption by region, Base Case, 2020-2030, IEA, Paris. https://www.iea.org/data-and-statistics/charts/data-centre-electricity-consumption-by-region-base-case-2020-2030, Licence: CC BY 4.0. Regional values derived from the published 2024 shares and 2024–2030 growth figures in Energy and AI. German reference: Destatis, gross electricity production, 509.2 TWh.
6

How modern AI developed

2012
AlexNet
ImageNet winner
2022
ChatGPT (GPT-3.5)
Fastest-growing consumer app
2026
GPT-6 Astra
“Anything you can do on a computer, Astra can do for you.”
60 m parameters
1.9 years
175 bn parameters
5,500 years
10 trn parameters
317,100 years
Model size — time to count the parameters, one per second
7

Two ways to the same conclusion

If it rained, the street is wet.  The street is not wet.
pq ,  ¬q  ⊢  ¬p
Rule-based
Six rules, written by hand
ONE STEP = ONE APPLIED RULE
if rainedwet← premise
not wet← premise
if not wet → not rained← contraposition
∴ not rained← modus ponens
It did not rain.
Learning-based
Billions of parameters, learned from data
ONE STEP = ONE PREDICTED WORD
If it rained , the street is wet . The street is not wet . It did ?
not0.61
probably0.22
certainly0.07
It did not rain.
8

Opacity blocks responsible decision making

EXAMPLE
A recruiter decides on candidates for a position and an opaque AI system recommends rejecting Candidate Alice.
WITHOUT AI SUPPORT
WITH AI SUPPORT
ATTRIBUTABILITY 
Does the decision reflect the decision-maker’s judgment?
Can the grounds for the decision be critically evaluated?
Why do I consider Candidate Alice unsuitable?
Can the grounds underlying the AI recommendation be critically evaluated?
Why does the AI consider Candidate Alice unsuitable?
ACCOUNTABILITY 
Is the decision-maker answerable to others?
Can the decision be justified to affected stakeholders?
Can I justify the rejection to Candidate Alice?
Can the decision be justified when part of its grounds is inaccessible?
Can I justify the rejection to Candidate Alice?
Opaque AI weakens the basis for responsible judgment without reducing responsibility for the decision.
9

Explanations pre-date AI

An explanation is an answer to a why-question that communicates the causes of an outcome. (Miller, 2019)
Philosophy, psychology and cognitive science have studied how people give and receive explanations for decades.
Contrastiveness
Why this outcome rather than that one?
Selectivity
Which of the possible causes belong in the answer?
Sociality
Who is explaining, to whom, and in what context?
An AI explanation is information about how an AI system arrived at its output, conveyed to a user so that they can answer such a why-question.
It is therefore socio-technical: a technical system provides the explanatory information, and a social process makes it meaningful to the recipient.
10

From opacity to explanation

Opacity is a feature of modern AI and impedes informed decision-making. 
An AI explanation conveys how the system arrived at its output, so the user can answer that why-question.
What makes an explanation effective depends on the model, the recipient, and the situation. Explanations must therefore be designed for their intended use.
Whether an explanation works is revealed in how people use it to make decisions. Explanations must therefore be empirically evaluated.
Research Objective: To improve human-AI interaction through the design and evaluation of explanations.
11

Six papers

#
SHORT TITLE
AUTHORS
OUTLET
RANK*
STATUS
1
Exploring XAI Users' Needs
P. Schröppel, M. Förster
European Conference on Information Systems
A
Published 2024
2
Navigating the Rashomon Effect
J. Rosenberger, P. Schröppel, S. Kruschel, M. Kraus, P. Zschech, M. Förster
European Conference on Information Systems
A
Published 2025
3
User-Centric Chain-of-Thought Reasoning
P. Schröppel
Hawaii International Conference on System Sciences
B
Published 2026
4
Let's Chat About Overreliance
K. Pitz, P. Schröppel, C. Tille, K. Züllig, S. Zimmermann
Internationale Tagung Wirtschaftsinformatik
B
Published 2026
5
Choose Wisely
M. Förster, P. Schröppel, C. Schwenke, L. Fink, M. Klier
International Conference on Information Systems
A
Published 2024
6
Co-Creation to Enhance Reflective Decision-Making
M. Förster, P. Schröppel, C. Schwenke, M. Klier, L. Fink
European Journal of Information Systems
A
Submitted 2026
*VHB-Rating 2024 Wirtschaftsinformatik — vhbonline.org
12

From interaction pattern to research objective

LESS
USER AGENCY
MORE
AI SYSTEM
OUTPUT
USER
TOPIC · RESEARCH OBJECTIVES
Receiving
Accept or reject a finished output
TOPIC I
Personalization of AI explanations
RO1
To design and evaluate a personalization approach for post-hoc AI explanations.
RO2
To design and evaluate a personalization approach for intrinsically interpretable models.
Contesting
Probe and push back on its reasoning
TOPIC II
Scrutinization of LLM outputs
RO3
To design and evaluate a user-centric approach to chain-of-thought reasoning that makes LLM reasoning traces structured, inspectable, and correctable by users.
RO4
To understand how AI literacy training and AI explanations jointly shape appropriate reliance on LLM-based AI systems.
Shaping
Feed inputs into how it is generated
TOPIC III
Reflective decision making with AI
RO5
To design and evaluate an interactive AI explanation approach to foster reflective decision making.
RO6
To empirically investigate how co-creation of AI outputs shapes reflective decision making with AI.
13

Three topics, six objectives

#
SHORT TITLE
TOPIC
RESEARCH OBJECTIVE
1
Exploring XAI Users' Needs
Topic I — Personalization of AI explanations
RO1 — To design and evaluate a personalization approach for post-hoc AI explanations.
2
Navigating the Rashomon Effect
RO2 — To design and evaluate a personalization approach for intrinsically interpretable models.
3
User-Centric Chain-of-Thought Reasoning
Topic II — Scrutinization of LLM outputs
RO3 — To design and evaluate a user-centric approach to chain-of-thought reasoning that makes LLM reasoning traces structured, inspectable, and correctable by users.
4
Let's Chat About Overreliance
RO4 — To understand how AI literacy training and AI explanations jointly shape appropriate reliance on LLM-based AI systems.
5
Choose Wisely
Topic III — Reflective decision making with AI
RO5 — To design and evaluate an interactive AI explanation approach to foster reflective decision making.
6
Co-Creation to Enhance Reflective Decision-Making
RO6 — To empirically investigate how co-creation of AI outputs shapes reflective decision making with AI.
14

Methodology: the IS perspective

ISR FRAMEWORK — HEVNER ET AL. 2004
Business needs
Relevance
Applicable knowledge
Rigour
Environment
People
Roles, capabilities, characteristics
Organisations
Strategies, structure and culture, processes
Technology
Infrastructure, applications, architecture
IS research
Develop / build
Theories, artifacts
Assess↑ ↓Refine
Justify / evaluate
Analytical, case study, experimental, field study, simulation
Knowledge base
Foundations
Theories, frameworks, constructs, models, methods, instantiations
Methodologies
Data analysis techniques, formalisms, measures, validation criteria
Application in the appropriate environment
Additions to the knowledge base
15

Paradigm and evaluation tier

#
SHORT TITLE
PARADIGM / TIER
TASK
MODEL / EXPLANATION
METHOD / DATA
1
Exploring XAI Users' Needs
DSR / Human-grounded
Image classification
CNN / Feature importance
Online experiment / interaction data
2
Navigating the Rashomon Effect
DSR / Functionally- & human-grounded
Demand forecasting
GAM / Feature shape plots
Online experiment / system + interaction data
3
User-Centric Chain-of-Thought Reasoning
DSR / Functionally- & human-grounded
Math word problem solving
LLM / Chain-of-thought traces
Offline benchmarking + online experiment
4
Let's Chat About Overreliance
BSR / Human-grounded
Hate speech detection
LLM / Citations, references
Online experiment / interaction data
5
Choose Wisely
DSR / Human-grounded
Educational decision making
Text embedding / Concept explanations
Online experiment / survey
6
Co-Creation to Enhance Reflective Decision-Making
BSR / Application-grounded
Educational decision making
Text embedding / Concept explanations
Online field experiment / interaction + survey
16

Choose Wisely: motivation and research objective

PREMISE
AI can induce reflection when used for augmentation rather than substitution of human decision-making (Abdel-Karim et al. 2023).
PROBLEM
In critical contexts, an efficiency-focused use of AI is harmful: hasty acceptance of recommendations yields sub-optimal decisions with irreversible consequences (Gati et al. 1996). Facing a black box, users can only follow blindly or disregard (Brasse et al. 2023).
CONTEXT
Educational choices are non-routine tasks (Gati & Kulcsár 2021) with meaningful, long-term impact (Slaten & Baskin 2014) — yet state-of-the-art educational recommender systems are black boxes (Khanal et al. 2020).
RESEARCH GAP
Existing XAI methods are designed to justify recommendations, focusing on advice acceptance (Vale et al. 2022) — not on users' critical examination and exploration of alternatives. No XAI-based approach has been proposed with the aim of increasing users' reflection.
RESEARCH OBJECTIVE 5
To design and evaluate an AI explanation approach to foster reflective decision making.
17

How the design was derived

DSR PROCESS · PEFFERS ET AL. 2007
1
Problem identification
2
Objectives of a solution
3
Design and development
4
Demonstration
5
Evaluation
6
Communication
An AI-based solution for reflective educational decisions must enhance:
Exploration
Investigate a wider set of educational alternatives
Self-reflection
Examine one’s own thoughts, priorities and behaviour
Autonomy
Decide on one’s own judgment instead of accepting advice
Confidence
Belief in one’s ability to decide well-informedly
Trust
Calibrated reliance on the recommender
18

Deriving the three design steps

1
STEP 1
PREMISE
Humans reason in abstract concepts — skills, learning goals. AI reasons in pixels and tokens.
OBSTACLE
No shared vocabulary: users can neither follow nor enter the process.
Ghorbani et al. 2019 · Mucha et al. 2021
DESIGN STEP 1
Concepts as a shared language
Prediction splits in two over a concept layer: profile → concepts → alternatives (Koh et al. 2020).
SERVES Exploration · Trust
2
STEP 2
PREMISE
Users can only reflect if they understand why an alternative is recommended.
OBSTACLE
Post-hoc methods approximate the model from outside — they justify, not expose.
Brasse et al. 2023 · Förster et al. 2020
DESIGN STEP 2
Concept-based explanations
Relevance scores are an actual intermediate result — ante-hoc, faithful by construction.
SERVES Self-reflection · Exploration · Autonomy
3
STEP 3
PREMISE
Explanations must support follow-up interaction on human-understandable variables.
OBSTACLE
An explanation alone leaves the user a spectator of the process.
Adadi & Berrada 2018 · Chromik & Butz 2021
DESIGN STEP 3
Interventions on the explanation
Users adjust concept relevance and see the recommendation change — a causal intervention (Pearl 2009).
SERVES Autonomy · Confidence · Self-reflection
19

Artifact: two-step prediction

USER PROFILE
01Concept relevance
02Course ranking
{{ edges }}
User Profile Business Administration Statistics
Digital Business & Analytics {{ sa1 }}
Market Analysis & Machine Learning {{ sa2 }}
Strategy & Corporate Finance {{ sa3 }}
External Accounting {{ sa4 }}
Sustainable Corporate Governance {{ sa5 }}
20

Evaluation: experimental design

PROCEDURE
Data input
Round 1
Continue
Leave
Round 2
Continue
Leave
Round 3
Leave
Overview
GROUP 1 · CONTROL
Recommendation
Concept explanation
Concept adjustment
GROUP 2 · EXPLANATION
Recommendation
Concept explanation
Concept adjustment
GROUP 3 · FULL ARTIFACT
Recommendation
Concept explanation
Concept adjustment
21
ai-scouty.skill-kompass.de
22

Results

Exploration
ANOVA: F = 6.461   p = 0.002
1.0
2.0
3.0
4.0
5.0
6.0
7.0
**
***
5.60
5.16
5.22
Explanation &
intervention
Explanation
Control
Self-reflection
ANOVA: F = 3.930   p = 0.020
1.0
2.0
3.0
4.0
5.0
6.0
7.0
*
**
5.98
5.65
5.72
Explanation &
intervention
Explanation
Control
Autonomy
ANOVA: F = 1.477   p = 0.229
1.0
2.0
3.0
4.0
5.0
6.0
7.0
5.55
5.32
5.39
Explanation &
intervention
Explanation
Control
Confidence
ANOVA: F = 5.516   p = 0.004
1.0
2.0
3.0
4.0
5.0
6.0
7.0
*
**
5.78
5.41
5.49
Explanation &
intervention
Explanation
Control
Trust
ANOVA: F = 4.215   p = 0.015
1.0
2.0
3.0
4.0
5.0
6.0
7.0
*
**
5.68
5.36
5.39
Explanation &
intervention
Explanation
Control
Usefulness
ANOVA: F = 5.573   p = 0.004
1.0
2.0
3.0
4.0
5.0
6.0
7.0
**
**
6.06
5.73
5.70
Explanation &
intervention
Explanation
Control
* p<0.05, ** p<0.01, *** p<0.001 (LSD test). Bars show means, whiskers ± 1 standard error.
23

Limitations and future research

Limitation
Future research
1
Single use case — generalisation to other domains untested
Apply and evaluate the approach in other domains and with other target groups
2
Human-grounded evaluation only; participants were not personally affected by their decisions
Move to application-grounded evaluation, e.g. field experiments
3
Subjective measures dominate — besides exploration, no objective data
Add objective measures such as completion rates, plus qualitative methods
4
Concept extraction was a first pass only
Let users contribute their own personally relevant concepts
5
Interventions tested only as built on explanations
Test alternative interventions and other ways of combining them with explanations
24

Thank you

Questions welcome.
Philipp Noah Schröppel
Institut für Business Analytics · Universität Ulm
philipp.schroeppel@uni-ulm.de
Universität Ulm
25

Appendix

Supporting material
Construct items
Distribution of drop-out rates
Paper 1 — method and results
Paper 2 — method and results
Paper 3 — artifact and results
26

Construct items

Exploration
Sun et al. (2023)
AI Scouty found me a broad set of courses.
AI Scouty found courses for me that offered a lot of variety.
AI Scouty found interesting courses for me that I had not considered before.
AI Scouty found courses for me that covered areas where I want to make progress.
Self-reflection
Grant et al. (2002)
AI Scouty has helped me to think about which course I would like to take.
AI Scouty has helped me to think about what areas of interest I want to progress in.
Confidence
Taylor and Betz (1983)
I am confident that with the help of AI Scouty, I have at least found one course that suits my interests.
I am confident that with the help of AI Scouty, I would not worry about whether the course fits my interests or not.
When interacting with AI Scouty, I get information about courses I am interested in.
I am confident that with the help of AI Scouty, I can figure out what an ideal course would be.
Autonomy
Sankaran et al. (2021)
I really enjoyed this way of finding courses.
I think I would actively use AI Scouty during my studies to find courses.
AI Scouty allowed me to act when the course recommendation was not a good fit for me.
AI Scouty let me keep control of the course recommendations.
Trust
Benbasat and Wang (2005)
AI Scouty has the expertise to understand my needs and preferences about courses.
AI Scouty has good knowledge about courses.
AI Scouty puts my interests first.
AI Scouty wants to understand my needs and preferences.
AI Scouty provides unbiased course recommendations.
Usefulness
Benbasat and Wang (2005)
Using AI Scouty would improve reflective decision-making on which course to take.
Using AI Scouty would make it easier to make reflective decisions about which course to select.
Using AI Scouty would enhance my effectiveness in assessing a subset of individually relevant courses based on areas such as interests.
Using Scouty would make it easier to find suitable courses.
27

Distribution of drop-out rates

PARTICIPANTS LEAVING (% WITHIN GROUP) EXPLANATION & INTERVENTION EXPLANATION CONTROL
After first round 4.9% 13.2% 11.9%
After second round 21.5% 17.4% 12.6%
After third round 73.6% 69.5% 75.5%
28

Paper 1 — XAI personalization approach

Exploring XAI Users' Needs · ECIS 2024
Five-step XAI personalization loop
29

Paper 1 — algorithm

Steps map to the loop
Algorithm: XAI personalization approach
30

Paper 1 — study task

Image geolocation with saliency explanations
Screenshot of the study interface
31

Paper 1 — results

33 participants · 10 iterations
Hyperparameter distributions, reward-model variance, and mean effectiveness ratings
32

Paper 2 — approach

Personalizing model interpretability
Six-step interpretability personalization loop
Steps (0)–(5): reward model, configuration, training, constraint check, user evaluation, update.
33

Paper 2 — model variants

Shape functions across successive configurations
Progression of shape-function sets across configurations
Excluding features and constraining shapes trades detail for interpretability.
34

Paper 2 — reward distributions

Two hyperparameters · four levels
Histograms of mean reward by hyperparameter level
Preferred levels differ between participants — no single interpretable configuration dominates.
35

Paper 2 — learning behaviour

10 rounds
Reward-model variance per round and weights against cumulative reward
(a) variance declines monotonically · (b) weights rise with cumulative reward.
36

Paper 3 — user-centric chain of thought

Tagged reasoning → rendered trace
Tagged reasoning output and its rendered user-facing trace
Facts, goals, premises and results are tagged, then surfaced as referenced reasoning steps.
37

Paper 3 — results

User-centric CoT (n = 37) vs. standard CoT (n = 43)
Construct means with standard errors and box plots by condition
Ease of use and usefulness differ significantly (*); trust is comparable across conditions.
38