<div class="csl-bib-body">
<div class="csl-entry">Azzolin, S., Teso, S., Lepri, B., Passerini, A., & Malhotra, S. (2026). GNN Explanations that do not Explain and How to find Them. In <i>The Fourteenth International Conference on Learning Representations : ICLR 2026</i>. The Fourteenth International Conference on Learning Representations (ICRL 2026), Rio de Janeiro, Brazil.</div>
</div>
-
dc.identifier.uri
http://hdl.handle.net/20.500.12708/229326
-
dc.description.abstract
Explanations provided by Self-explainable Graph Neural Networks (SE-GNNs) are fundamental for understanding the model's inner workings and for identifying potential misuse of sensitive attributes. Although recent works have highlighted that these explanations can be suboptimal and potentially misleading, a characterization of their failure cases is unavailable. In this work, we identify a critical failure of SE-GNN explanations: explanations can be unambiguously unrelated to how the SE-GNNs infer labels. We show that, on the one hand, many SE-GNNs can achieve optimal true risk while producing these degenerate explanations, and on the other, most faithfulness metrics can fail to identify these failure modes. Our empirical analysis reveals that degenerate explanations can be maliciously planted (allowing an attacker to hide the use of sensitive attributes) and can also emerge naturally, highlighting the need for reliable auditing. To address this, we introduce a novel faithfulness metric that reliably marks degenerate explanations as unfaithful, in both malicious and natural settings. Our code is available on GitHub.
en
dc.description.sponsorship
FWF - Österr. Wissenschaftsfonds
-
dc.language.iso
en
-
dc.subject
Graph Neural Networks
en
dc.subject
Explainability
en
dc.subject
Interpretability
en
dc.subject
Positive Existential First-order logic
en
dc.subject
Logic and Explainability
en
dc.title
GNN Explanations that do not Explain and How to find Them
en
dc.type
Inproceedings
en
dc.type
Konferenzbeitrag
de
dc.contributor.affiliation
University of Trento, Italy
-
dc.contributor.affiliation
University of Trento, Italy
-
dc.contributor.affiliation
Fondazione Bruno Kessler, Italy
-
dc.contributor.affiliation
University of Trento, Italy
-
dc.relation.grantno
I 6728
-
dc.type.category
Full-Paper Contribution
-
tuw.booktitle
The Fourteenth International Conference on Learning Representations : ICLR 2026
-
tuw.peerreviewed
true
-
tuw.project.title
NanoX
-
tuw.researchTopic.id
I1
-
tuw.researchTopic.name
Logic and Computation
-
tuw.researchTopic.value
100
-
tuw.publication.orgunit
E194-06 - Forschungsbereich Machine Learning
-
dc.description.numberOfPages
49
-
tuw.author.orcid
0009-0005-3418-0585
-
tuw.author.orcid
0000-0002-2340-9461
-
tuw.event.name
The Fourteenth International Conference on Learning Representations (ICRL 2026)
en
tuw.event.startdate
23-04-2026
-
tuw.event.enddate
27-04-2026
-
tuw.event.online
On Site
-
tuw.event.type
Event for scientific audience
-
tuw.event.place
Rio de Janeiro
-
tuw.event.country
BR
-
tuw.event.presenter
Azzolin, Steve
-
wb.sciencebranch
Informatik
-
wb.sciencebranch
Wirtschaftswissenschaften
-
wb.sciencebranch.oefos
1020
-
wb.sciencebranch.oefos
5020
-
wb.sciencebranch.value
90
-
wb.sciencebranch.value
10
-
item.languageiso639-1
en
-
item.cerifentitytype
Publications
-
item.openairetype
conference paper
-
item.grantfulltext
none
-
item.openairecristype
http://purl.org/coar/resource_type/c_5794
-
item.fulltext
no Fulltext
-
crisitem.project.funder
FWF - Österr. Wissenschaftsfonds
-
crisitem.project.grantno
I 6728
-
crisitem.author.dept
University of Trento, Italy
-
crisitem.author.dept
University of Trento, Italy
-
crisitem.author.dept
Fondazione Bruno Kessler, Italy
-
crisitem.author.dept
University of Trento, Italy
-
crisitem.author.dept
E194-06 - Forschungsbereich Machine Learning
-
crisitem.author.orcid
0009-0005-3418-0585
-
crisitem.author.orcid
0000-0002-2340-9461
-
crisitem.author.parentorg
E194 - Institut für Information Systems Engineering