Abstract
Traditional benchmarking approaches for evaluating AI systems face limitations when applied to complex, real-world domains such as legal decision-making, medical research, and scientific discovery. These limitations include missing or incorrect ground truth data, data leakage, confirmation bias, and deployment mismatches that make systematic evaluation challenging. Here, we propose a novel approach explanation as evaluation, which uses the quality of AI-generated explanations as a proxy for overall system performance. Our process operates across three phases: (1) Design and Development, where explanation quality dimensions are selected; (2) Deployment, where explanations are generated alongside system outputs; and (3) Evaluation, where explanations are systematically assessed against the selected dimensions. We argue that this approach addresses key desiderata for AI evaluation: the ability to evaluate multifaceted outputs without ground truth, support for continuous evaluation, efficient use of human expertise, and ready application across diverse domains. We demonstrate the feasibility of our approach through a proof-of-concept evaluation in clinical trial outcome prediction. Using explanations for AI evaluation builds on the growing regulatory requirement for explainable AI systems, potentially enabling more robust and continuous assessment of AI performance in complex, dynamic domains.
| Original language | English |
|---|---|
| Journal | CEUR Workshop Proceedings |
| Volume | 4210 |
| State | Published - 2026 |
| Externally published | Yes |
| Event | 2nd Workshop on Trust, Autonomy and Accountability in PKG-Based Agentic AI, TAAPAAI 2026 - Dubrovnik, Croatia Duration: May 10 2026 → May 10 2026 |
Keywords
- AI system evaluation
- XAI
- explainable AI
Fingerprint
Dive into the research topics of 'Explanation as Evaluation: Using explanation quality to measure AI system performance'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver