Explainable Artificial Intelligence for Black-Box Machine Learning: A Taxonomy of Model-Agnostic and Deep Explanation Methods
Keywords:
Explainable artificial intelligence, model transparency, XAIAbstract
The rapid adoption of increasingly complex machine-learning models has intensified concerns regarding transparency, interpretability, accountability, and user trust. Deep neural networks, ensemble learners, and other high-capacity predictive models frequently provide superior predictive performance while offering limited insight into the mechanisms underlying individual predictions. Explainable artificial intelligence (XAI) has consequently emerged as an important research area concerned with making algorithmic decisions understandable to developers, domain specialists, regulators, and end users. This review examines the principal explainability approaches available through 2020 and organizes them according to model dependence, explanation scope, explanation representation, and underlying computational strategy. Particular attention is given to model-agnostic approaches such as Local Interpretable Model-Agnostic Explanations and SHapley Additive exPlanations, together with neural-network-specific approaches including saliency mapping, layer-wise relevance propagation, Grad-CAM, DeepLIFT, integrated gradients, and concept-based explanations. The review discusses the distinction between local and global interpretability, post-hoc and intrinsically interpretable models, and feature-attribution and example-based explanations. It further considers major limitations involving explanation fidelity, instability, computational cost, human interpretability, and the absence of standardized evaluation frameworks. The paper concludes that explainability should not be regarded as a single algorithmic property but as a multidimensional requirement involving the predictive model, explanation method, intended audience, and decision context.
References
Barredo Arrieta A, Díaz-Rodríguez N, Del Ser J, Bennetot A, Tabik S, Barbado A, et al. Explainable artificial intelligence (XAI): concepts, taxonomies, opportunities and challenges toward responsible AI. Inf Fusion. 2020;58:82-115.
Guidotti R, Monreale A, Ruggieri S, Turini F, Giannotti F, Pedreschi D. A survey of methods for explaining black box models. ACM Comput Surv. 2018;51(5):93.
Adadi A, Berrada M. Peeking inside the black-box: a survey on explainable artificial intelligence (XAI). IEEE Access. 2018;6:52138-52160.
Ribeiro MT, Singh S, Guestrin C. “Why should I trust you?” Explaining the predictions of any classifier. In: Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining. New York: ACM; 2016. p. 1135-1144.
Lundberg SM, Lee SI. A unified approach to interpreting model predictions. Adv Neural Inf Process Syst. 2017;30:4765-4774.
Selvaraju RR, Cogswell M, Das A, Vedantam R, Parikh D, Batra D. Grad-CAM: visual explanations from deep networks via gradient-based localization. In: Proceedings of the IEEE International Conference on Computer Vision. 2017. p. 618-626.
Sundararajan M, Taly A, Yan Q. Axiomatic attribution for deep networks. Proc Mach Learn Res. 2017;70:3319-3328.
Shrikumar A, Greenside P, Kundaje A. Learning important features through propagating activation differences. Proc Mach Learn Res. 2017;70:3145-3153.
Bach S, Binder A, Montavon G, Klauschen F, Müller KR, Samek W. On pixel-wise explanations for non-linear classifier decisions by layer-wise relevance propagation. PLoS One. 2015;10(7).
Montavon G, Samek W, Müller KR. Methods for interpreting and understanding deep neural networks. Digit Signal Process. 2018;73:1-15.
Lipton ZC. The mythos of model interpretability. Queue. 2018;16(3):31-57.
Doshi-Velez F, Kim B. Towards a rigorous science of interpretable machine learning. arXiv. 2017;1702.08608.
Adebayo J, Gilmer J, Muelly M, Goodfellow I, Hardt M, Kim B. Sanity checks for saliency maps. Adv Neural Inf Process Syst. 2018;31:9505-9515.
Kim B, Wattenberg M, Gilmer J, Cai C, Wexler J, Viégas F, Sayres R. Interpretability beyond feature attribution: quantitative testing with concept activation vectors. Proc Mach Learn Res. 2018;80:2668-2677.
Fong RC, Vedaldi A. Interpretable explanations of black boxes by meaningful perturbation. In: Proceedings of the IEEE International Conference on Computer Vision. 2017. p. 3429-3437.
Zeiler MD, Fergus R. Visualizing and understanding convolutional networks. In: Fleet D, Pajdla T, Schiele B, Tuytelaars T, editors. Computer Vision—ECCV 2014. Cham: Springer; 2014. p. 818-833.
Simonyan K, Vedaldi A, Zisserman A. Deep inside convolutional networks: visualising image classification models and saliency maps. arXiv. 2013;1312.6034.
Ribeiro MT, Singh S, Guestrin C. Anchors: high-precision model-agnostic explanations. Proc AAAI Conf Artif Intell. 2018;32(1):1527-1535.
Hooker S, Erhan D, Kindermans PJ, Kim B. A benchmark for interpretability methods in deep neural networks. Adv Neural Inf Process Syst. 2019;32:9737-9748.
Carvalho DV, Pereira EM, Cardoso JS. Machine learning interpretability: a survey on methods and metrics. Electronics. 2019;8(8):832.