A Graph Path Interpretability Metric for Explainable Classification of Dermoscopic Skin Lesions
DOI:
https://doi.org/10.18287/JBPE26.12.030304Keywords:
dermoscopy, skin cancer, explainable artificial intelligence, interpretability, knowledge graph, deep learningAbstract
Deep neural networks can classify dermoscopic images of skin lesions at a level comparable to that of medical experts. However, their lack of transparency limits the interpretability of model decisions, while existing explainable artificial intelligence techniques often produce pixel-level attributions that are difficult to relate to clinical diagnostic criteria. This study proposes the Graph Path Interpretability (GPI) metric, a quantitative model-independent framework for assessing whether a model decision can be explained through a clinically meaningful path of the form “image – dermoscopic feature – diagnosis”. The study was conducted using the HAM10000 dataset containing 10015 images from seven diagnostic classes. An EfficientNet-B0 classifier achieved an accuracy of 0.81 and an AUC of 0.96 and was combined with five ABCD-rule features and a knowledge graph linking dermoscopic features and diagnoses through statistically significant associations identified using the Benjamini–Hochberg correction. Interpretability was evaluated using the path coverage metric GPIcov, the weighted path metric GPIw, and a faithfulness test based on sequential concept removal. Three reasoning modules – graph attention, graph convolution, and multilayer perceptron (MLP), were compared under similar classification performance. The value of GPIcov was close to 1.0, indicating that nearly all model decisions could be associated with conceptual paths in the knowledge graph. Under comparable predictive performance, graph-based models generated decisions supported by clinically meaningful paths, whereas the MLP model did not exploit conceptual relationships. The concept-removal test was consistent with the generated explanations. Evaluation on the independent ISIC 2018 dataset yielded an accuracy of 0.78 and an AUC of 0.94, while GPIcov remained close to 1.0. The results indicate that GPI can serve as a complementary quantitative measure of interpretability for dermoscopic image classification models.
References
1. G. Argenziano, H. P. Soyer, “Dermoscopy of pigmented skin lesions – a valuable tool for early,” The Lancet Oncology 2(7), 443–449 (2001). DOI: https://doi.org/10.1016/S1470-2045(00)00422-8
2. A. Esteva, B. Kuprel, R. A. Novoa, J. Ko, S. M. Swetter, H. M. Blau, and S. Thrun, “Dermatologist-level classification of skin cancer with deep neural networks,” Nature 542(7639), 115–118 (2017). DOI: https://doi.org/10.1038/nature21056
3. T. J. Brinker, A. Hekler, A. H. Enk, et al., “Deep learning outperformed 136 of 157 dermatologists in a head-to-head dermoscopic melanoma image classification task,” European Journal of Cancer 113, 47–54 (2019).
4. P. Tschandl, C. Rosendahl, and H. Kittler, “The HAM10000 dataset, a large collection of multi-source dermatoscopic images of common pigmented skin lesions,” Scientific Data 5(1), 180161 (2018). DOI: https://doi.org/10.1038/sdata.2018.161
5. B. V. Grechkin, V. O. Vinokurov, Yu. A. Khristoforova, and I. A. Matveeva, “VGG convolutional neural network classification of hyperspectral images of skin neoplasms,” Journal of Biomedical Photonics & Engineering 9(4), 040304 (2023). DOI: https://doi.org/10.18287/JBPE23.09.040304
6. S. N., V. A. R., S. R., K. N., P. M. M., and J. K., “Highly Accurate Skin Cancer Diagnosis Using HMT-NET and Vision Transformer Models,” Journal of Biomedical Photonics & Engineering 11(4), 040308 (2025). DOI: https://doi.org/10.18287/JBPE25.11.040308
7. I. A. Matveeva, A. I. Komlev, O. I. Kaganov, A. A. Moryatov, and V. P. Zakharov, “Multidimensional analysis of dermoscopic images and spectral information for the diagnosis of skin tumors,” Journal of Biomedical Photonics & Engineering 10(1), 010307 (2024). DOI: https://doi.org/10.18287/JBPE24.10.010307
8. I. A. Bratchenko, Yu. A. Khristoforova, L. A. Bratchenko, A. A. Moryatov, S. V. Kozlov, E. G. Borisova, and V. P. Zakharov, “Optical biopsy of amelanotic melanoma with Raman and autofluorescence spectra stimulated by 785 nm laser excitation,” Journal of Biomedical Photonics & Engineering 7(2), 020308 (2021). DOI: https://doi.org/10.18287/JBPE21.07.020308
9. A. Barredo Arrieta, N. Díaz-Rodríguez, J. Del Ser, A. Bennetot, S. Tabik, A. Barbado, S. Garcia, S. Gil-Lopez, D. Molina, R. Benjamins, R. Chatila, and F. Herrera, “Explainable Artificial Intelligence (XAI): Concepts, taxonomies, opportunities and challenges toward responsible AI,” Information Fusion 58, 82–115 (2020). DOI: https://doi.org/10.1016/j.inffus.2019.12.012
10. R. R. Selvaraju, M. Cogswell, A. Das, R. Vedantam, D. Parikh, and D. Batra, “Grad-CAM: visual explanations from deep networks via gradient-based localization,” in Proceedings of the IEEE International Conference on Computer Vision (ICCV), IEEE, 618–626 (2017). DOI: https://doi.org/10.1109/ICCV.2017.74
11. S. M. Lundberg, S.-I. Lee, “A unified approach to interpreting model predictions,” in Advances in Neural Information Processing Systems 30, 4765–4774 (2017).
12. F. Nachbar, W. Stolz, T. Merkle, A. B. Cognetta, T. Vogt, M. Landthaler, P. Bilek, O. Braun-Falco, and G. Plewig, “The ABCD rule of dermatoscopy: high prospective value in the diagnosis of doubtful melanocytic skin lesions,” Journal of the American Academy of Dermatology 30(4), 551–559 (1994). DOI: https://doi.org/10.1016/S0190-9622(94)70061-3
13. A. N. Zhdanova, “Ontological approach to predicting psychological characteristics of social media users using knowledge graphs,” Artificial Intelligence and Decision Making 3, 40–59 (2025). [in Russian]
14. M. Tan, Q. V. Le, “EfficientNet: rethinking model scaling for convolutional neural networks,” in Proceedings of the 36th International Conference on Machine Learning (ICML), PMLR 97, 6105–6114 (2019).
15. P. W. Koh, T. Nguyen, Y. S. Tang, S. Mussmann, E. Pierson, B. Kim, and P. Liang, “Concept bottleneck models,” in Proceedings of the 37th International Conference on Machine Learning (ICML), Proceedings of Machine Learning Research 119, 5338–5348 (2020).
16. P. Veličković, G. Cucurull, A. Casanova, A. Romero, P. Liò, and Y. Bengio, “Graph Attention Networks,” arXiv preprint arXiv:1710.10903 (2017).
17. V. Petsiuk, A. Das, and K. Saenko, “RISE: Randomized Input Sampling for Explanation of Black-box Models,” in Proceedings of the British Machine Vision Conference (BMVC) (2018)
18. K. Simonyan, A. Vedaldi, and A. Zisserman, “Deep Inside Convolutional Networks: Visualising Image Classification Models and Saliency Maps,” arXiv preprint arXiv:1312.6034 (2013).
19. M. T. Ribeiro, S. Singh, and C. Guestrin, “"Why Should I Trust You?": Explaining the Predictions of Any Classifier,” in Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (ACM, 2016), pp. 1135–1144. DOI: https://doi.org/10.1145/2939672.2939778
20. B. Zhou, A. Khosla, A. Lapedriza, A. Oliva, and A. Torralba, “Learning deep features for discriminative localization,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2921–2929 (2016). DOI: https://doi.org/10.1109/CVPR.2016.319
21. M. Sundararajan, A. Taly, and Q. Yan, “Axiomatic attribution for deep networks,” in Proceedings of the 34th International Conference on Machine Learning (ICML), PMLR 70, 3319–3328 (2017).
22. S. Bach, A. Binder, G. Montavon, F. Klauschen, K.-R. Müller, and W. Samek, “On Pixel-Wise Explanations for Non-Linear Classifier Decisions by Layer-Wise Relevance Propagation,” PLoS ONE 10(7), e0130140 (2015). DOI: https://doi.org/10.1371/journal.pone.0130140
23. B. Kim, M. Wattenberg, J. Gilmer, C. Cai, J. Wexler, F. Viégas, and R. Sayres, “Interpretability beyond feature attribution: quantitative testing with concept activation vectors (TCAV),” in Proceedings of the 35th International Conference on Machine Learning (ICML), PMLR 80, 2668–2677 (2018).
24. C. Chen, O. Li, D. Tao, A. Barnett, C. Rudin, and J. Su, “This looks like that: deep learning for interpretable image recognition,” in Advances in Neural Information Processing Systems 32, 8930–8941 (2019).
25. A. Ghorbani, J. Wexler, J. Y. Zou, and B. Kim, “Towards automatic concept-based explanations,” in Advances in Neural Information Processing Systems 32, 9273–9282 (2019).
26. A. Lucieri, M. N. Bajwa, S. A. Braun, M. I. Malik, A. Dengel, and S. Ahmed, “ExAID: A multimodal explanation framework for computer-aided diagnosis of skin lesions,” Computer Methods and Programs in Biomedicine 215, 106620 (2022). DOI: https://doi.org/10.1016/j.cmpb.2022.106620
27. C. Barata, M. E. Celebi, and J. S. Marques, “Explainable skin lesion diagnosis using taxonomies,” Pattern Recognition 110, 107413 (2021). DOI: https://doi.org/10.1016/j.patcog.2020.107413
28. C. Barata, M. E. Celebi, and J. S. Marques, “A survey of feature extraction in dermoscopy image analysis of skin cancer,” IEEE Journal of Biomedical and Health Informatics 23(3), 1096–1109 (2019). DOI: https://doi.org/10.1109/JBHI.2018.2845939
29. E. Tjoa, C. Guan, “A survey on explainable artificial intelligence (XAI): toward medical XAI,” IEEE Transactions on Neural Networks and Learning Systems 32(11), 4793–4813 (2021). DOI: https://doi.org/10.1109/TNNLS.2020.3027314
30. T. N. Kipf, M. Welling, “Semi-Supervised Classification with Graph Convolutional Networks,” arXiv preprint arXiv:1609.02907 (2016).
31. S. Hooker, D. Erhan, P.-J. Kindermans, and B. Kim, “A benchmark for interpretability methods in deep neural networks,” in Advances in Neural Information Processing Systems 32, 9737–9748 (2019).
32. J. Adebayo, J. Gilmer, M. Muelly, I. Goodfellow, M. Hardt, and B. Kim, “Sanity checks for saliency maps,” in Advances in Neural Information Processing Systems 31, 9505–9515 (2018).
33. J. DeYoung, S. Jain, N. F. Rajani, E. Lehman, C. Xiong, R. Socher, and B. C. Wallace, “ERASER: a benchmark to evaluate rationalized NLP models,” in Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics (Association for Computational Linguistics, 2020), pp. 4443–4458. DOI: https://doi.org/10.18653/v1/2020.acl-main.408
34. W. Samek, A. Binder, G. Montavon, S. Lapuschkin, and K.-R. Muller, “Evaluating the Visualization of What a Deep Neural Network Has Learned,” IEEE transactions on neural networks and learning systems 28(11), 2660–2673 (2017). DOI: https://doi.org/10.1109/TNNLS.2016.2599820
35. G. Argenziano, H. P. Soyer, S. Chimenti, et al., “Dermoscopy of pigmented skin lesions: Results of a consensus meeting via the Internet,” Journal of the American Academy of Dermatology 48(5), 679–693 (2003).
36. C. Rudin, “Stop explaining black box machine learning models for high stakes decisions and use interpretable models instead,” Nature Machine Intelligence 1(5), 206–215 (2019). DOI: https://doi.org/10.1038/s42256-019-0048-x
37. A. Krizhevsky, I. Sutskever, and G. E. Hinton, “ImageNet classification with deep convolutional neural networks,” in Advances in Neural Information Processing Systems 25, 1097–1105 (2012).
38. K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 770–778 (2016). DOI: https://doi.org/10.1109/CVPR.2016.90
39. A. Dosovitskiy, L. Beyer, A. Kolesnikov, “An image is worth 16×16 words: transformers for image recognition at scale,” arXiv preprint arXiv:2010.11929 (2021).
40. H. A. Haenssle, C. Fink, R. Schneiderbauer, et al., “Man against machine: diagnostic performance of a deep learning convolutional neural network for dermoscopic melanoma recognition in comparison to 58 dermatologists,” Annals of Oncology 29(8), 1836–1842 (2018).
41. P. Tschandl, N. Codella, B. N. Akay, G. Argenziano, R. P. Braun, H. Cabo, D. Gutman, A. Halpern, B. Helba, R. Hofmann-Wellenhof, A. Lallas, J. Lapins, C. Longo, J. Malvehy, M. A. Marchetti, A. Marghoob, S. Menzies, A. Oakley, J. Paoli, S. Puig, C. Rinner, C. Rosendahl, A. Scope, C. Sinz, H. P. Soyer, L. Thomas, I. Zalaudek, and H. Kittler, “Comparison of the accuracy of human readers versus machine-learning algorithms for pigmented skin lesion classification: an open, web-based, international, diagnostic study,” The Lancet Oncology 20(7), 938–947 (2019). DOI: https://doi.org/10.1016/S1470-2045(19)30333-X
42. S. S. Han, M. S. Kim, W. Lim, G. H. Park, I. Park, and S. E. Chang, “Classification of the Clinical Images for Benign and Malignant Cutaneous Tumors Using a Deep Learning Algorithm,” Journal of Investigative Dermatology 138(7), 1529–1538 (2018). DOI: https://doi.org/10.1016/j.jid.2018.01.028
43. L. Yu, H. Chen, Q. Dou, J. Qin, and P. A. Heng, “Automated melanoma recognition in dermoscopy images via very deep residual networks,” IEEE Transactions on Medical Imaging 36(4), 994–1004 (2017). DOI: https://doi.org/10.1109/TMI.2016.2642839
44. P. Tschandl, C. Rinner, Z. Apalla, G. Argenziano, N. Codella, A. Halpern, M. Janda, A. Lallas, C. Longo, J. Malvehy, J. Paoli, S. Puig, C. Rosendahl, H. P. Soyer, I. Zalaudek, and H. Kittler, “Human-computer collaboration for skin cancer recognition,” Nature Medicine 26(8), 1229–1234 (2020). DOI: https://doi.org/10.1038/s41591-020-0942-0
45. N. Codella, V. Rotemberg, P. Tschandl, M. E. Celebi, S. Dusza, D. Gutman, B. Helba, A. Kalloo, K. Liopyris, M. Marchetti, H. Kittler, and A. Halpern, “Skin Lesion Analysis Toward Melanoma Detection 2018: A Challenge Hosted by the International Skin Imaging Collaboration (ISIC),” arXiv preprint arXiv:1902.03368 (2019).
46. M. Combalia, N. C. F. Codella, V. Rotemberg, B. Helba, V. Vilaplana, O. Reiter, C. Carrera, A. Barreiro, A. C. Halpern, S. Puig, and J. Malvehy, “BCN20000: Dermoscopic Lesions in the Wild,” arXiv preprint arXiv:1908.02288 (2019).
47. J. Kawahara, S. Daneshvar, G. Argenziano, and G. Hamarneh, “Seven-point checklist and skin lesion classification using multitask multimodal neural nets,” IEEE Journal of Biomedical and Health Informatics 23(2), 538–546 (2019). DOI: https://doi.org/10.1109/JBHI.2018.2824327
48. F. Doshi-Velez, B. Kim, “Towards A Rigorous Science of Interpretable Machine Learning,” arXiv preprint arXiv:1702.08608 (2017).
Downloads
Published
Issue
Section
License
Copyright (c) 2026 A. N. Zhdanova

This work is licensed under a Creative Commons Attribution 4.0 International License.
Distribution or reproduction of this work in whole or in part requires full attribution of the original publication, including its DOI.













