A Comparison on Cross-Domain Representation Learning for Skin-Lesion Classification
Published in Proceedings, Italian Workshop on Artificial Life and Evolutionary Computation (WIVACE), Springer, 2025
Recommended citation: Paolo Andreini, Simone Bonechi, Monica Bianchini, Barbara Toniella Corradini, Franco Scarselli, F. Manni, M. Quercioli. A Comparison on Cross-Domain Representation Learning for Skin-Lesion Classification. Communications in Computer and Information Science, 3001 (pp. 47-60). 2025. (BibTex)
Abstract
The advent of large–scale pre–training has produced powerful computer vision models, yet their effectiveness as feature extractors on highly specialized, out–of–distribution domains remains a critical issue. To fill this gap, we systematically compare the feature extraction capabilities of different, state–of–the–art architectures on the challenging task of dermoscopic skin–lesion classification—a domain substantially different from the natural images used during pre–training. We evaluate four distinct backbone architectures for feature extraction, encompassing both unimodal and multimodal models. Our selection includes EfficientNet B0 and Vision Transformer as representative unimodal models, along with Stable Diffusion Model and CLIP as state–of–the–art multimodal models. Using a subset of the ISIC archive, we assess the quality of extracted features via both linear probing and classification with a Multi Layer Perceptron. Our experiments suggest that all architectures produce highly effective features, achieving competitive performance. Notably, multimodal models, although pre–trained for tasks other than image classification, extract features that are remarkably competitive in this context. This work can serve as a comparative guide for the selection of feature extractors when tackling classification in specialized domains, such as medical imaging. The results highlight not only the generalizability of modern architectures, but also the surprising versatility of multimodal models as powerful feature extractors for interdisciplinary tasks
You can find the full paper here
