Mel-DEPTHS: a benchmark dataset for epidermis and tumor segmentation for melanoma staging


TOPUZ Y., GÖKCAN M. T., Men A. M. Ö., Yıldız S., Sertbudak İ., Kaymaz S., ...More

Medical and Biological Engineering and Computing, 2026 (SCI-Expanded, Scopus)

  • Publication Type: Article / Article
  • Publication Date: 2026
  • Doi Number: 10.1007/s11517-026-03656-3
  • Journal Name: Medical and Biological Engineering and Computing
  • Journal Indexes: Science Citation Index Expanded (SCI-EXPANDED), Scopus, ABI/INFORM, Aerospace Database, Applied Science & Technology Source, BIOSIS, CINAHL, Compendex, EMBASE, INSPEC, MEDLINE, Academic Search Ultimate (EBSCO), Natural Science Collection (ProQuest), Biological Science Database (ProQuest), Biomedical Reference Collection: Corporate Edition (EBSCO), Business Source Ultimate (EBSCO), Engineering Source (EBSCO), Health Research Premium Collection (ProQuest), Pharma Collection (ProQuest), Technology Collection (ProQuest)
  • Keywords: Digital pathology, Iterative self-training, Melanoma, Pixel level annotation
  • Yıldız Technical University Affiliated: Yes

Abstract

Abstract: Accurate delineation of epidermis and tumor boundaries is central to melanoma staging, yet pixel-level annotation on whole-slide images (WSIs) is labor-intensive and inconsistent across observers. Advancing this field requires standardized, publicly available benchmarks with expert-validated labels. We introduce Mel-DEPTHS, a new benchmark dataset for epidermis and tumor segmentation, designed to accelerate and standardize research for automated melanoma staging. Mel-DEPTHS comprises 50 anonymized melanoma WSIs (40x, 0.25m/pixel) with pixel-level masks for epidermis and tumor regions. Clinical variables such as invasion depth, ulceration, and pT stage are provided alongside fixed train/test partitions to ensure reproducibility. To mitigate annotation burden, we developed an Expert-Supervised Iterative Self-Training (ESIST) protocol: a pretrained model generates pseudo-labels, which dermatopathologists iteratively refine for retraining. We benchmarked six state-of-the-art segmentation models (UNet, UNet++, UNet3+, UPerNet, TransUNet, ConvUNeXt) using WSI-level precision, recall, IoU, and Dice. TransUNet achieved the best performance, closely followed by ConvUNeXt and UperNet. Three-fold cross-validation also confirmed consistent model rankings and label robustness. Mel-DEPTHS provides the fidelity and diversity necessary for clinically meaningful segmentation. It establishes a standardized benchmark and fosters reproducibility in computational pathology.