
International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 03 | Mar 2026 www.irjet.net p-ISSN: 2395-0072
![]()

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 03 | Mar 2026 www.irjet.net p-ISSN: 2395-0072
M. Pravallika1, MD. Kouser Ali2, M. Dinesh Kumar3, P. Sai Mehar4
1 Student, Dept. of Electronics and Communication Engineering, Seshadri Rao Gudlavalleru Engineering College, Gudlavalleru, Andhra Pradesh, India
2 Student, Dept. of Electronics and Communication Engineering, Seshadri Rao Gudlavalleru Engineering College, Gudlavalleru, Andhra Pradesh, India
3 Student, Dept. of Electronics and Communication Engineering, Seshadri Rao Gudlavalleru Engineering College, Gudlavalleru, Andhra Pradesh, India
4 Student, Dept. of Electronics and Communication Engineering, Seshadri Rao Gudlavalleru Engineering College, Gudlavalleru, Andhra Pradesh, India
Abstract - Skin disease, especially melanoma, is among the most dangerous health conditions worldwide. Early detection is useful for successful treatment. However, diagnosis of skin lesions requires dermatological expertise, which may be unavailable in rural or resource-limited regions, and can be highly time-consuming. There is a gap that can be fulfilled by accurate and accessible diagnostic support systems. A computer-aided detection tool can assist clinicians by providing a fast, reliable second opinion, reducing delays, and increasing diagnostic consistency for skin cancer detection. This project implements a deep learning-basedskincancer classification systemthat is used for the classification of melanoma and benign classes using convolutional neural network (CNN) models, namely MobileNetV2_S and EfficientNetB3. It is implemented in TensorFlow through transfer learning. Skin lesion images taken from the ISIC 2020 dataset are resized and normalized before being fed to the CNN models for classification into benign and melanoma. CNNs' models are used for feature extraction, which learns spatial lesion patterns. The proposed method also includes data augmentation, class balancing, fine-tuning, threshold optimization,hyperparameters,andtesttimeaugmentation to improve classification performance. The performance of the models is evaluated by using metrics such as accuracy, precision, recall, specificity, F1-score, and AUC. These experimental results show EfficientNetB3 achieved the best overall test performance with 94.72% accuracy, 96.26% specificity, 94.06% melanoma F1-score, and 0.9860 AUC, while MobileNetV2_S provided competitive performance with 92.24% accuracy and lower computational cost, which resultsinlowtrainingtime.
Key Words: melanoma, deep learning, transfer learning, computer-aided diagnosis, MobileNetV2, EfficientNetB3
Skin cancer is one of the most common forms of cancer, and melanoma is clinically significant because of its high metastatic potential when diagnosis is delayed. Visual examination and dermoscopy remain central to routine
screening, but interpretation is influenced by clinical experience, lesion diversity, imaging variability, and time constraints. As a result, automated systems that can analyze dermoscopic images consistently and rapidly are increasinglyvaluableasdecision-supporttools[1],[2].
Recent advances in deep learning have transformed medical image analysis by allowing convolutional neural networks to learn discriminative patterns directly from images. In dermatology, these networks can capture texture, pigment distribution, asymmetry, border irregularity, and other lesion characteristics that are difficult to encode manually. Public datasets such as ISIC have accelerated this research by providing large-scale dermoscopiccollectionsformelanomaanalysis[3],[7].
The objective ofthiswork istodevelopandcomparetwo transfer-learning based binary classifiers, MobileNetV2_S and EfficientNetB3, for skin lesion classification into benign and melanoma categories. The implementation emphasizes practical deployment features: lightweight preprocessing, data augmentation, class balancing, twostage fine-tuning, threshold optimization constrained by specificity, and test-time augmentation. The main contributions of the paper are: (i) a reproducible TensorFlow-based binary classification pipeline for ISIC2020 images, (ii) a detailed comparison between a lightweight and a stronger backbone under the same training protocol, and (iii) quantitative analysis of classification quality and inference cost for clinicianassistancescenarios.
Existing literature shows clear progression from handcrafted feature pipelines to end-to-end deep neural networks for melanoma detection. The review by Kaur et al. highlights that high-performing computer-aided diagnosis systems often combine careful preprocessing, lesion-focused analysis, and deep classification models to improve melanoma discrimination [1]. Their study also underlines the relevance of ISIC2020 as a challenging benchmarkformodernmelanomaCADresearch.

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 03 | Mar 2026 www.irjet.net p-ISSN: 2395-0072
Practical systems have used pretrained CNN backbones such as VGG-based networks and MobileNet to enable accessible skin cancer screening. Mohankumar et al. presented a web-oriented diagnostic workflow for benign and malignant classification and showed that transfer learning can support real-time clinician assistance with satisfactory validation performance [2]. These studies confirm that transfer learning is effective even when medical datasets are smaller than large natural-image corpora.
Dataset quality and evaluation design remain equally important. Rotemberg et al. described the patient-centric SIIM-ISIC 2020 dataset and its value for melanoma researchinclinicallymeaningfulsettings[3].Cassidyetal. latershowedthatISIC imagecollectionscontainduplicate andnear-duplicatesamples thatcanbiasevaluationifnot handledcarefully[7].Motivatedbytheseobservations,the present work focuses on a controlled split-based evaluation with threshold tuning and robust test-time averaging, while comparing MobileNetV2 [4] and EfficientNet[5]underidenticalexperimentalconditions.
3.1
TheimplementationusestheISIC2020dermoscopicimage collectionorganizedintotrain,validation,andtestfolders. After extraction, the training pipeline processed 9,017 imagesfortraining,1,287imagesforvalidation,and2,578 images for final testing. The task is binary classification, where class label 0 corresponds to benign lesions, and class label 1 corresponds to melanoma. The test-set confusion matrices show 1,419 benign and 1,159 melanoma images, indicating that the evaluation set remainsclinicallymeaningfulforbothclasses.
All images were resized to 224 × 224 × 3. Backbonespecific preprocessing from TensorFlow Keras applicationswasappliedbeforemodel input.Thetraining generator used moderate augmentation to improve generalization:rotationupto20degrees,widthandheight shifts of 0.03, zoom up to 0.08, horizontal flipping, and nearest-neighborfillingforemptypixels.
3.2
Two pretrained CNN backbones were investigated. The first was MobileNetV2_S, implemented with MobileNetV2 andwidthmultiplieralpha=0.75toreducecomputational cost.ThesecondwasEfficientNetB3,selectedasastronger backbone with higher representational capacity. Both models were initialized with ImageNet weights and used globalaveragepoolingatthebackboneoutput.
A common custom classification head was attached to each backbone. This head consisted of batch normalization, a dropout layer of 0.35, a fully connected layer with 512 units, batch normalization, LeakyReLU activation, a dropout layer of 0.40, a second dense layer
with 256 units, batch normalization, LeakyReLU activation, a dropout layer of 0.30, and a final sigmoid classifier for binary prediction. The total parameter count was approximately 2.18 million for MobileNetV2_S and 11.71millionforEfficientNetB3.
Training was performed in two phases. In Phase 1, the backbone was frozen and only the custom classifier head was trained for 15 epochs. In Phase 2, the top 30% of the backbone layers, excluding batch-normalization layers, wereunfrozenforfine-tuningwithanadditionalschedule of up to 35 epochs. The Adam optimizer was used with learningrate1e-3inPhase1and1e-5duringfine-tuning.
Binary cross-entropy was used as the loss function. Accuracy, precision, recall, and AUC were monitored duringtraining.Tohandle classimbalance,balancedclass weightswerecomputedfromthetraininglabels,resulting inweightsof0.9081forbenignand1.1127formelanoma.
Model Checkpoint, Reduce LROnPlateau, and Early Stopping callbacks were configured to maximize validationofAUCandpreventoverfitting.
Instead of using a fixed decision threshold of 0.50 for all predictions,thevalidationsetprobabilitiesweresearched over thresholds from 0.10 to 0.90 in 161 steps. Only thresholds satisfying a minimum validation specificity of 0.95wereconsidered,andthefinalthresholdwasselected using the highest F1 score within that feasible set. This design favors lower false-positive rates during tuning; however, varying generalization on the unseen test set resulted in final test specificities of 91.05% for MobileNetV2_Sand96.26%forEfficientNetB3.
Forfinalinference,afive-passtest-timeaugmentationwas used. Each test image was evaluated once in its original form and four additional times with horizontal flipping, rotation up to 10 degrees, and zoom up to 0.05. The probabilities from all passes were averaged to obtain a morestablefinalscore.
The models were evaluated using accuracy, weighted precision, weighted recall, specificity, weighted F1-score, melanoma-specific precision/recall/F1, and ROC-AUC. Recall for the melanoma class is equivalent to sensitivity in this binary formulation. In addition to classification quality, training time and average per-image inference timewererecordedtoassessdeploymentfeasibility.
The principal hyperparameters and implementation settings used in the final experiments are summarized in Table 1. All values in this table were taken directly from

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 03 | Mar 2026 www.irjet.net p-ISSN: 2395-0072
the Python notebook used for model training and evaluation.
Table -1: Hyperparametersettingsusedinmodeltraining andevaluation
Parameter Value / Setting
Randomseed 42
Inputimagesize 224×224×3
Training/validation batchsize 32
Testbatchsize 1
Phase1training 15epochswithfrozen backbone
Phase2fine-tuning 35epochswithtop30%of backbonelayersunfrozen; BatchNormlayerskeptfrozen
Optimizer Adam
Headlearningrate 1e-3
Fine-tuninglearning rate 1e-5
Lossfunction Binarycross-entropy(label smoothing=0.0)
Classimbalance handling Balancedclassweights enabled
Trainingaugmentation
Rotation=20,widthshift= 0.03,heightshift=0.03,zoom =0.08,horizontalflip
Thresholdsearch Thresholdsfrom0.10to0.90 in161steps;minimum validationspecificity=0.95
Test-time augmentation 5passestotal:1original+4 augmentedpasseswith horizontalflip,rotation=10, zoom=0.05
The experimental results are illustrated in Figures 1-10, summarized numerically in Table 2 and Table 3, and finally synthesized in Figure 11. EfficientNetB3 achieved the highest overall accuracy, weighted F1-score,
specificity, and ROC-AUC, whereas MobileNetV2_S remained competitive while requiring fewer parameters andlowerinferencetime.
Thisindicatesapracticaltrade-offforclinicaldeployment. EfficientNetB3 is the stronger overall model when maximum predictive performance and balanced accuracy are required. Conversely, MobileNetV2_S achieved a slightly higher melanoma recall (93.70% vs. 92.84%), making it a highly sensitive, lightweight option suited for first-line screening environments where minimizing false negatives (missed melanomas) is the highest clinical priority.
Furthermore, the proposed EfficientNetB3 pipeline demonstrated competitive performance relative to recent literature. Kaur et al. [1] reported 93.40% classification accuracy in their melanoma CAD framework, while Mohankumar et al. [2] reported 92% validation accuracy in a deep learning system using VGG19 and MobileNetV2. In comparison, the proposed EfficientNetB3 model achieved 94.72% test accuracy and 0.9860 ROC-AUC on ISIC2020. Although direct comparison across studies should be interpreted carefully because of differences in preprocessing,splits,andevaluationsettings,theseresults indicatethattheproposedcombinationofclassweighting, two-stage fine-tuning, threshold optimization, and fivepasstest-timeaugmentationprovidesarobustframework formelanomaclassification.

-1: MobileNetV2_SaccuracyhistoryfromthePython notebookoutput.

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 03 | Mar 2026 www.irjet.net p-ISSN: 2395-0072

Fig -2: MobileNetV2_SlosshistoryfromthePython notebookoutput.

Fig -3:MobileNetV2_SAUChistoryfromthePython notebookoutput.

Fig -4: MobileNetV2_SROCcurvefromthePython notebookoutput.

Fig -5: MobileNetV2_SconfusionmatrixfromthePython notebookoutput.

Fig -6: EfficientNetB3accuracyhistoryfromthePython notebookoutput.

Fig -7: EfficientNetB3losshistoryfromthePython notebookoutput.

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 03 | Mar 2026 www.irjet.net p-ISSN: 2395-0072

Fig -8: EfficientNetB3AUChistoryfromthePython notebookoutput.

Fig -9: EfficientNetB3ROCcurvefromthePython notebookoutput.

Fig -10:EfficientNetB3confusionmatrixfromthePython notebookoutput.
Table -2: Comparative Test Performance of the Proposed Models
Table -3: Melanoma-Specific and Computational Performance of the Proposed Models

Fig -11: Comparisonofkeyperformancemetricsderived fromtheprojectoutputs.

Volume: 13 Issue: 03 | Mar 2026 www.irjet.net p-ISSN: 2395-0072
MobileNetV2_S reached 92.24% test accuracy, 91.05% specificity, and 93.70% melanoma recall at an optimized threshold of 0.350. Its confusion matrix shows 1,292 benign lesions correctly identified, 127 benign lesions misclassified as melanoma, 1,086 melanoma lesions correctlyidentified,and73melanomalesionsmissed.The best validation accuracy observed during training was approximately 93.01%, with a best validation AUC of 97.90%.
EfficientNetB3 produced the best overall test result with 94.72% accuracy, 96.26% specificity, 95.31% melanoma precision,and0.9860ROC-AUCatathresholdof0.495.Its confusion matrix shows 1,366 benign lesions correctly classified, only 53 benign false alarms, 1,076 melanoma lesions correctly classified, and 83 melanoma misses. The model also achieved a higher best validation accuracy of about93.86%andabestvalidationAUCof98.26%.
The training curves further show that both models converged smoothly after transfer learning, with gradual improvementintrainingaccuracyandAUC.EfficientNetB3 maintained a stronger validation trend and lower falsepositive burden, which explains its superior specificity. However, MobileNetV2_S recorded a slightly higher melanoma recall, meaning it missed fewer melanoma casesontheheld-outtestset.Thisdistinctionisimportant in clinical settings: depending on whether the application prioritizessensitivityoroverallbalance,eithermodelmay be preferred. The slightly higher final test AUC values relativetothebestvalidationofAUCareplausiblebecause final testing used five-pass test-time augmentation, which stabilizedpredictionscores.
From a computational perspective, MobileNetV2_S completed training in about 59.43 minutes and required 0.0805 seconds per test image, whereas EfficientNetB3 required 80.21 minutes and 0.1125 seconds per image. The additional cost of EfficientNetB3 is acceptable for workstation deployment, but MobileNetV2_S remains advantageous when memory and latency constraints are strict.
This work presented a study on skin cancer detection using deep learning with two transfer-learning classifiers, MobileNetV2_S and EfficientNetB3, trained on ISIC2020 dermoscopic images. The implemented pipeline combines data augmentation, class balancing, staged fine-tuning, validation-based threshold selection, and test-time augmentationtoimproverobustness.
Amongtheevaluatedmodels,EfficientNetB3deliveredthe best overall classification performance with 94.72% accuracy, 96.26% specificity, and 0.9860 ROC-AUC, while MobileNetV2_S offered a lighter alternative with lower computational demand and slightly higher melanoma recall. These results show that transfer learning can provide practical support for fast and consistent
melanoma screening from dermoscopic images. Despite these promising results, the present evaluation is limited to binary benign-versus-melanoma classification on a single public dataset and should be complemented by externalclinicalvalidationbeforedeployment.
Future work can extend the present system by incorporatinglesionsegmentation,explainabilitymethods such as Grad-CAM, external clinical validation, and multiclass skin lesion classification. Integration into a secure clinicalinterfaceormobile-assistedworkflowmayfurther improveaccessibilityinunderservedregions.
The authors acknowledge the ISIC archive for providing access to dermoscopic image data and the open-source TensorFlow ecosystem used for model development and evaluation.
[1]R.Kaur,H.GholamHosseini,andM.Linden,"Advanced Deep Learning Models for Melanoma Diagnosis in Computer-Aided Skin Cancer Detection," Sensors, vol. 25, no.3,Art.no.594,2025.
[2] L. Mohankumar, K. Lakshmi Saraswathi, and K. V. Ramana,"SkinCancerDetection byUsing DeepLearning," International Journal of Innovative Research in Technology,vol.11,no.11,pp.4775-4779,2025.
[3] V. Rotemberg et al., "A patient-centric dataset of images and metadata for identifying melanomas using clinicalcontext,"ScientificData,vol.8,Art.no.34,2021.
[4]M.Sandler,A.Howard,M.Zhu,A.Zhmoginov,andL.-C. Chen, "MobileNetV2: Inverted Residuals and Linear Bottlenecks,"inProc.IEEE/CVFConf.Comput.Vis.Pattern Recognit.,2018,pp.4510-4520.
[5] M. Tan and Q. V. Le, "EfficientNet: Rethinking Model Scaling for Convolutional Neural Networks," in Proc. 36th Int. Conf. Mach. Learn. (ICML), PMLR, vol. 97, 2019, pp. 6105-6114.
[6]M. Abadi etal.,"TensorFlow:A Systemfor Large-Scale MachineLearning,"inProc.12thUSENIXSymp.Operating Systems Design and Implementation (OSDI), 2016, pp. 265-283.
[7] B. Cassidy, C. Kendrick, A. Brodzicki, J. JaworekKorjakowska, and M. H. Yap, "Analysis of the ISIC image datasets: Usage, benchmarks and recommendations," MedicalImageAnalysis,vol.75,Art.no.102305,2022.