Skip to main content

PHISHING URLS DETECTION USING MACHINE LEARNING AND FLASH FRAME WORK

Page 1

International Research Journal of Engineering and Technology (IRJET)

e-ISSN: 2395-0056

Volume: 12 Issue: 04 | Apr 2025

p-ISSN: 2395-0072

www.irjet.net

PHISHING URLS DETECTION USING MACHINE LEARNING AND FLASH FRAME WORK Mr. SANTHOSH M1, Mr. RAJADURAI2 1Mr. SANTHOSH M, M.sc CFIS, Department of Computer Science Engineering,

Dr. MGR UNIVERSITY, Chennai, India

2Mr. RAJADURAI, Assistant Professor, Center of Excellence in Digital Forensics, Chennai, India

---------------------------------------------------------------------***--------------------------------------------------------------------Abstract - The rapid advancement of Artificial Intelligence 1. INTRODUCTION (AI) has significantly propelled the growth of the Internet of Things (IoT). However, as this technology becomes more integrated with internet connectivity, it also faces heightened cybersecurity risks—particularly from malicious websites. Detecting these threats is crucial, and machine learning algorithms have shown strong potential in identifying anomalous patterns within large volumes of network traffic. In this project, we leverage several machine learning models— such as Random Forest, Support Vector Machine (SVM), Decision Tree, Extra Trees Classifier, K-Nearest Neighbors (kNN), XGBoost, CatBoost, Multilayer Perceptron (MLP), and Gradient Boosting—to effectively detect and classify malicious URLs.

The increasing reliance on the Internet has led to a surge in cyber-attacks, with malicious URLs being a primary method for phishing, malware, and spam attacks. These malicious URLs compromise data security by threatening the confidentiality, integrity, and availability of sensitive information. Traditional methods of detecting malicious websites rely heavily on manually defined rules and thresholds, which are often subjective and inflexible. As a result, these systems struggle to keep up with evolving threats and fail to detect new malicious URLs[4]. One of the key limitations of traditional detection methods is the reliance on blacklists, which record known malicious URLs. However, blacklists cannot identify newly generated malicious URLs in real-time, leaving users vulnerable to attacks. Over 90% of malicious links are clicked before they even appear on blacklists. Additionally, the maintenance of these lists depends on human feedback, making them both labor-intensive and prone to delays in identifying new threats[5]. As the volume of data and the frequency of attacks continue to rise, traditional methods become increasingly ineffective.

A significant aspect of our methodology involves robust feature engineering, as the success of a machine learning model is deeply rooted in the quality of its input features. To enhance performance further, we propose an unsupervised learning approach that learns URL embeddings. Additionally, we developed a web application using the Flask framework to detect and flag potentially harmful URLs in real time. Malicious URLs continue to be a dominant cyber threat vector, commonly used in phishing attacks, malware distribution, and spam campaigns. While traditional blacklist-based methods are effective for previously known threats, they often fall short in identifying newly generated malicious URLs [1]. To overcome this limitation, our study introduces a machine learning-based detection system using URL-based features in a multiclass classification setup. We focus specifically on three prevalent attack types: phishing, spam, and malware [2].

In contrast, machine learning (ML) techniques offer a more adaptive and scalable approach. Algorithms such as Decision Trees (DT), Support Vector Machines (SVM), Extra Trees Classifiers (ETC), and Random Forest (RF) can automatically learn patterns from large datasets and identify malicious URLs without human intervention. These models are able to detect both known and previously unseen threats, improving detection accuracy and efficiency[6].

To evaluate performance, we compared four widely-used ensemble learning algorithms: Extreme Gradient Boosting (XGBoost), Adaptive Boosting (AdaBoost), Light Gradient Boosting (LightGBM), and Categorical Boosting (CatBoost). Our results demonstrate the effectiveness of these models in identifying threats based on engineered URL characteristics, offering a valuable supplement to existing anti-phishing, antispam, and anti-malware systems [3].

This paper explores the use of machine learning for detecting malicious URLs, focusing on phishing, spam, and malware attacks. By leveraging ensemble learning methods like XGBoost, AdaBoost, LightGBM, and CatBoost, we aim to enhance the performance and scalability of malicious URL detection systems[7]. Our goal is to provide a more reliable and adaptive solution to combat the growing challenges posed by cyber threats, ultimately improving cybersecurity practices across various sectors[8].

Key Words: Malicious URLs, Machine Learning, Detection, Phishing, Spam, Malware, Ensemble Learning

© 2025, IRJET

|

Impact Factor value: 8.315

|

ISO 9001:2008 Certified Journal

|

Page 1429


Turn static files into dynamic content formats.

Create a flipbook
PHISHING URLS DETECTION USING MACHINE LEARNING AND FLASH FRAME WORK by IRJET Journal - Issuu