ELEVATING PHISHING DETECTION PERFORMANCE WITH MACHINE LEARNING AND DEEP LEARNING-ENABLED FEATURE SELECTION
DOI:
https://doi.org/10.64751/ajaccm.2026.v6.n3.751Abstract
Phishing remains one of the most persistent and rapidly evolving cybersecurity threats, exploiting deceptive websites, malicious URLs, fraudulent messages, compromised domains, and social-engineering strategies to obtain sensitive information such as usernames, passwords, financial credentials, personal records, and authentication tokens. Conventional phishing detection mechanisms based on blacklists, manually defined rules, static signatures, and heuristic filters provide useful protection against previously identified attacks but often exhibit limited effectiveness against zero-day phishing websites, short-lived malicious domains, obfuscated URLs, and dynamically changing attack patterns. Furthermore, machine learning-based phishing detection models frequently process large and redundant feature spaces containing irrelevant, correlated, or noisy attributes, which may increase computational overhead and reduce generalization capability. This research proposes an intelligent phishing detection framework that integrates machine learning, deep learning-enabled feature selection, multi-source phishing feature extraction, hybrid classification, and real-time risk assessment. The proposed framework extracts URL lexical characteristics, domain and host-based properties, webpage content indicators, Hypertext Markup Language and JavaScript features, security certificate attributes, redirection behavior, and contextual metadata. A deep learningenabled feature selection module employs representation learning and importance estimation to identify the most discriminative phishing indicators while eliminating redundant and low-contribution attributes. The selected feature subset is subsequently evaluated using machine learning classifiers such as Random Forest, Support Vector Machine, XGBoost, and Logistic Regression, together with deep learning architectures including Multilayer Perceptron, Convolutional Neural Network, and Long Short-Term Memory networks. A hybrid decision engine combines model confidence, anomaly indicators, and contextual risk information to classify web resources as legitimate, suspicious, or phishing. The proposed architecture consists of five interconnected layers: Data Acquisition, Preprocessing and Feature Engineering, Deep Learning-Enabled Feature Selection and Intelligent Detection, Risk Assessment and Response, and Application/User layers. Illustrative conceptual evaluation demonstrates that the proposed hybrid framework can achieve higher detection accuracy, precision, recall, F1-score, and lower response latency than blacklist-based, conventional machine learning, and standalone deep learning approaches. The framework provides a scalable foundation for intelligent phishing protection across browsers, email gateways, enterprise networks, financial platforms, educational environments, and cloud-based security services.
Downloads
Published
Issue
Section
License

This work is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License.







