Skip to main content

Evaluation Metrics for Sentiment Analysis: A Comprehensive Review and Future Directions

Page 1

International Research Journal of Engineering and Technology (IRJET)

e-ISSN: 2395-0056

Volume: 12 Issue: 04 | Apr 2025

p-ISSN: 2395-0072

www.irjet.net

Evaluation Metrics for Sentiment Analysis: A Comprehensive Review and Future Directions Palak Kaushal1, Anita Ganpati2 1 Palak Kaushal, Himachal Pradesh University, Shimla, India 2Anita Ganpati, Himachal Pradesh University, Shimla, India

---------------------------------------------------------------------***---------------------------------------------------------------------

Abstract - Evaluation metrics are crucial for assessing the performance and reliability of sentiment analysis models in various

applications. Evaluation metrics are critical for appraising sentiment analysis models performance and guaranteeing their dependability in various applications. The research thoroughly examined the classification, regression, ranking, and explainability metrics. Every measure has advantages and disadvantages that affect how well sentiment categorisation and forecasting assignments work. These measures are compared, providing insight into their effectiveness in various sentiment analysis contexts. Future studies should concentrate on fairness-driven and context-aware assessment methods to expand the reliability and interpretability of classification models.

Key Words: Sentiment Analysis, Evaluation Metrics, Classification Metrics, Regression Metrics, Ranking Metrics. 1. INTRODUCTION Sentiment Analysis, a branch of Natural Language Processing (NLP)[1], aims to construe and assess emotions, opinions, and attitudes conveyed in textual data[2]. Sentiment analysis is becoming a crucial tool in many fields, such as business intelligence, consumer feedback analysis, and social media monitoring, due to the explosive expansion of digital material on social media, ecommerce platforms, and online reviews. Businesses utilise sentiment analysis to gain insight into public opinion, make better decisions, and improve user experience. Assessing the usefulness and dependability of sentiment categorisation models is a crucial component of sentiment analysis. A model's performance in numerous tasks is resolute by the assessment measures it uses. Different criteria for evaluation are needed depending on whether the task requires classification, regression, or ranking. Moreover, explainability and fairness measures have become more crucial for ensuring transparency and objective judgements because advanced learning and black-box models are used more often in sentiment analysis. Classification metrics, regression metrics, ranking metrics, and explainability/fairness measurements are the four primary categories into which this study divides sentiment analysis assessment metrics. This study intends to assist researchers in choosing suitable assessment techniques depending on the nature of their sentiment analysis jobs by offering an organised summary of various measures.

2. LITERATURE REVIEW Sentiment analysis model assessment has been deeply studied and several measures have been put out to evaluate performance on tasks involving classification, regression, and ranking. The important research that has influenced the creation of sentiment analysis assessment metrics is reviewed in this section. Most popular sentiment analysis responsibility is sentiment investigation, which is classifying text into predetermined sentiment categories. Using accuracy as the core assessment criterion, [2]used conventional learning classifiers, such as Naive Bayes, Support Vector Machines (SVM), and Maximum Entropy. Precision, recall, and F1-score have been adopted as more informative metrics because of the criticism of accuracy's shortcomings in unbalanced datasets [3]. Deep learning-based sentiment classification has further emphasized the need for robust evaluation metrics. Long ShortTerm Memory (LSTMs) networks, introduced by[4]established strong performance in capturing contextual dependencies in sentiment classification. The paper[5] evaluates sentiment classification using accuracy, precision, recall, F1-score, and ROC-AUC. The ensemble model outperformed individual classifiers, achieving high F1-score and AUC, reducing misclassification, and improving sentiment detection in Arabic social media text. The paper[6] evaluates sentiment analysis models using accuracy, precision, recall, and F1-score to compare their effectiveness in recognizing emotional content. The results highlight that deep learning models beat traditional processes, achieving higher precision and recall, making them more suitable for sentiment classification tasks.

© 2025, IRJET

|

Impact Factor value: 8.315

|

ISO 9001:2008 Certified Journal

|

Page 876


Turn static files into dynamic content formats.

Create a flipbook
Evaluation Metrics for Sentiment Analysis: A Comprehensive Review and Future Directions by IRJET Journal - Issuu