Peer Reviewed Open Access Journal
ISSN: 3139-3349
This paper presents a comparative analysis of five advanced Natural Language Processing (NLP) models: GPT-4, GPT-3.5, BERT, RoBERTa, and DistilBERT, specifically trained and evaluated for sentiment classification on Twitter data. The study emphasizes the development of these models and assesses their performance using standard metrics, including accuracy, precision, recall, and F1-score. The results indicate that BERT achieved the highest F1-score of 0.8% with a balanced focus on accuracy and efficiency, while DistilBERT delivered a competitive accuracy of 0.86% with significantly reduced inference times. Although GPT-based models excelled in contextual understanding, they exhibited higher latency. These findings highlight the trade-off between predictive accuracy and computational efficiency when deploying AI models for real-time sentiment analysis applications.
Sentiment Analysis, Natural Language Processing, Transformer Models, OpenAI, GPT, Twitter
Chumakov, S., Kovantsev, A., & Surikov, A. (2023). Generative approach to aspect-based sentiment analysis with GPT language models. Procedia Computer Science, 229, 284–293. https://doi.org/10.1016/j.procs.2023.12.030
Fei, Hao, Bobo Li, Qian Liu, Lidong Bing, Fei Li, and Tat-Seng Chua. "Reasoning implicit sentiment with chain-of-thought prompting." arXiv preprint arXiv:2305.11255 (2023).
Gowda, K. P., Porwal, R., Ramesh, C., Tiwari, S. S., Srivastava, K., Rambabu, R., & Govinda Rao, S. (2025). Transformers in sentiment analysis: A paradigm shift in deep learning research.Journal of Information Systems Engineering and Management, 10(5s), 262–280.https://doi.org/10.52783/jisem.v10i5s.612:contentReference[oaicite:13]{index=13}
https://doi.org/10.1109/SMART50582.2020.9337081
Kang, J.-W., & Choi, S.-Y. (2025). Comparative investigation of GPT and FinBERT's sentiment analysis performance in news across different sectors. Electronics, 14(6), 1090. https://doi.org/10.3390/electronics14061090
Kong, Y., Xu, Z., & Mei, M. (2023). Cross-domain sentiment analysis based on feature projection and multi-source attention in IoT. Sensors, 23(16), 7282. https://doi.org/10.3390/s23167282 Wu, Q., Xia, C., & Tian, S. (2025). AI-driven sentiment analytics: Unlocking business value in the e-commerce landscape. arXiv. https://arxiv.org/abs/2504.08738
Lin, Y., & Liu, T. (2024). Impact of NLP Algorithms on Sentiment Analysis Efficiency and Accuracy. Journal of Information Systems and Informatics.
Liu, Y., et al. (2022). RoBERTa: A Robustly Optimized BERT Pretraining Approach. arXiv preprint arXiv:1907.11692.
Mathebula, M., Modupe, A., &Marivate, V. (2024). Fine-tuning retrieval-augmented generation with an auto-regressive language model for sentiment analysis in financial reviews. Applied Sciences, 14(23), 10782. https://doi.org/10.3390/app142310782
Miranda, C. H., Sanchez-Torres, G., &Salcedo, D. (2023). Exploring the evolution of sentiment in Spanish pandemic tweets: A data analysis based on a fine-tuned BERT architecture. Data, 8(6), 96.
Pak, A., &Paroubek, P. (2010, May). Twitter as a corpus for sentiment analysis and opinion mining. In LREc (Vol. 10, No. 2010, pp. 1320-1326).
Perikos, I., & Diamantopoulos, A. (2024). Explainable Aspect-Based Sentiment Analysis Using Transformer Models. Big Data and Cognitive Computing, 8(11), 141. https://doi.org/10.3390/bdcc8110141
Prova, N. (2025). Multilingual Emotion Classification in E-Commerce Customer Reviews Using GPT and Deep Learning-Based Meta-Ensemble Model. Available at SSRN 5161505.
Qin, L., Chen, Q., Wei, F., Huang, S., & Che, W. (2023). Cross-lingual prompting: Improving zero-shot chain-of-thought reasoning across languages. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing (EMNLP). https://aclanthology.org/2023.emnlp-main.163/
Reddy, S., Kumar, A., & Sharma, P. (2025). Leveraging sentiment analysis in the digital era: Uncovering insights from unstructured data for enhanced customer engagement. Journal of Modern Technology, 21(3), 212–219Fei, H., Li, B., Liu, Q., Bing, L., Li, F., &Chua, T.-S. (2023). Reasoning implicit sentiment with chain-of-thought prompting. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing (EMNLP). https://arxiv.org/abs/2305.11255
Semary, N. A., Ahmed, W., Amin, K., Pławiak, P., & Hammad, M. (2023). Improving sentiment classification using a RoBERTa-based hybrid model. Frontiers in human neuroscience, 17, 1292010.
Shafikuzzaman, M., Islam, M. R., Rolli, A. C., Akhter, S., &Seliya, N. (2025). An empirical evaluation of the zero-shot, few-shot, and traditional fine-tuning based pretrained language models for sentiment analysis in software engineering. IEEE Access. https://doi.org/10.1109/ACCESS.2024.3439450
Tan, K. L., Lee, C. P., & Lim, K. M. (2023). A survey of sentiment analysis: Approaches, datasets, and future research. Applied Sciences, 13(7), 4550.
Turpin, M., Michael, J., Perez, E., & Bowman, S. R. (2023). Language models don't always say what they think: Unfaithful explanations in chain-of-thought prompting. Proceedings of the 2023 Conference on Neural Information Processing Systems (NeurIPS). https://arxiv.org/abs/2305.04388
Vamvourellis, D., & Mehta, D. (2025). Reasoning or overthinking: Evaluating large language models on financial sentiment analysis. arXiv. https://arxiv.org/abs/2506.04574
Widyananda, W., Maskur, & Fauzi, A. (2025). Machine Learning and Transformer-based Model for Sentiment Analysis of Indonesian E-Commerce Reviews. The Indonesian Journal of Computer Science, 14(4). (ai-soc.org)
Wiguna, B. S., Hudiyanti, C. V., Rausanfita, A., & Arifin, A. Z. (2021). Sarcasm Detection Engine for Twitter Sentiment Analysis using Textual and Emoji Features. Jurnal Ilmu Komputer dan Informasi, 14(1). (jiki.cs.ui.ac.id)
Yunitasari, Y., Musdholifah, A., & Sari, A. K. (2019). Sarcasm Detection For Sentiment Analysis in Indonesian Tweets. Indonesian Journal of Computing and Cybernetics Systems, 13(1). (Jurnal Universitas Gadjah Mada)
