EVALUATING CUSTOMER EXPERIENCE IN E-COMMERCE: MULTILINGUAL SENTIMENT ANALYSIS OF USER REVIEWS USING TRANSFORMER MODELS
DOI:
https://doi.org/10.60022/sis.2.(02).6Keywords:
Sentiment analysis, Natural Language Processing, e-commerce, customer experience, multilingual transformer models, XLM-RoBERTa, class imbalanceAbstract
This study addresses the challenge of multilingual sentiment analysis in e-commerce, with a focus on Ukrainian and Russian book reviews. We propose a hybrid framework based on transformer architectures that accounts for the linguistic complexity of Slavic languages and the significant class imbalance often present in customer feedback data. Using a dataset of approximately 70,000 user reviews from the Ukrainian online bookstore Yakaboo, we evaluate four model variants based on the XLM-RoBERTa architecture and compare their performance to a monolingual Ukrainian RoBERTa baseline.
The classification task is formulated as binary: identifying low-rated reviews (1–3 stars) versus high-rated ones (4–5 stars), with the minority class comprising only 4–5% of the dataset. Our best-performing model — combining partial fine-tuning, a deep classifier, and focal loss — achieves a macro F1 score of 0.73 and an F1 score of 0.47 for the minority class, outperforming the baseline (0.66 and 0.37, respectively). The application of oversampling and focal loss proved effective in mitigating class imbalance.
Beyond technical performance, the findings underline the practical utility of accurate sentiment detection for e-commerce platforms. Effective identification of negative feedback enables companies to address customer dis- satisfaction, refine product offerings, and inform personalized marketing strategies. This research contributes to advancing multilingual NLP in low-resource settings and provides a scalable solution for real-world sentiment classification in morphologically rich languages.
The scientific novelty of this research lies in the integration of multilingual transformer architectures with adaptive classifiers and class imbalance mitigation techniques, specifically tailored for real-world e-commerce review data in morphologically rich languages. The proposed framework demonstrates how domain-specific fine-tuning and architectural customization can significantly enhance sentiment classification performance in low-resource language settings.
References
Statista. (2024). E-commerce worldwide — Statista Digital Market Outlook 2024. https://statista.com/outlook/dmo/ecommerce/worldwide.
Turlakova, S.S., & Shumilo, Ya.M. (2025). Infl uence of AI Tools on Consumer Behavior Ma-nagement in Digital Marketing. Sci. innov., 21(1), 67–81. https://doi.org/10.15407/scine21.01.067.
Oleksiuk, O., & Shafalyuk, O. (2023). Management of pharmaceutical online retail through a regional market- place with neural network and statistical analytical tools. Neuro-Fuzzy Modeling Techniques in Economics, 12, 155–174. http://doi.org/10.33111/nfmte.2023.155.
Bigne, E., Ruiz, C., Perez-Cabañero, C., et al. (2023). Are customer star ratings and sentiments aligned? A deep learning study of the customer service experience in tourism destinations. Service Business, 17, 281–314. https://doi. org/10.1007/s11628–023–00524–0.
Prytula, M. (2024). Fine-tuning BERT, DistilBERT, XLM-RoBERTa and Ukr-RoBERTa models for sentiment anal- ysis of ukrainian language reviews. Artificial Intelligence, 2, 85–98. https://doi.org/10.15407/jai2024.02.085.
Al-Natour, S., & Turetken, O. (2020). A comparative assessment of sentiment analysis and star ratings for con- sumer reviews. International Journal of Information Management, 54, Article 102132. https://doi.org/10.1016/j.ijinfo- mgt.2020.102132.
Devlin, J., Chang, M. W., Lee, K., & Toutanova, K. (2019). BERT: Pre-training of deep bidirectional transformers for language understanding. Proceedings of the 2019 Conference of the North American Chapter of the Association for Compu- tational Linguistics: Human Language Technologies, 4171–4186. https://doi.org/10.18653/v1/N19–1423.
Conneau, A., Khandelwal, K., Goyal, N., Chaudhary, V., Wenzek, G., Guzmán, F., Grave, E., Ott, M., Zettlemoyer, L., & Stoyanov, V. (2020). Unsupervised cross-lingual representation learning at scale. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, 8440–8451. https://doi.org/10.18653/v1/2020.acl-main.747.
Huang, H., Asemi Zavareh, A., & Mustafa, M. B. (2023). Sentiment Analysis in E-Commerce Platforms: A Review of Current Techniques and Future Directions. IEEE Access, 11, 93141–93162. https://doi.org/10.1109/ACCESS.2023.3307308. [10] Agüero-Torales, M.M., Abreu Salas, J.I., & López-Herrera, A.G. (2021). Deep learning and multilingual senti- ment analysis on social media data: An overview. Applied Soft Computing, 107, Article 107373. https://doi.org/10.1016/j.
asoc.2021.107373.
Derbentsev, V.D., Bezkorovainyi, V.S., Matviychuk, A.V., Pomazun, O.M., Hrabariev, A.V., & Hostryk, A.M.
(2023). A comparative study of deep learning models for sentiment analysis of social media texts. CEUR Workshop Pro- ceedings, 3465, 168–188. https://ceur-ws.org/Vol-3465/paper18.pdf.
Mabokela, K. R., Celik, T., & Raborife, M. (2024). Multilingual Sentiment Analysis for Under-Resourced Languag- es: A Systematic Review of the Landscape. IEEE Access, 12, 13100–13122. https://doi.org/10.1109/ACCESS.2022.3224136. [13] Ge, J., Alonso-Vazquez, M., & Gretzel, U. (2018). Sentiment analysis: a review. Advances in social media for travel,
tourism and hospitality: new perspectives, practice and cases. 243–261. https://doi.org/10.4324/9781315565736.
Pan, Y., & Zhang, J.Q. (2009). Is a “star” worth a thousand words? The interplay between product-review texts
and rating valences. European Journal of Marketing, 43(11/12), 1269–1280. https://doi.org/10.1108/03090560910989876. [15] Cho, H., Hasija, S., & Sosa, M. (2021). Reading Between the Stars: Understanding the Effects of Online Customer
Reviews on Product Demand. INSEAD Working Paper No. 2021/07/TOM. https://doi.org/10.2139/ssrn.3240453.
Altab, H. M., Yinping, M., & Adu-Yeboah, S. S. (2022). Understanding Online Consumer Textual Reviews and Rat- ing: Review Length With Moderated Multiple Regression Analysis Approach. SAGE Open, 12(2). https://doi.org/10.1177/
Thakkar, G., Mikelić Preradović, N., & Tadić, M. (2024). Examining Sentiment Analysis for Low-Resource Lan- guages with Data Augmentation Techniques. Eng, 5(4), 2920–2942. https://doi.org/10.3390/eng5040152.
Aliyu, Y., Sarlan, A., Danyaro, K.U., B.A. Rahman, A.S., & Abdullahi, M. (2024). Sentiment Analysis in Low- Resource Settings: A Comprehensive Review of Approaches, Languages, and Data Sources. IEEE Access, 12, 10300–10323. https://doi.org/10.1109/ACCESS.2024.3398635.
Barbieri, F., Espinosa Anke, L., & Camacho-Collados, J. (2022). XLM-T: Multilingual Language Models in Twitter for Sentiment Analysis and Beyond. In Proceedings of the 13th Conference on Language Resources and Evaluation, LREC 2022, 258–266. https://aclanthology.org/2022.lrec-1.27
Manias, G., Mavrogiorgou, A., Kiourtis, A., Symvoulidis, C., et al. (2023). Multilingual text categorization and sen- timent analysis: a comparative analysis of the utilization of multilingual approaches for classifying twitter data. Neural Computing and Applications, 35(29), 1–17. https://doi.org/10.1007/s00521-023-08629-3.
Miah, M.S.U., Kabir, M.M., Sarwar, T.B., Safran, M., Alfarhood, S., & Mridha, M. F. (2024). A multimodal ap- proach to cross-lingual sentiment analysis with ensemble of transformer and LLM. Scientific Reports, 14, 19603. https://doi. org/10.1038/s41598-024-60210-7.
Desai, T., & Meva, D. (2023). Comparative Analysis of Different Machine Learning Approaches for Sentiment Analysis. In H. Sharma, V. Shrivastava, K.K. Bharti, & L. Wang (Eds.), Lecture Notes in Networks and Systems: Vol. 686. Communication and Intelligent Systems (pp. 175–185). Springer, Singapore. https://doi.org/10.1007/978-981-99-2100-3_15.
Matviychuk, A., Derbentsev, V., Bezkorovainyi, V., Kmytiuk, T., & Hostryk, A. (2024). Leveraging Artificial In- telligence and Large Language Models for Fake Content Detection in Digital Media. CEUR Workshop Proceedings, 3933, 75–92. https://ceur-ws.org/Vol-3933/Paper_7.pdf.
Akhmedov, R.R., Bezkorovainyi, V.S., & Danylchenko, T.V. (2021). Methodology of content analysis of electronic mass media. Ekonomichnyi Prostir (Economic Scope), 176, 141–145. https://doi.org/10.32782/2224-6282/176-25.
Derbentsev, V., Bezkorovainyi, V., & Akhmedov, R. (2020). Machine learning approach of analysis of emotional polarity of electronic social media. Neuro-Fuzzy Modeling Techniques in Economics, 9, 95–137. http://doi.org/10.33111/ nfmte.2020.095.
Lemaitre, G., Nogueira, F., & Aridas, C.K. (2017). Imbalanced-learn: A Python toolbox to tackle the curse of im- balanced datasets in machine learning. Journal of Machine Learning Research, 18(1), 559–563. https://doi.org/10.48550/ arXiv.1609.06570.
He, H., & Garcia, E. A. (2009). Learning from imbalanced data. IEEE Transactions on Knowledge and Data Engi- neering, 21(9), 1263–1284. https://doi.org/10.1109/TKDE.2008.239.
Chawla, N.V., Bowyer, K.W., Hall, L.O., & Kegelmeyer, W.P. (2002). SMOTE: Synthetic Minority Over-sampling Technique. Journal of Artificial Intelligence Research, 16, 321–357. https://doi.org/10.1613/jair.953.
Firza, N., Bakiu, A., & Monaco, A. (2025). Machine Learning for Quality Diagnostics: Insights into Consumer Electronics Evaluation. Electronics, 14(5), Article 939. https://doi.org/10.3390/electronics14050939.
Mao, Y., Liu, Q., & Zhang, Y. (2024). Sentiment analysis methods, applications, and challenges: A systematic lit- erature review. Journal of King Saud University — Computer and Information Sciences, 36(4), Article 102048. https://doi. org/10.1016/j.jksuci.2024.102048.
Zheng, L., Wang, H., & Gao, S. (2018). Sentimental feature selection for sentiment analysis of Chinese online re- views. International Journal of Machine Learning and Cybernetics, 9(1), 75–84. https://doi.org/10.1007/s13042-015-0347-4.
Radchenko, V. (2020). ukr-roberta-base [Model]. Hugging Face. https://huggingface.co/youscan/ukr-roberta-base
Yakaboo Book Reviews. (2020) Curated list of Ukrainian natural language processing (NLP) resources (corpora, pretrained models, libraries, etc.). https://github.com/osyvokon/awesome-ukrainian-nlp
Kingma, D. P., & Ba, J. (2015). Adam: A Method for Stochastic Optimization. International Conference on Learning Representations (ICLR). https://doi.org/10.48550/arXiv.1412.6980.
Lin, T.-Y., Goyal, P., Girshick, R., He, K., & Dollár, P. (2017). Focal Loss for Dense Object Detection. In Proceedings of the IEEE International Conference on Computer Vision (pp. 2980–2988). IEEE. https://doi.org/10.1109/ICCV.2017.324.
Hugging Face Team. (2023). Transformers [Software library]. Hugging Face. https://huggingface.co/transformers/
Downloads
Published
How to Cite
Issue
Section
License
Copyright (c) 2024 Василь Дербенцев, Віталій Безкоровайний, Ренат Ахмедов, Микола Бондарчук

This work is licensed under a Creative Commons Attribution 4.0 International License.