MARATTO

conference paper · 2022 3rd International Conference on Intelligent Engineering and Management (ICIEM)

The Sentiment Analysis of Human Behavior on Products and Organizations using K-Means Clustering and SVM Classifier

202217 citationsBeni Suef University

In plain language

This research examines sentiment analysis and opinion mining across multiple languages, including English, French, and Dutch. The study evaluates text from sources such as reviews and blogs to assess consumer sentiment regarding specific consumption products. Using manually labelled sentences categorised into positive, negative, and neutral classes, the work addresses text issues including noise. Unigram features enriched with linguistic information achieved roughly 83 percent accuracy when classifying English texts. To categorise sentiment from text, the approach applies K-means clustering alongside a support vector machine classifier. The investigation also explores active learning techniques to reduce the volume of data requiring manual annotation, while providing experimental data on how well learned models transfer across different subject domains and languages.

Key takeaways

  • Machine learning models were tested on multilingual texts, specifically Dutch, English, and French.
  • Unigram features combined with linguistic information reached approximately 83 percent accuracy for English text classification.
  • K-means clustering and a support vector machine classifier were used to identify individual sentiment from text.
  • Active learning methods helped minimise the manual annotation needed for training data.
  • The study generated data regarding the transferability of trained sentiment models across different domains and languages.

Why it matters

Understanding customer opinions across diverse languages and informal text sources like blogs is vital for evaluating how organisations and products are perceived. Developing automated systems that require less manual data labelling and can adapt across languages helps organisations analyse large volumes of consumer feedback more efficiently.

Commercialisation angle

The methodology could enable automated opinion mining tools for businesses tracking consumer responses to products and brand reputation. Potential users include market research teams and customer experience analysts operating in multilingual environments. Given that the work focuses on experimental validation and feature testing achieving 83 percent accuracy in English, the technology represents early-stage applied research that requires further operational development before direct market deployment.

AI-generated from the published abstract. Always read the original work before citing.

Abstract

Sentiment analysis Sentiment analysis is the process of extracting information from the text and it is considered as opinion mining. Machine learning experiments have been shown in the study that involves review, blog, for some texts that are written in different languages including Dutch, English and French. Set of example sentence have been set that are manually labeled, neutral, positive or negative. The study covers the curiosity of the consumer regarding the specific consumption products. Categorization models have been developing that has been used in the study. Number of issues that includes noisy nature of the text has been discussed in the study. With an accuracy of roughly 83 percent for English texts, we can determine positive, negative, and neutral sentiments toward the subject under investigation using unigram features augmented with linguistic information. The role of active learning approaches in minimizing the number of instances that need to be manually annotated is discussed in this article. Our studies also give data on the transferability of learned models between domains and languages. Here, Sentiment analysis of a particular person has studied using a K-Means Clustering and SVM classifier to classify the sentiment of a person from the text

Research topics

  • Sentiment Analysis and Opinion Mining
  • Spam and Phishing Detection
  • Advanced Text Analysis Techniques

Sustainable Development Goals

Read the original research

This page summarises published work. The authoritative version sits with the publisher.

DOI: 10.1109/iciem54221.2022.9853128

Is something wrong with this record? Report it or request removal.

Discussion

Discuss this research

Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.

No discussion yet. Open the first thread.