Ensemble Machine Learning Model for Classification of Spam Product Reviews

Complexity 2020:1-10 (2020)
  Copy   BIBTEX

Abstract

Nowadays, online product reviews have been at the heart of the product assessment process for a company and its customers. They give feedback to a company on improving product quality, planning, and monitoring its business schemes in order to increase sale and gain more profit. They are also helpful for customers to select the right products in less effort and time. Most companies make spam reviews of products in order to increase the products sales and gain more profit. Detecting spam product reviews is a challenging issue in NLP. Numerous machine learning approaches have attempted to detect and classify the product reviews as spam or nonspam. However, in order to improve the classification accuracy, this study has introduced an ensemble machine learning model that combines predictions from multilayer perceptron, K-Nearest Neighbour, and Random Forest and predicts the outcome of the review as spam or real, based on the majority vote of the contributing models. In order to accomplish the task of spam review classification, the proposed ensemble and other benchmark boosting approaches are tested with 25 statistical features extracted from mobile application reviews of Yelp Dataset. Then, three different selection techniques are exploited to diminish the feature space and filter out the top 10 optimal features. The effectiveness of the proposed ensemble, the individual models, and other benchmark boosting approaches is again evaluated with 10 optimal features in terms of classification accuracy. Experimental outcomes illustrate that the proposed ensemble model outperformed the individual classifiers and state-of-the-art boosting approaches like Generalized Boost Regression Model, Extreme Gradient Boost, and AdaBoost Regression Model in terms of classification accuracy.

Links

PhilArchive



    Upload a copy of this work     Papers currently archived: 91,202

External links

Setup an account with your affiliations in order to access resources via your University's proxy server

Through your library

Similar books and articles

Machine Learning and Job Posting Classification: A Comparative Study.Ibrahim M. Nasser & Amjad H. Alzaanin - 2020 - International Journal of Engineering and Information Systems (IJEAIS) 4 (9):06-14.
The ethical status of non-commercial spam.Emma Rooksby - 2007 - Ethics and Information Technology 9 (2):141-152.

Analytics

Added to PP
2020-12-22

Downloads
52 (#292,437)

6 months
42 (#88,821)

Historical graph of downloads
How can I increase my downloads?

Citations of this work

No citations found.

Add more citations

References found in this work

No references found.

Add more references