Search results
1 – 1 of 1Mariam Elhussein and Samiha Brahimi
This paper aims to propose a novel way of using textual clustering as a feature selection method. It is applied to identify the most important keywords in the profile…
Abstract
Purpose
This paper aims to propose a novel way of using textual clustering as a feature selection method. It is applied to identify the most important keywords in the profile classification. The method is demonstrated through the problem of sick-leave promoters on Twitter.
Design/methodology/approach
Four machine learning classifiers were used on a total of 35,578 tweets posted on Twitter. The data were manually labeled into two categories: promoter and nonpromoter. Classification performance was compared when the proposed clustering feature selection approach and the standard feature selection were applied.
Findings
Radom forest achieved the highest accuracy of 95.91% higher than similar work compared. Furthermore, using clustering as a feature selection method improved the Sensitivity of the model from 73.83% to 98.79%. Sensitivity (recall) is the most important measure of classifier performance when detecting promoters’ accounts that have spam-like behavior.
Research limitations/implications
The method applied is novel, more testing is needed in other datasets before generalizing its results.
Practical implications
The model applied can be used by Saudi authorities to report on the accounts that sell sick-leaves online.
Originality/value
The research is proposing a new way textual clustering can be used in feature selection.
Details