Back to Search Start Over

Feature extraction using LR-PCA hybridization on twitter data and classification accuracy using machine learning algorithms.

Authors :
Murugan, N. Senthil
Devi, G. Usha
Source :
Cluster Computing; Nov2019 Supplement 6, Vol. 22, p13965-13974, 10p
Publication Year :
2019

Abstract

Twitter, a social blogging site which became the tremendous topic in today's environment, which made several organizations and public to develop their identity and overwhelming through this social website. But unfortunately, twitter facing great challenges due to spammers who break the reputation of the website from deliberate users to stop using it. Researchers have proposed many techniques to overcome the issues faced by the spammers. As far researchers find a new path so as the spammers develop new techniques to travel in that path. So far, many algorithms were proposed to detect the spammers and some extraction techniques have developed to increase the potential of detection rate. In this paper, the main focus is about feature extraction of our data with a hybrid approach of combining logistic regression with dimensional reduction technique using principal component analysis. Our dataset contains 17 million users' tweets with 159 features included in it. Then we are going to extract particular features from it which would be helpful for the further process of increasing the classification accuracy. For the classification process, our work extended for the process of classification of data using some machine learning techniques. From the proposed work the detection rate could be increased by using particular features for the classification process. [ABSTRACT FROM AUTHOR]

Details

Language :
English
ISSN :
13867857
Volume :
22
Database :
Complementary Index
Journal :
Cluster Computing
Publication Type :
Academic Journal
Accession number :
139866527
Full Text :
https://doi.org/10.1007/s10586-018-2158-3