Publication:
Big data feature selection and projection for gender prediction based on user web behaviour

Loading...
Thumbnail Image

Advisor

Journal Title

Journal ISSN

Volume Title

Publisher

IEEE

Research Projects

Organizational Units

Journal Issue

Abstract

Prediction of a visitors' gender and other demographic information helps with the presentation of the appropriate content to the user. In this paper, we perform gender prediction based on Turkish users' web log data. Our methods use three different sets of features, namely the URLs (Uniform Resource Locator), the textual contents and the DMOZ (from directory.mozilla.org) hierarchies of the pages visited by each user. Since we have a sparse high-dimensional input dataset, first we apply Information Gain and Chi-square based feature selection. We use a MapReduce based approach to compute these feature relevance measures. We also apply stochastic singular value decomposition (SSVD) feature projection method. We present gender classification results, based on these feature selection and projection methods, using the Logistic Regression classifier. Using the Logistic Regression classifier on the selected URL features results in the best performance.

Description

Subject

Citation

Collections

Endorsement

Review

Supplemented By

Referenced By

Related Goal

0

Views

0

Downloads
View PlumX Details