Open Access
Tweets Classification Using Corpus Dependent Tags, Character and POS N-grams뀀Ƞ뀀Ƞ Notebook for PAN at CLEF 2015
Carlos E. González-Gallardo,Azucena Montes,Gerardo Sierra,J. Antonio Núñez-Juárez,Adolfo Jonathan Salinas-López,Juan Ek +5 more
- 01 Jan 2015
TL;DR: This paper takes into account stylistic features represented by character N- grams and POS N-grams to classify tweets and obtains results that are very satisfactory.
read more
Abstract: This paper is part of the Author Profiling task at PAN 2015 contest; in witch participants had to predict the gender, age and personality traits of Twitter users in four different languages (Spanish, English, Italian and Dutch). Our approach takes into account stylistic features represented by character N- grams and POS N-grams to classify tweets. The main idea of using character N- grams is to extract as much information as possible that is encoded inside the tweet (emoticons, character flooding, use of capital letters, etc.). POS N-grams were obtained using Freeling and certain token were relabeled with Twitter de- pendent tags. Obtained results were very satisfactory; our global ranking score was of 83.46%.
read more
Chat with Paper
AI Agents for this Paper
Find similar papers on Google Scholar, PubMed and Arxiv
Write a critical review of this paper
Analyze citations of this paper to find unaddressed research gaps
Citations
What demographic attributes do our digital footprints reveal? A systematic review
Joanne Hinds,Adam Joinson +1 more
TL;DR: A systematic review that synthesises current evidence on predicting demographic attributes from online digital traces and provides a database containing the platforms and digital traces examined, sample sizes, accuracy measures and the classification methods applied.
100
Distance measures in author profiling
Mirco Kocher,Jacques Savoy +1 more
TL;DR: Empirically, the empirical evaluations indicate that the Canberra or Clark distance measures tend to produce better effectiveness than the rest, at least in the context of an author profiling task.
68
A Language-independent and Compositional Model for Personality Trait Recognition from Short Texts
Fei Liu,Julien Perez,Scott Nowson +2 more
- 01 Apr 2017
TL;DR: This work proposes to use deep-learning-based models with atomic features of text – the characters – to build hierarchical, vectorial word and sentence representations for the task of trait inference, and shows state-of-the-art performance on a corpus of tweets.
•Proceedings Article
Building a Corpus for Personality-dependent Natural Language Understanding and Generation.
Ricelli Moreira Silva Ramos,Georges Basile Stavracas Neto,Barbara Barbosa Claudino Silva,Danielle Sampaio Monteiro,Ivandré Paraboni,Rafael Dias +5 more
- 01 May 2018
TL;DR: The b5 corpus is described, a collection of controlled and free (non-topic specific) texts produced in different communicative tasks, and accompanied by inventories of personality of their authors and additional demographics, which aims to provide support for a wide range of NLP studies based on personality information.
34
Multilingual author profiling using word embedding averages and SVMs
Roy Khristopher Bayot,Teresa Gonçalves +1 more
- 01 Jan 2016
TL;DR: An experiment done to investigate author profiling of tweets in English and Spanish, particularly for cross genre evaluation shows that using average of word vectors outperforms tfidf in most cross genre problems for age and gender.
28
References
•Journal Article
Scikit-learn: Machine Learning in Python
Fabian Pedregosa,Gaël Varoquaux,Alexandre Gramfort,Vincent Michel,Bertrand Thirion,Olivier Grisel,Mathieu Blondel,Peter Prettenhofer,Ron Weiss,Vincent Dubourg,Jake Vanderplas,Alexandre Passos,David Cournapeau,Matthieu Brucher,Matthieu Perrot,Edouard Duchesnay +15 more
TL;DR: Scikit-learn is a Python module integrating a wide range of state-of-the-art machine learning algorithms for medium-scale supervised and unsupervised problems, focusing on bringing machine learning to non-specialists using a general-purpose high-level language.
Comparative study on methods of detecting research fronts using different types of citation
TL;DR: Direct citation, which could detect large and young emerging clusters earlier, shows the best performance in detecting a research front, and co-citation shows the worst, which suggests that the content similarity of papers connected by direct citations is the greatest and that direct citation networks have the least risk of missing emerging research domains.
1.7K
•Proceedings Article
Authorship Attribution of Micro-Messages
Roy Schwartz,Oren Tsur,Ari Rappoport,Moshe Koppel +3 more
- 01 Oct 2013
TL;DR: The concept of an author’s unique “signature” is introduced, and it is shown that such signatures are typical of many authors when writing very short texts.
120
Overview of the 2nd Author Profiling Task at PAN 2014
Francisco Rangel,Paolo Rosso,Moshe Koppel,Efstathios Stamatatos,Giacomo Inches +4 more
- 01 Jan 2013
TL;DR: The framework and results for the Author Pro- filing task at PAN 2013 are presented and the evaluation framework used to measure the participants performance to solve the problem of identifying age and gender from anonymous texts is described.