TY - JOUR
T1 - Author verification using a semantic space model
AU - Hernández-Castañeda, Ángel
AU - Calvo, Hiram
PY - 2017
Y1 - 2017
N2 - In this work we propose to solve the author verification problem using a semantic space model through Latent Dirichlet Allocation (LDA). We experiment with the corpus used in the author identification tasks at PAN 2014 and PAN 2015. These datasets consist of subsets in the following languages: English, Spanish, Dutch and Greek. Each problem contained in these corpora is formed by one to five known documents which were written by one author and one unknown document. The task is to predict whether the unknown document was written by the author who wrote the known documents. We processed the documents in the dataset and captured the fingerprint of authors by generating a probabilistic distribution of words in the documents. In PAN 2015 classification, we achieved 81.6%, 75.4%, 74.1%, 67.1%accuracy for each English, Spanish, Dutch and Greek subset respectively. In particular for the English subset, we outreached the best result reported in both competitions.
AB - In this work we propose to solve the author verification problem using a semantic space model through Latent Dirichlet Allocation (LDA). We experiment with the corpus used in the author identification tasks at PAN 2014 and PAN 2015. These datasets consist of subsets in the following languages: English, Spanish, Dutch and Greek. Each problem contained in these corpora is formed by one to five known documents which were written by one author and one unknown document. The task is to predict whether the unknown document was written by the author who wrote the known documents. We processed the documents in the dataset and captured the fingerprint of authors by generating a probabilistic distribution of words in the documents. In PAN 2015 classification, we achieved 81.6%, 75.4%, 74.1%, 67.1%accuracy for each English, Spanish, Dutch and Greek subset respectively. In particular for the English subset, we outreached the best result reported in both competitions.
KW - Author verification
KW - Cross-genre
KW - Cross-topic
KW - Latent dirichlet allocation
KW - Semantic space model
UR - http://www.scopus.com/inward/record.url?scp=85021806484&partnerID=8YFLogxK
U2 - 10.13053/CyS-21-2-2732
DO - 10.13053/CyS-21-2-2732
M3 - Artículo
SN - 1405-5546
VL - 21
SP - 167
EP - 179
JO - Computacion y Sistemas
JF - Computacion y Sistemas
IS - 2
ER -