A Novel Method for Unsupervised and Supervised Conversational Message Thread Detection

Giacomo Domeniconi, Konstantinos Semertzidis, Vanessa Lopez, Elizabeth M. Daly, Spyros Kotoulas, Gianluca Moro

2016

Abstract

Efficiently detecting conversation threads from a pool of messages, such as social network chats, emails, comments to posts, news etc., is relevant for various applications, including Web Marketing, Information Retrieval and Digital Forensics. Existing approaches focus on text similarity using keywords as features that are strongly dependent on the dataset. Therefore, dealing with new corpora requires further costly analyses conducted by experts to find out new relevant features. This paper introduces a novel method to detect threads from any type of conversational texts overcoming the issue of previously determining specific features for each dataset. To automatically determine the relevant features of messages we map each message into a three dimensional representation based on its semantic content, the social interactions in terms of sender/recipients and its timestamp; then clustering is used to detect conversation threads. In addition, we propose a supervised approach to detect conversation threads that builds a classification model which combines the above extracted features for predicting whether a pair of messages belongs to the same thread or not. Our model harnesses the distance measure of a message to a cluster representing a thread to capture the probability that a message is part of that same thread. We present our experimental results on seven datasets, pertaining to different types of messages, and demonstrate the effectiveness of our method in the detection of conversation threads, clearly outperforming the state of the art and yielding an improvement of up to a 19%.

Download


Paper Citation


in Harvard Style

Domeniconi G., Semertzidis K., Lopez V., Daly E., Kotoulas S. and Moro G. (2016). A Novel Method for Unsupervised and Supervised Conversational Message Thread Detection . In Proceedings of the 5th International Conference on Data Management Technologies and Applications - Volume 1: DATA, ISBN 978-989-758-193-9, pages 43-54. DOI: 10.5220/0006001100430054

in Bibtex Style

@conference{data16,
author={Giacomo Domeniconi and Konstantinos Semertzidis and Vanessa Lopez and Elizabeth M. Daly and Spyros Kotoulas and Gianluca Moro},
title={A Novel Method for Unsupervised and Supervised Conversational Message Thread Detection},
booktitle={Proceedings of the 5th International Conference on Data Management Technologies and Applications - Volume 1: DATA,},
year={2016},
pages={43-54},
publisher={SciTePress},
organization={INSTICC},
doi={10.5220/0006001100430054},
isbn={978-989-758-193-9},
}


in EndNote Style

TY - CONF
JO - Proceedings of the 5th International Conference on Data Management Technologies and Applications - Volume 1: DATA,
TI - A Novel Method for Unsupervised and Supervised Conversational Message Thread Detection
SN - 978-989-758-193-9
AU - Domeniconi G.
AU - Semertzidis K.
AU - Lopez V.
AU - Daly E.
AU - Kotoulas S.
AU - Moro G.
PY - 2016
SP - 43
EP - 54
DO - 10.5220/0006001100430054