Research output: Chapter in Book/Report/Conference proceeding › Conference contribution › Research › peer-review
Till today, classification of documents into negative, neutral, or positive remains a key task within the analysis of text tonality/sentiment. There are several methods for the automatic analysis of text sentiment. The method based on network models, the most linguistically sound, to our viewpoint, allows us take into account the syntagmatic connections of words. Also, it utilizes the assumption that not all words in a text are equivalent; some words have more weight and cast higher impact upon the tonality of the text than others. We see it natural to represent a text as a network for sentiment studies, especially in the case of short texts where grammar structures play a higher role in formation of the text pragmatics and the text cannot be seen as just “a bag of words”. We propose a method of text analysis that combines using a lexical mask and an efficient clustering mechanism. In this case, cluster analysis is one of the main methods of typology which demands obtaining formal rules for calculating the number of clusters. The choice of a set of clusters and the moment of completion of the clustering algorithm depend on each other. We show that cluster analysis of data from an n-dimensional vector space using the “single linkage” method can be considered a discrete random process. Sequences of “minimum distances” define the trajectories of this process. “Approximation-estimating test” allows establishing the Markov moment of the completion of the agglomerative clustering process.
Original language | English |
---|---|
Title of host publication | Internet Science. 6th International Conference, INSCI 2019 |
Subtitle of host publication | Proceedings |
Editors | Samira El Yacoubi, Franco Bagnoli, Giovanna Pacini |
Place of Publication | Cham |
Publisher | Springer Nature |
Pages | 18-31 |
Number of pages | 14 |
ISBN (Electronic) | 9783030347703 |
ISBN (Print) | 9783030347697 |
DOIs | |
State | Published - 1 Dec 2019 |
Event | 6th International Conference on Internet Science, INSCI 2019 - Perpignan, France Duration: 2 Dec 2019 → 5 Dec 2019 |
Name | Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics) |
---|---|
Volume | 11938 LNCS |
ISSN (Print) | 0302-9743 |
ISSN (Electronic) | 1611-3349 |
Conference | 6th International Conference on Internet Science, INSCI 2019 |
---|---|
Abbreviated title | INSCI'2019 |
Country/Territory | France |
City | Perpignan |
Period | 2/12/19 → 5/12/19 |
ID: 49784768