This paper addresses the automatic classification of Russian sentences from legal documents (laws) into those with and without legal vagueness. A training dataset of 6,000 annotated sentences with vagueness loci was developed through collaboration between linguists and legal experts. The study focuses exclusively on vagueness introduced by gradable adjectives. We evaluate both classical machine learning models and transformer-based architectures. Data augmentation applied to RuBERT successfully resolves class imbalance, achieving an F1-score of 0.89. Analysis of linguistic features reveals, that adjectives with the negative prefix “ne-” predominantly occur in sentences without vagueness.
Translated title of the contributionAutomatic Detection of Adjectival Vagueness in Russian Legal Texts: Dataset, Models, and Results
Original languageRussian
Pages (from-to)74-84
JournalInternational Journal of Open Information Technologies
Volume13
Issue number12
StatePublished - 2025

ID: 157209901