Combining WordNet and Word Embeddings in Data Augmentation for Legal Texts

Abstract

Creating balanced labeled textual corpora for complex tasks, like legal analysis, is a challenging and expensive process that often requires the collaboration of domain experts. To address this problem, we propose a data augmentation method based on the combination of GloVe word embeddings and the WordNet ontology. We present an example of application in the legal domain, specifically on decisions of the Court of Justice of the European Union. Our evaluation with human experts confirms that our method is more robust than the alternatives.

Publication
Proceedings of the Natural Legal Language Processing Workshop 2022
Andrea Galassi
Andrea Galassi
Junior Assistant Professor

He is an expert in deep learning architectures for natural language processing.

Federico Ruggeri
Federico Ruggeri
Postdoctoral Research Fellow

His research aims to devise Natural Language Processing (NLP) systems that learn to generate, distill, and use knowledge from unstructured text.

Paolo Torroni
Paolo Torroni
Associate Professor

Head of the Language Technologies lab. His main research focus is in artificial intelligence, and in particular natural language processing, multi-agent systems, and computational logics.