TU Darmstadt / ULB / TUbiblio

Label modification and bootstrapping for zero-shot cross-lingual hate speech detection

Bigoulaeva, Irina ; Hangya, Viktor ; Gurevych, Iryna ; Fraser, Alexander (2023)
Label modification and bootstrapping for zero-shot cross-lingual hate speech detection.
In: Language Resources and Evaluation
doi: 10.1007/s10579-023-09637-4
Artikel, Bibliographie

Kurzbeschreibung (Abstract)

The goal of hate speech detection is to filter negative online content aiming at certain groups of people. Due to the easy accessibility and multilinguality of social media platforms, it is crucial to protect everyone which requires building hate speech detection systems for a wide range of languages. However, the available labeled hate speech datasets are limited, making it difficult to build systems for many languages. In this paper we focus on cross-lingual transfer learning to support hate speech detection in low-resource languages, while highlighting label issues across application scenarios, such as inconsistent label sets of corpora or differing hate speech definitions, which hinder the application of such methods. We leverage cross-lingual word embeddings to train our neural network systems on the source language and apply them to the target language, which lacks labeled examples, and show that good performance can be achieved. We then incorporate unlabeled target language data for further model improvements by bootstrapping labels using an ensemble of different model architectures. Furthermore, we investigate the issue of label imbalance in hate speech datasets, since the high ratio of non-hate examples compared to hate examples often leads to low model performance. We test simple data undersampling and oversampling techniques and show their effectiveness.

Typ des Eintrags: Artikel
Erschienen: 2023
Autor(en): Bigoulaeva, Irina ; Hangya, Viktor ; Gurevych, Iryna ; Fraser, Alexander
Art des Eintrags: Bibliographie
Titel: Label modification and bootstrapping for zero-shot cross-lingual hate speech detection
Sprache: Englisch
Publikationsjahr: 18 Februar 2023
Verlag: Springer
Titel der Zeitschrift, Zeitung oder Schriftenreihe: Language Resources and Evaluation
DOI: 10.1007/s10579-023-09637-4
Kurzbeschreibung (Abstract):

The goal of hate speech detection is to filter negative online content aiming at certain groups of people. Due to the easy accessibility and multilinguality of social media platforms, it is crucial to protect everyone which requires building hate speech detection systems for a wide range of languages. However, the available labeled hate speech datasets are limited, making it difficult to build systems for many languages. In this paper we focus on cross-lingual transfer learning to support hate speech detection in low-resource languages, while highlighting label issues across application scenarios, such as inconsistent label sets of corpora or differing hate speech definitions, which hinder the application of such methods. We leverage cross-lingual word embeddings to train our neural network systems on the source language and apply them to the target language, which lacks labeled examples, and show that good performance can be achieved. We then incorporate unlabeled target language data for further model improvements by bootstrapping labels using an ensemble of different model architectures. Furthermore, we investigate the issue of label imbalance in hate speech datasets, since the high ratio of non-hate examples compared to hate examples often leads to low model performance. We test simple data undersampling and oversampling techniques and show their effectiveness.

Fachbereich(e)/-gebiet(e): 20 Fachbereich Informatik
20 Fachbereich Informatik > Ubiquitäre Wissensverarbeitung
Hinterlegungsdatum: 13 Mär 2023 08:55
Letzte Änderung: 04 Jul 2023 10:53
PPN: 509266339
Export:
Suche nach Titel in: TUfind oder in Google
Frage zum Eintrag Frage zum Eintrag

Optionen (nur für Redakteure)
Redaktionelle Details anzeigen Redaktionelle Details anzeigen