Offensive language dataset of Croatian, English and Slovenian comments FRENK 1.0

Dataset

PID

The FRENK dataset consists of comments to Facebook posts (news articles) of mainstream media outlets from Croatia, Great Britain, and Slovenia, on the topics of migrants and LGBT. The dataset contains whole discussion threads. Each comment is annotated by the type of socially unacceptable discourse (e.g., inappropriate, offensive, violent speech) and its target (e.g., migrants/LGBT, commenters, media). The annotation schema in its details is described in https://arxiv.org/pdf/1906.02045.pdf. Usernames in the metadata are pseudo-anonymised and removed from the comments.

The data in each language (Croatian (hr), English (en), Slovenian (sl), and topic (migrants, LGBT) is divided into a training and a testing portion. The training and testing data consist of separate discussion threads, i.e., there is no cross-discussion-thread contamination between training and testing data. The sizes of the splits are the following: Croatian, migrants: 4356 training comments, 978 testing comments; Croatian LGBT: 4494 training comments, 1142 comments; English, migrants: 4540 training comments, 1285 testing comments; English, LGBT: 4819 training comments, 1017 testing comments; Slovenian, migrants: 5145 training comments, 1277 testing comments; Slovenian, LGBT: 2842 training comments, 900 testing comments.

Identifier
PID	http://hdl.handle.net/11356/1433
Related Identifier	https://arxiv.org/pdf/1906.02045.pdf
Related Identifier	http://hdl.handle.net/11356/1462
Related Identifier	http://nl.ijs.si/frenk/
Metadata Access	http://www.clarin.si/repository/oai/request?verb=GetRecord&metadataPrefix=oai_dc&identifier=oai:www.clarin.si:11356/1433

Provenance
Creator	Ljubešić, Nikola; Fišer, Darja; Erjavec, Tomaž
Publisher	Jožef Stefan Institute
Publication Year	2021
Rights	CLARIN.SI Licence ACA ID-BY-NC-INF-NORED 1.0; https://clarin.si/repository/xmlui/page/licence-aca-id-by-nc-inf-nored-1.0; ACA
OpenAccess	true
Contact	info(at)clarin.si

Representation
Language	Croatian; English; Slovenian; Slovene
Resource Type	corpus
Format	text/plain; charset=utf-8; application/zip; downloadable_files_count: 1
Discipline	Linguistics