The CLASSLA-Stanza model for morphosyntactic annotation of standard Croatian 2.1

Dataset

PID

The model for morphosyntactic annotation of standard Croatian was built with the CLASSLA-Stanza tool (https://github.com/clarinsi/classla) by training on the hr500k training corpus (http://hdl.handle.net/11356/1792) and using the CLARIN.SI-embed.hr word embeddings (http://hdl.handle.net/11356/1790). The model produces simultaneously UPOS, FEATS and XPOS (MULTEXT-East) labels. The estimated F1 of the XPOS annotations is ~94.87.

The difference to the previous version of the model is that this version was trained using the new version of the hr500k corpus and the new version of the Croatian word embeddings.

Identifier
PID	http://hdl.handle.net/11356/1832
Related Identifier	http://dx.doi.org/10.18653/v1/W19-3704
Related Identifier	http://hdl.handle.net/11356/1348
Related Identifier	https://github.com/clarinsi/classla
Metadata Access	http://www.clarin.si/repository/oai/request?verb=GetRecord&metadataPrefix=oai_dc&identifier=oai:www.clarin.si:11356/1832

Provenance
Creator	Terčon, Luka; Ljubešić, Nikola
Publisher	Jožef Stefan Institute
Publication Year	2023
Rights	Creative Commons - Attribution-ShareAlike 4.0 International (CC BY-SA 4.0); https://creativecommons.org/licenses/by-sa/4.0/; PUB
OpenAccess	true
Contact	info(at)clarin.si

Representation
Language	Croatian
Resource Type	toolService
Format	text/plain; charset=utf-8; application/zip; downloadable_files_count: 2
Discipline	Linguistics