The CLASSLA-Stanza model for morphosyntactic annotation of standard Croatian 2.1

PID

The model for morphosyntactic annotation of standard Croatian was built with the CLASSLA-Stanza tool (https://github.com/clarinsi/classla) by training on the hr500k training corpus (http://hdl.handle.net/11356/1792) and using the CLARIN.SI-embed.hr word embeddings (http://hdl.handle.net/11356/1790). The model produces simultaneously UPOS, FEATS and XPOS (MULTEXT-East) labels. The estimated F1 of the XPOS annotations is ~94.87.

The difference to the previous version of the model is that this version was trained using the new version of the hr500k corpus and the new version of the Croatian word embeddings.

Identifier
PID http://hdl.handle.net/11356/1832
Related Identifier http://dx.doi.org/10.18653/v1/W19-3704
Related Identifier http://hdl.handle.net/11356/1348
Related Identifier https://github.com/clarinsi/classla
Metadata Access http://www.clarin.si/repository/oai/request?verb=GetRecord&metadataPrefix=oai_dc&identifier=oai:www.clarin.si:11356/1832
Provenance
Creator Terčon, Luka; Ljubešić, Nikola
Publisher Jožef Stefan Institute
Publication Year 2023
Rights Creative Commons - Attribution-ShareAlike 4.0 International (CC BY-SA 4.0); https://creativecommons.org/licenses/by-sa/4.0/; PUB
OpenAccess true
Contact info(at)clarin.si
Representation
Language Croatian
Resource Type toolService
Format text/plain; charset=utf-8; application/zip; downloadable_files_count: 2
Discipline Linguistics