ORTOFON v3: corpus of informal spoken Czech with multi-tier transcription (transcriptions & audio)

Dataset

PID

ORTOFON v3 is a corpus of authentic spoken Czech used in informal situations (private environment, spontaneity, unpreparedness etc.) that covers the area of the whole Czech Republic. The corpus is composed of 697 recordings from 2012–2020 and contains 2 445 793 orthographic words (i.e. a total of 2 976 742 tokens including punctuation); a total of 1 121 different speakers appear in the probes. ORTOFON v3 is partially balanced regarding the basic sociolinguistic speaker categories (gender, age group, level of education and region of childhood residence). The transcription is linked to the corresponding audio track. Unlike the ORAL-series corpora, the transcription was carried out on two main tiers, orthographic and phonetic, supplemented by an additional metalanguage tier. The (anonymized) transcriptions are provided in the XML Elan Annotation format, audio (with corresponding anonymization beeps) is in uncompressed 16-bit PCM WAV, mono, 16 kHz format. Another format option of the transcriptions is also available under less restrictive CC BY-NC-SA license at http://hdl.handle.net/11234/1-5687

Identifier
PID	http://hdl.handle.net/11234/1-5686
Related Identifier	http://hdl.handle.net/11234/1-2579
Related Identifier	https://wiki.korpus.cz/doku.php/en:cnk:ortofon
Metadata Access	http://lindat.mff.cuni.cz/repository/oai/request?verb=GetRecord&metadataPrefix=oai_dc&identifier=oai:lindat.mff.cuni.cz:11234/1-5686

Provenance
Creator	Lukeš, David; Kopřivová, Marie; Laubeová, Zuzana; Poukarová, Petra; Horký, Václav; Jelínek, Tomáš; Křivan, Jan; Waclawičová, Martina; Benešová, Lucie; Škarpová, Marie
Publisher	Charles University, Faculty of Arts, Institute of the Czech National Corpus
Publication Year	2024
Rights	License Agreement for Czech National Corpus Data; ACA; https://lindat.mff.cuni.cz/repository/xmlui/page/license-cnc-data
OpenAccess	true
Contact	lindat-help(at)ufal.mff.cuni.cz

Representation
Language	Czech
Resource Type	corpus
Format	text/plain; charset=utf-8; application/x-gzip; downloadable_files_count: 1
Discipline	Linguistics