The Old Czech Corpus (OCC) is a digitally curated collection of texts written in the Czech language from the 13th to the 16th centuries. It brings together manuscripts, printed books, legal documents, religious treatises, and lyrical poetry, providing scholars with a unique window into the early development of Czech as a literarylanguage and its sociocultural context.
Old Czech (also called prehistoric Czech or medieval Czech) emerged during a period of profound political and cultural change in Central Europe. During the reign of the Pemyslid dynasty and later the Jagiellonian kings, Czech began to replace Latin in many public domains, especially after the Hus reform movement of the early 15thcentury. This linguistic shift is reflected in the corpus, which documents:
These documents demonstrate the gradual standardisation of orthography, the introduction of diacritics, and the influence of German, Latin, and later Polish on the Czech lexicon.
The OCC currently contains roughly 12million words, arranged into three major subcorpora:
Over 200 handwritten sources, digitised at high resolution and transcribed using the Text Encoding Initiative (TEI) guidelines. The manuscript subcorpus includes:
Printed material from the first Czech printing presses (14751520). The most significant titles are:
Inscriptions, municipal statutes, and court records, many of which survive only as fragments. This subcorpus is essential for research on dialectal variation and sociolegal language.
All texts are available through a webbased interface that supports:
Opensource tools such as UDPipe and Sketch Engine have been adapted for the corpus, giving researchers the ability to generate frequency lists, collocations, and concordances in real time.
The OCC has been used in a wide variety of projects, ranging from philology to computational linguistics. Some notable examples include:
By tracking the occurrence of specific phoneme representations (e.g., the shift from to ), scholars have mapped the gradual phonological changes that led to modern Czech vowel harmony.
Statistical analysis of loanwords shows three major periods of influence:
Using the corpuss temporal tagging, researchers have traced how words such as svoboda (originally freedom in the legal sense) evolved to acquire political connotations during the early modern period.
Machinelearning classifiers trained on the OCC successfully differentiate between authors of anonymous sermons, helping to attribute previously unknown texts to known scribes.
If you are new to the OCC, follow these steps:
kiovatka AND 14* for all occurrences in the 1400s.NNFS1----A---- (noun, feminine, singular, nominative).Č for ).Most tutorials are available under the Help section, and a community forum allows you to ask questions about tagsets, data cleaning, or integration with external tools.
The OCC project aims to expand both its quantitative size and its analytical depth. Upcoming initiatives include:
These developments will further solidify the OCC as an essential resource for anyone studying the linguistic, cultural, or historical landscape of medieval Central Europe.
Bernek, Ji. (2018). Old Czech Corpus Methodology and First Results. Prague: Charles University Press.
Novk, Petra. (2020). Phonological Change in Early Czech: Evidence from the OCC. Journal of Slavic Linguistics 28(2): 145170.
Stbrn, Karel. (2022). Lexical Borrowing in Medieval Czech Texts. Central European Review 9(1): 3358.
Vesel, Marta & Dvok, Luk. (2024). Automatic Tagging of Old Czech: A Neural Approach. Proceedings of the 12th International Conference on Historical Linguistics, pp. 7184.
The Old Czech Corpus not only preserves the words of our ancestors; it makes them accessible for the digital age, allowing us to ask questions that were impossible a decade ago.
Jan Hruka, Director of the Czech Institute of Language Studies
