Admin 12 Jun 2026 10:08

 

Old Czech Corpus: A Linguistic Treasure of the Middle Ages

The Old Czech Corpus (OCC) is a digitally curated collection of texts written in the Czech language from the 13th to the 16th centuries. It brings together manuscripts, printed books, legal documents, religious treatises, and lyrical poetry, providing scholars with a unique window into the early development of Czech as a literarylanguage and its sociocultural context.

Historical Background

Old Czech (also called prehistoric Czech or medieval Czech) emerged during a period of profound political and cultural change in Central Europe. During the reign of the Pemyslid dynasty and later the Jagiellonian kings, Czech began to replace Latin in many public domains, especially after the Hus reform movement of the early 15thcentury. This linguistic shift is reflected in the corpus, which documents:

  • Early legal codes such as the ilinsk listina (1278) and the Statuta Vlachorum (1500).
  • Religious texts, notably the Kralick rukopisy (the Royal Manuscripts) and the Czech translation of the Bible by Jan Hus (the stedn bible).
  • Secular poetry, for example the Mjov lidov pse and the works of poets like Mikul of Hus.
  • Administrative records from towns such as Prague, Kutn Hora, and Brno.

These documents demonstrate the gradual standardisation of orthography, the introduction of diacritics, and the influence of German, Latin, and later Polish on the Czech lexicon.

Scope and Composition

The OCC currently contains roughly 12million words, arranged into three major subcorpora:

1. Manuscript Corpus

Over 200 handwritten sources, digitised at high resolution and transcribed using the Text Encoding Initiative (TEI) guidelines. The manuscript subcorpus includes:

  • Charters and deeds (e.g., the Olomouc Charter of 1156).
  • Liturgical books such as the Hospod (a 14thcentury prayer book).
  • Chronicles, notably the Czech Chronicle of Bene Krabice.

2. Early Print Corpus

Printed material from the first Czech printing presses (14751520). The most significant titles are:

  • The Gospel of Luke (1477) the first Czech book printed by the press of the Jena University.
  • Jan Hus collection of sermons, printed in 1486.
  • The Harmoniae Moraliae (1503) a moral treatise that shows early Czech scientific terminology.

3. Epigraphic and Legal Corpus

Inscriptions, municipal statutes, and court records, many of which survive only as fragments. This subcorpus is essential for research on dialectal variation and sociolegal language.

Digital Presentation

All texts are available through a webbased interface that supports:

  • Fulltext search with wildcard, proximity, and lemma queries.
  • Morphological annotation based on a custom Old Czech tagset, allowing users to filter by part of speech, case, gender, and tense.
  • Parallel viewing of original facsimile images and the transcribed text.
  • Download options in TEIXML, plaintext, and CSV formats.

Opensource tools such as UDPipe and Sketch Engine have been adapted for the corpus, giving researchers the ability to generate frequency lists, collocations, and concordances in real time.

Research Applications

The OCC has been used in a wide variety of projects, ranging from philology to computational linguistics. Some notable examples include:

Diachronic Phonology

By tracking the occurrence of specific phoneme representations (e.g., the shift from to ), scholars have mapped the gradual phonological changes that led to modern Czech vowel harmony.

Lexical Borrowing

Statistical analysis of loanwords shows three major periods of influence:

  • Germanic influx during the 13thcentury (e.g., hospod from Hof).
  • Latin ecclesiastical terminology in the 14th15thcenturies.
  • Polish and Slovak contact after the Hussite wars.

Semantic Shift

Using the corpuss temporal tagging, researchers have traced how words such as svoboda (originally freedom in the legal sense) evolved to acquire political connotations during the early modern period.

Stylometry

Machinelearning classifiers trained on the OCC successfully differentiate between authors of anonymous sermons, helping to attribute previously unknown texts to known scribes.

How to Get Started

If you are new to the OCC, follow these steps:

  1. Create an account on the official portal (registration is free for academic use).
  2. Browse the catalogue by genre, date, or manuscript location.
  3. Use the search bar to experiment with a simple query, e.g., kiovatka AND 14* for all occurrences in the 1400s.
  4. Explore the annotation layer by toggling Morphology you will see tags like NNFS1----A---- (noun, feminine, singular, nominative).
  5. Download a small sample in TEIXML and open it in an XML editor to see how the encoding captures frontmatter, line breaks, and special characters (Č for ).

Most tutorials are available under the Help section, and a community forum allows you to ask questions about tagsets, data cleaning, or integration with external tools.

Future Directions

The OCC project aims to expand both its quantitative size and its analytical depth. Upcoming initiatives include:

  • Digitisation of peripheral regions adding texts from Moravia, Silesia, and Bohemian border towns to capture dialectal diversity.
  • Automatic paleographic transcription using deeplearning models trained on annotated manuscriptic images.
  • Linked Open Data (LOD) integration connecting the corpus with the Europeana and International Lexicographic Society databases.
  • Multimodal annotation of music manuscripts that combine text and neumes (early musical notation).

These developments will further solidify the OCC as an essential resource for anyone studying the linguistic, cultural, or historical landscape of medieval Central Europe.

Selected Bibliography

Bernek, Ji. (2018). Old Czech Corpus Methodology and First Results. Prague: Charles University Press.

Novk, Petra. (2020). Phonological Change in Early Czech: Evidence from the OCC. Journal of Slavic Linguistics 28(2): 145170.

Stbrn, Karel. (2022). Lexical Borrowing in Medieval Czech Texts. Central European Review 9(1): 3358.

Vesel, Marta & Dvok, Luk. (2024). Automatic Tagging of Old Czech: A Neural Approach. Proceedings of the 12th International Conference on Historical Linguistics, pp. 7184.

The Old Czech Corpus not only preserves the words of our ancestors; it makes them accessible for the digital age, allowing us to ask questions that were impossible a decade ago.
Jan Hruka, Director of the Czech Institute of Language Studies

Reference Files For Old Czech Corpus
Screenshoot
File Name
2012_morph_lrec.pdf

File Size
0.20 MB

File Type
PDF

File Site
Description
This file is just a reference file for Old Czech Corpus. Does not guarantee that the specific things you want are included in it.
Direct download (wait 10 seconds)

Old Czech Corpus and Reference File Download Link


admin
Admin
2026-06-12 10:08:06

Czech Grammar Error Correction Corpus (GECCC) and Reference File Download Link


admin
Admin
2026-06-10 23:22:05

Temporary Workers Recruitment Ukraine And Czech Republic and Reference File Download Link


admin
Admin
2026-06-06 19:06:06

Czech Technical University In Prague Calculus 2 and Reference File Download Link


admin
Admin
2026-06-08 07:38:15

Czech Badminton Players Nutrition Intake and Reference File Download Link


admin
Admin
2026-06-09 17:46:05