The Russian National Corpus is a representative collection of texts in Russian, counting more than 17 bln tokens and completed with linguistic annotation and search tools
Search in corpora
News
Show allWe continue to expand the corpus functionality for teaching Russian at school. The Practice Example Generator has been updated with rules for spelling consonants in prefixes. These include invariant prefixes such as в-, от-, над-, под, меж-, and others; prefixes ending in з-/с-, such as без-/бес-, из-/ис-, раз-/роз and рас-/рос-, and others; and borrowed prefixes such as экс-, суб-. The new rules cover 16 groups of words with prefixes.
You can access the generator page from the RNC for Schools section by clicking on the corresponding banner.
The GICR (VK) corpus has been expanded by 4.3 billion word tokens. This segment of GICR is now fully available to users. Annotation has been added for the author’s age at the time the text was written. The display of statistics by sociolinguistic parameters has been improved. In the corpus and subcorpus overviews, the distribution of texts by year of publication is now shown as a chart.
The Russian Classics corpus has been expanded by 1.6 million word tokens. The update includes works by authors already represented in the corpus: Baratynsky’s prose and letters, Zhukovsky’s poems not included in the Soviet four-volume edition, selected works and translations by Pushkin, Dostoevsky, and Turgenev, as well as Chekhov’s journalism and notebooks.
The update also restores technical losses of text and graphic formatting elements in works by several authors.