Facebook
TwitterODC Public Domain Dedication and Licence (PDDL) v1.0http://www.opendatacommons.org/licenses/pddl/1.0/
License information was derived automatically
This repository contains a collection of Russian literature in txt format (all in UTF-8 encoding). In addition, for each author there is a csv file containing information about the year of writing of each work.
This dataset was created for a project to determine the authorship of a piece of text, but I'm sure that you can use this dataset for anything 😉.
The main feature that allows this dataset to be used for any purpose is that the data is not processed at all. The text has not been pre-processed in any way, the designations of authors, chapters and references to the translation of foreign inserts have not been removed.
Thanks Ilibrary, LitLib, Wikisource and all-all-all.
Facebook
TwitterApache License, v2.0https://www.apache.org/licenses/LICENSE-2.0
License information was derived automatically
This dataset was created by Matvei Danilov
Released under Apache 2.0
Facebook
TwitterAttribution-NonCommercial 4.0 (CC BY-NC 4.0)https://creativecommons.org/licenses/by-nc/4.0/
License information was derived automatically
Complete dataset containing the academic profile, journal ranking, indexing metrics, and publication metadata for Hermeneutics of Old Russian Literature Journal (Arts) [ISSN: 1607-6192].
Facebook
TwitterMIT Licensehttps://opensource.org/licenses/MIT
License information was derived automatically
• articles.txt – Texts of popular articles on various topics published on dzen.ru (~20 million characters)
• books-A.txt – Fragments of various works of world-class Russian and foreign literature (~20 million characters)
• books-B.txt – Fragments of various works of literature, both world-famous and little-known (~20 million characters)
• fanfiction.txt – Texts of popular fanfiction on various topics published on ficbook.net (~20 million characters)
• jokes.txt – Texts of various jokes and puns (~6.7 million characters)
• poems.txt – Texts of various poems by world-famous authors (~40 million characters)
Facebook
Twitterhttps://rascasse.com/termshttps://rascasse.com/terms
Demographic, psychographic, geographic and brand-affinity data for the Russian literature audience in United States, sourced from Rascasse's panel of 12+ social and digital signals.
Facebook
TwitterThe data material consists of a detailed description of a review corpus used in order to analyze the reception of Russian literature in Sweden. The investigations that have and will be conducted based on the review corpus analyze for example translation visibility, translation criticism and the image of Russian literature in the Swedish literary system. The review corpus consists of 430 reviews of post-Soviet Russian novels published in Swedish translation between 1992 and 2020. The reviews are protected by copyright and may not be made available. Therefore, the data instead contains a complete specification of the review database, information regarding how the reviews have been classified, and finally, information about thematic coding related to specific investigations (articles).
Facebook
TwitterThe repository contains data, scripts, and main figures from the article "The Rise and Fall of Poetry in Nineteenth-Century Russian Literature: Age-Period-Cohort Analysis."
Facebook
TwitterJoined corpus of russian books. Can be good for text generation networks like gpt I do not own any of these texts and they should be used for educational purposes only, i guess
Facebook
TwitterMIT Licensehttps://opensource.org/licenses/MIT
License information was derived automatically
Recent advances in the field of universal language models and transformers require the development of a methodology for their broad diagnostics and testing for general intellectual skills - detection of natural language inference, commonsense reasoning, ability to perform simple logical operations regardless of text subject or lexicon. For the first time, a benchmark of nine tasks, collected and organized analogically to the SuperGLUE methodology, was developed from scratch for the Russian language. We provide baselines, human level evaluation, an open-source framework for evaluating models and an overall leaderboard of transformer models for the Russian language.
Facebook
TwitterAttribution 4.0 (CC BY 4.0)https://creativecommons.org/licenses/by/4.0/
License information was derived automatically
This article examines the representation of nature in Russian literature as a complex semiotic, philosophical, and cultural system. Drawing on a wide range of literary periods from Old Russian texts to contemporary works, the study analyzes how natural imagery functions not merely as a descriptive element but as a structural and meaning-generating component of the художественный текст. Particular attention is paid to the evolution of the landscape motif, its psychological, symbolic, and ontological dimensions, and its role in expressing the relationship between human beings and the surrounding world. The research integrates approaches from structuralism, semiotics, cognitive poetics, and cultural studies to demonstrate that nature in Russian literature encodes national identity, collective memory, and universal existential concepts. The article also highlights the growing relevance of ecological discourse in modern literary contexts, emphasizing the role of literature in shaping environmental consciousness and ethical attitudes toward nature.
Facebook
TwitterThe data material contains editions (in book form) of Swedish literature published in Russian translation on Russian (and Soviet) publishing houses between 1946–2021. The data material aims to create a general overview of translation and publication of Swedish literature in Russia, and will form the basis for further analyses of the publication and reception of Swedish literature in Russia.
Data has been gathered using Swedish and Russian library resources (The Swedish Royal Library’s catalog Libris, the digital catalog of The National Library of Russia, and the digital catalog of The Russian State Library. Since neither catalog is complete, additional searches have been conducted on various literary webpages (livelib.ru, fantlab.ru), and web sites of Russian publishing houses.
Version 1 of the data material is restricted to editions of adult prose fiction, poetry and drama. Editions of non-fiction and children's literature will be added in following versions. Furthermore, Version 1 only includes a limited number of variables. A version containing more variables (about translators, types of translation, contents of anthologies etc.) will be published when the research project is finalized and the related articles have been published.
The data is available as an Excel file and as tab-separated text (.xlsx, .tsv).
Facebook
Twitterhttps://choosealicense.com/licenses/unknown/https://choosealicense.com/licenses/unknown/
RusLit Corpus
A corpus of Russian literature in clean text format.
Description
This dataset contains cleaned text (stripped of extraneous artifacts) collected from works by authors who passed away more than 70 years ago, placing them in the public domain. Some texts may contain semantic nonsense (e.g. OCR or digitization artifacts), as well as fragments of French, German, English, or Japanese text mixed in with the Russian.
Structure
Each record… See the full description on the dataset page: https://huggingface.co/datasets/RafaelUI/russian_literature.
Facebook
TwitterThe article explores the development of Russian literature from ancient chronicles to contemporary prose. It examines its role as a moral, philosophical, and historical phenomenon reflecting the nation’s spiritual path. Special attention is given to key writers and works that shape moral values and national identity.
Facebook
TwitterThis dataset was created by Kosarevsky Dmitry
Facebook
TwitterAttribution-NonCommercial-ShareAlike 4.0 (CC BY-NC-SA 4.0)https://creativecommons.org/licenses/by-nc-sa/4.0/
License information was derived automatically
This is a study for Aleksandr Herzen's tombstone in Nice. Aleksandr Ivanovich Herzen (Russian: Алекса́ндр Ива́нович Ге́рцен; April 6 [O.S. 25 March] 1812 – January 21 [O.S. 9 January] 1870) was a Russian writer and thinker known as the father of Russian socialism and one of the main fathers of agrarian populism (being an ideological ancestor of the Narodniki, Socialist-Revolutionaries, Trudoviks and the agrarian American Populist Party). He is held responsible for creating a political climate leading to the emancipation of the serfs in 1861. His autobiography, My Past and Thoughts, is often considered the best specimen of that genre in Russian literature. He also published the important social novel Who is to Blame? (1845–46).
Facebook
Twitterhttps://rascasse.com/termshttps://rascasse.com/terms
Demografische, psychografische, geografische und Marken-Affinitätsdaten zur Zielgruppe Russian literature in Deutschland – aus über 12 digitalen Signalquellen von Rascasse.
Facebook
Twitter🇷🇺 Russian Public Domain 🇷🇺
Russian-Public Domain or Russian-PD is a large collection aiming to aggregate all Russian monographies and periodicals in the public domain.
Dataset summary
The collection contains 8525 titles making up 995,163,165 words recovered from the Internet Archive. Each parquet file has the full text of 2,000 books selected at random.
Curation method
The composition of the dataset adheres to the criteria for public domain works in the… See the full description on the dataset page: https://huggingface.co/datasets/PleIAs/Russian-PD.
Facebook
TwitterAttribution 4.0 (CC BY 4.0)https://creativecommons.org/licenses/by/4.0/
License information was derived automatically
Russian LibriSpeech (RuLS)
Identifier: SLR96 from openslr.org Summary: This dataset is based on LibriVox audiobooks Category: Speech License: The dataset is Public Domain in the USA. About this resource: Russian LibriSpeech (RuLS) dataset is based on LibriVox's public domain audio books (see BOOKS.TXT for the list of included books) and contains about 98 hours of audio data.
Facebook
TwitterThe data material consists of a detailed description of a review corpus used in order to analyze the reception of Russian literature in Sweden. The reviews are protected by copyright and may not be made available. Therefore, the data instead contains a complete specification of the review database, information regarding how the reviews have been classified, and finally, information about the authors, translators, critics and media sources related to the material.
Facebook
TwitterApache License, v2.0https://www.apache.org/licenses/LICENSE-2.0
License information was derived automatically
The Multilingual Literature Parallel Corpus is designed for translation tasks, containing parallel text pairs from literature in three languages: Kazakh (kaz_Cyrl), Russian (rus_Cyrl), and English (eng_Latn).
The Multilingual Literature Parallel Corpus provides parallel text pairs for translation tasks across Kazakh, Russian, and English. The dataset is curated to support the development of machine translation models by offering high-quality parallel sentences from literature.
The dataset is suitable for developing and benchmarking translation models, cross-linguistic analysis, and linguistic research in the context of Kazakh, Russian, and English literature.
The dataset is not suitable for non-literary text translation tasks, real-time translation applications, or any use cases requiring domain-specific jargon outside literature.
The dataset was created to enhance the quality and accessibility of machine translation models for Kazakh, Russian, and English, specifically within the literary domain.
The source data consists of original literary texts in Kazakh, Russian, and English.
The source data consists of translated literary texts in Kazakh, Russian, and English.
The data was collected from publicly available literary sources, preprocessed to align translations accurately, and normalized to maintain consistency in formatting and structure.
Facebook
TwitterODC Public Domain Dedication and Licence (PDDL) v1.0http://www.opendatacommons.org/licenses/pddl/1.0/
License information was derived automatically
This repository contains a collection of Russian literature in txt format (all in UTF-8 encoding). In addition, for each author there is a csv file containing information about the year of writing of each work.
This dataset was created for a project to determine the authorship of a piece of text, but I'm sure that you can use this dataset for anything 😉.
The main feature that allows this dataset to be used for any purpose is that the data is not processed at all. The text has not been pre-processed in any way, the designations of authors, chapters and references to the translation of foreign inserts have not been removed.
Thanks Ilibrary, LitLib, Wikisource and all-all-all.