Facebook
TwitterThis is the first Russian news summarization dataset. A paper about this dataset: https://arxiv.org/pdf/2006.11063.pdf Additional files and notebooks: https://github.com/IlyaGusev/gazeta/ Previous datasets for headline generation: https://github.com/RossiyaSegodnya/ria_news_dataset https://www.kaggle.com/yutkin/corpus-of-russian-news-articles-from-lenta
This is the second version of the dataset. The data structure is pretty straightforward. Every line of a file is a JSON object with 5 fields: URL, title, text, summary, and date. The dataset consists of 74126 examples. The first 60964 examples by date are in the training dataset, the proceeding 6369 examples are in the validation dataset, and the remaining 6793 pairs are in the test dataset.
Legal basis for distribution of the dataset: https://www.gazeta.ru/credits.shtml, paragraph 2.1.2. All rights belong to "www.gazeta.ru". This dataset can be removed at the request of the copyright holder. Usage of this dataset is possible only for personal purposes on a non-commercial basis.
Facebook
TwitterThis dataset was introduced at READi workshop (LREC-COLING 2024). This dataset was proposed to train a model for Russian legal text simplification. This dataset has already helped to train GPT and T5 models for this task, as well as in Dynamic Topic Modelling task to analyze the history of Russian law from 2009 to 2022 .
We have collected our data from Rossiyskaya Gazeta website. It's a Russian newspaper published by the Government of Russia. The daily newspaper serves as the official government gazette of the Government of the Russian Federation, publishing government-related affairs such as official decrees, statements and documents of state bodies, the promulgation of newly approved laws, Presidential decrees, and government announcements. Rossiyskaya Gazeta provides legal text descriptions for common people called "comments". But these descriptions are made only for important documents, so while there are hundreds of thousands of legal documents in Russia, only a couple of thousands has a "comment". We used this "comment" as simplified version of the document.
Overall there are 2963 pairs of original documents and simplified ones. Dataset contains documents from December 31, 2008 up to November 28, 2022 - thus it contains COVID-19-related laws too.
Dataset has 5 columns: 1. Название документа (Document Title) 2. Ссылка (Link to the original document) 3. Текст (Original document text) 4. Комментарий РГ (Rossiyskaya Gazeta comment) 5. Дата (Publication date)
Example of original document text (2nd article in row):
https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F11047041%2F625daafe6319319700f754f5ac1d89e4%2Forig.png?generation=1676324147067453&alt=media" alt="">
Example of Rossiyskaya Gazeta comment:
https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F11047041%2F3230f4df60713b417826434ddacc8882%2Fcomm.png?generation=1676324185146892&alt=media" alt="">
The photo on the headline was taken from Roscosmos official website. `
Facebook
TwitterThe repository contains an ongoing collection of tweets IDs associated with the current conflict in Ukraine and Russia, which we commenced collecting on Februrary 22, 2022. To comply with Twitter’s Terms of Service, we are only publicly releasing the Tweet IDs of the collected Tweets. The data is released for non-commercial research use. Note that the compressed files must be first uncompressed in order to use included scripts. This dataset is release v1.3 and is not actively maintained -- the actively maintained dataset can be found here: https://github.com/echen102/ukraine-russia. This release contains Tweet IDs collected from 2/22/22 - 1/08/23. Please refer to the README for more details regarding data, data organization and data usage agreement. This dataset is licensed under the Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International Public License . By using this dataset, you agree to abide by the stipulations in the license, remain in compliance with Twitter’s Terms of Service, and cite the following manuscript: Emily Chen and Emilio Ferrara. 2022. Tweets in Time of Conflict: A Public Dataset Tracking the Twitter Discourse on the War Between Ukraine and Russia. arXiv:cs.SI/2203.07488
Facebook
TwitterThis dataset is designed for research on audio deepfake detection, focusing specifically on generated speech in Russian. It contains TTS-generated audio, paired with transcriptions, and a mixed set for real vs fake classification tasks.
The main goal is to support research on audio deepfake detection in underrepresented languages, especially Russian. The dataset simulates real-world scenarios using multiple state-of-the-art TTS systems to generate fakes and includes clean, real audio data.
We used three high-quality TTS models to synthesize Russian speech:
XTTS-v2: Cross-lingual, zero-shot voice cloning with multilingual support.
Silero TTS: Lightweight, real-time Russian TTS model.
VITS RU Multispeaker: VITS-based Russian model with speaker variability.
For real human speech, we used a part of SOVA dataset, which contains clean Russian utterances recorded by multiple speakers.
Facebook
TwitterCC0 1.0 Universal Public Domain Dedicationhttps://creativecommons.org/publicdomain/zero/1.0/
License information was derived automatically
Russian Election Data: data (.csv, .rds) and codebook. Replication material are available at: https://github.com/georgytarasenko/RED-replication-package
Facebook
TwitterSupplement to the Lenta.Ru news dataset (until December 2019)
Updated script to download data in current format: GitHub
Facebook
TwitterThe set of over 2,250 files archived here comprises a database of the Russian Constructicon, an open-access electronic resource freely available at https://constructicon.github.io/russian/. The Russian Constructicon is a searchable database of constructions accompanied with thorough descriptions of their properties and annotated illustrative examples.
Facebook
TwitterSOVA (Balalaika)
[!IMPORTANT] Official dataset for our INTERSPEECH 2026 paper "A Data-Centric Framework for Addressing Phonetic and Prosodic Challenges in Russian Speech Generative Models" (arXiv:2507.13563). Part of the Balalaika Russian speech data-processing pipeline — code: https://github.com/lab260ru/balalaika. If you use this resource, please cite it.
Part of the Balalaika Russian speech data-processing pipeline. See the code repository for details.… See the full description on the dataset page: https://huggingface.co/datasets/lab260/sova_balalaika.
Facebook
TwitterPoeTree (Poetry Treebanks) is a dataset comprising over 300,000 poems / 84,000,000 tokens in nine languages (Czech, English, French, German, Hungarian, Italian, Portuguese, Spanish, and Russian). Each corpus has been deduplicated, enriched with Universal Dependencies, provided with additional metadata and converted into a unified JSON structure (schema available at https://versologie.cz/poetree/json-schema).
Facebook
TwitterComprehensive dataset of Tweets containing the keyword 'bucha' around the Ukraine Invasion in February 2022. The user handle column has been excluded to protect deleted accounts that have not been retweeted or replied to. Tweets have been collected via the Academic API using the Search endpoint in four languages: Languages English search term: "bucha AND lang='en'" German search term: "(bucha OR butscha) AND lang='de'" Russian search term: (Бу́ча OR bucha) AND lang:ru Ukrainian search term: (Бу́ча OR bucha) AND lang:uk Timeframe Data starts 1. March 2022 and ends on 27. May 2023. Collection dates Details on collection dates per Tweet (e.g. to compare with creation dates) as well as the IDs of Tweets for consistency checks can be found here: https://github.com/Leibniz-HBI/ukraine_twitter_data (https://doi.org/10.17605/OSF.IO/RTQXN)
Facebook
TwitterAttribution 4.0 (CC BY 4.0)https://creativecommons.org/licenses/by/4.0/
License information was derived automatically
Kyrgyz comments are still new things to examine. Defining whether comment is toxic or not is a big challenge. We used YouTube API to get comments. Code can be found here - https://github.com/abdulra7ma/ycc
Facebook
TwitterApache License, v2.0https://www.apache.org/licenses/LICENSE-2.0
License information was derived automatically
Mathematical Reasoning Dataset (English & Russian)
A bilingual collection of synthetic school-level mathematics questions and answers, based on the DeepMind mathematics_dataset generator. This dataset contains two language splits:
en — the original English data, taken as-is from the official mathematics_dataset-v1.0 release published by Google DeepMind (github.com/google-deepmind/mathematics_dataset). ru — a Russian version generated from scratch with a translated fork of the… See the full description on the dataset page: https://huggingface.co/datasets/d0rj/mathematics_dataset.
Facebook
TwitterAttribution 4.0 (CC BY 4.0)https://creativecommons.org/licenses/by/4.0/
License information was derived automatically
This dataset accompanies the article "2D ADAPTIVE MULTI-INTERVAL SCALE (AMIS): METHOD FOR NORMALIZATION AND VISUALIZATION OF SPATIAL DATA" (https://doi.org/10.5281/zenodo.20577673). The dataset contains: 1. Example data: - Reference dataset (32 teams, Russian Championship 2012-13, 25×40 grid, 32,000 observations) - Match data (France vs Croatia, 2018 World Cup final; England vs Germany, 1966 World Cup final) 2. Software: - Executable tool (2D AMIS normalizer for Windows, packaged as 2D_AMIS_Tool_eng.zip) - Python source code (2D_AMIS_Tool_eng.py) 3. Results (figures from the article): - AMIS-normalized heatmaps (cellwise, 25×40 grid) - AMIS transformation curve for central cell (row 13, column 20) - Spatial profile (row 13, Croatia) Method summary: The 2D AMIS method normalizes each spatial cell individually using adaptive multi-interval scaling, transforming raw data into a unified [0, 100] scale where 50 represents the reference norm. Unlike min-max or z-score normalization, AMIS accounts for the local density of data distribution. Software notes: - The executable file is provided as .zip due to archiving policies. Unzip before use. - First launch may take 10-30 seconds due to Python initialization. License: - Source code: MIT - Data: CC BY 4.0 Related links: - GitHub: https://github.com/Famimot/2D_AMIS_Tool - Related preprint (SSRN): https://ssrn.com/abstract=6792479
Facebook
TwitterMIT Licensehttps://opensource.org/licenses/MIT
License information was derived automatically
win32k-cot-dataset
200 Chain-of-Thought reasoning examples for Windows kernel vulnerability analysis.Teach your LLM to think like a kernel security researcher, not just pattern-match.
🔗 GitHub Repository: Cooma-sys/win32k-cot-dataset🔗 Hugging Face Dataset: CoomasX/win32k-cot-dataset
What is this?
A hand-curated dataset of PhD-level, Russian-language Chain-of-Thought (CoT) reasoning traces covering Windows kernel (win32k.sys / win32kfull.sys) vulnerability… See the full description on the dataset page: https://huggingface.co/datasets/CoomasX/win32k-cot-dataset.
Facebook
TwitterMIT Licensehttps://opensource.org/licenses/MIT
License information was derived automatically
IndustryBench: Probing the Industrial Knowledge Boundaries of LLMs
💻Github | 📝Paper IndustryBench is a multi-lingual benchmark for evaluating the industrial domain knowledge of large language models. It comprises 2,049 expert-curated QA pairs spanning 12 industrial sectors, with human-reviewed translations in Chinese, English, Russian, and Vietnamese.
Overview
Dimension Details
Total questions 2,049
Languages Chinese (zh), English (en), Russian (ru)… See the full description on the dataset page: https://huggingface.co/datasets/alibaba-multimodal-industrial-ai/IndustryBench.
Facebook
TwitterMIT Licensehttps://opensource.org/licenses/MIT
License information was derived automatically
Description
This dataset is a translation of the Google GoEmotions emotion classification dataset. All features remain unchanged, except for the addition of a new ru_text column containing the translated text in Russian. For the translation process, I used the Deep translator with the Google engine. You can find all the details about translation, raw .csv files and other stuff in this Github repository. For more information also check the official original dataset card.… See the full description on the dataset page: https://huggingface.co/datasets/seara/ru_go_emotions.
Facebook
TwitterAttribution 4.0 (CC BY 4.0)https://creativecommons.org/licenses/by/4.0/
License information was derived automatically
This dataset consists of high-quality parallel English-Russian sentences from medical and scientific literature, specifically curated for training language models in professional medical translation in ophthalmology. The corpus focuses on medical research abstracts, ensuring domain specificity and professional-level translation quality.
The dataset was compiled from medical research abstracts published in peer-reviewed journals, ensuring high-quality source material. Each abstract was professionally translated, making this dataset particularly valuable for training medical translation models.
The dataset is created from the Russian Journal of Clinical Ophthalmology: https://clinopht.com/en/
"Russian Journal of Clinical Ophthalmology" is a peer-reviewed journal publishing clinical and basic science research and other relevant manuscripts that relate to epidemiology, etiology, pathogenesis, diagnosis and treatment of eye diseases.
We implemented a rigorous multi-stage processing pipeline to ensure the highest possible quality of parallel sentences:
Utilized BERTAlign (https://github.com/bfsujason/bertalign) for accurate sentence alignment
Manually validated alignment results for accuracy.
Applied multiple layers of quality control:
* Basic Filtering
* Length ratio validation
* Minimum/maximum length thresholds
* Special character consistency
* COMET QE Scoring:
* Implemented quality estimation using COMET (https://unbabel.github.io/COMET/html/index.html) with the Unbabel/wmt22-cometkiwi-da metric
* Manually estimated the quality thresholds for train and test splits
* Removed pairs below quality threshold
Facebook
TwitterAttribution-ShareAlike 4.0 (CC BY-SA 4.0)https://creativecommons.org/licenses/by-sa/4.0/
License information was derived automatically
GOLOS Annotated by Balalaika
[!IMPORTANT] Official dataset for our INTERSPEECH 2026 paper "A Data-Centric Framework for Addressing Phonetic and Prosodic Challenges in Russian Speech Generative Models" (arXiv:2507.13563). Part of the Balalaika Russian speech data-processing pipeline — code: https://github.com/lab260ru/balalaika. If you use this resource, please cite it.
A curated Russian speech dataset for advanced speech generative tasks.
Overview
GOLOS… See the full description on the dataset page: https://huggingface.co/datasets/lab260/golos_balalaika.
Facebook
TwitterПериод: Сентябрь 1999 - декабрь 2019
Скрипт для скачивания новостей.
Dates: Sept. 1999 - Dec 2019
Script for news downloading.
Facebook
TwitterAttribution-NonCommercial 4.0 (CC BY-NC 4.0)https://creativecommons.org/licenses/by-nc/4.0/
License information was derived automatically
Data Description
This HF data repository contains the Russian Alpaca dataset used in our study of monolingual versus multilingual instruction tuning.
GitHub Paper
Creation
Machine-translated from yahma/alpaca-cleaned into Russian.
Usage
This data is intended to be used for Russian instruction tuning. The dataset has roughly 52K instances in the JSON format. Each instance has an instruction, an output, and an optional input. An example is shown below:
{… See the full description on the dataset page: https://huggingface.co/datasets/pinzhenchen/alpaca-cleaned-ru.
Facebook
TwitterThis is the first Russian news summarization dataset. A paper about this dataset: https://arxiv.org/pdf/2006.11063.pdf Additional files and notebooks: https://github.com/IlyaGusev/gazeta/ Previous datasets for headline generation: https://github.com/RossiyaSegodnya/ria_news_dataset https://www.kaggle.com/yutkin/corpus-of-russian-news-articles-from-lenta
This is the second version of the dataset. The data structure is pretty straightforward. Every line of a file is a JSON object with 5 fields: URL, title, text, summary, and date. The dataset consists of 74126 examples. The first 60964 examples by date are in the training dataset, the proceeding 6369 examples are in the validation dataset, and the remaining 6793 pairs are in the test dataset.
Legal basis for distribution of the dataset: https://www.gazeta.ru/credits.shtml, paragraph 2.1.2. All rights belong to "www.gazeta.ru". This dataset can be removed at the request of the copyright holder. Usage of this dataset is possible only for personal purposes on a non-commercial basis.