45 datasets found
  1. Gazeta Summaries

    • kaggle.com
    zip
    Updated Sep 5, 2021
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Ilya Gusev (2021). Gazeta Summaries [Dataset]. https://www.kaggle.com/phoenix120/gazeta-summaries
    Explore at:
    zip(193749591 bytes)Available download formats
    Dataset updated
    Sep 5, 2021
    Authors
    Ilya Gusev
    Description

    Context

    This is the first Russian news summarization dataset. A paper about this dataset: https://arxiv.org/pdf/2006.11063.pdf Additional files and notebooks: https://github.com/IlyaGusev/gazeta/ Previous datasets for headline generation: https://github.com/RossiyaSegodnya/ria_news_dataset https://www.kaggle.com/yutkin/corpus-of-russian-news-articles-from-lenta

    Content

    This is the second version of the dataset. The data structure is pretty straightforward. Every line of a file is a JSON object with 5 fields: URL, title, text, summary, and date. The dataset consists of 74126 examples. The first 60964 examples by date are in the training dataset, the proceeding 6369 examples are in the validation dataset, and the remaining 6793 pairs are in the test dataset.

    Legal issues

    Legal basis for distribution of the dataset: https://www.gazeta.ru/credits.shtml, paragraph 2.1.2. All rights belong to "www.gazeta.ru". This dataset can be removed at the request of the copyright holder. Usage of this dataset is possible only for personal purposes on a non-commercial basis.

  2. Rossiyskaya Gazeta Papers (Russian legal texts)

    • kaggle.com
    zip
    Updated Feb 13, 2023
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    athugodage (2023). Rossiyskaya Gazeta Papers (Russian legal texts) [Dataset]. https://www.kaggle.com/datasets/athugodage/russian-legal-text-parallel-corpus
    Explore at:
    zip(27158041 bytes)Available download formats
    Dataset updated
    Feb 13, 2023
    Authors
    athugodage
    Area covered
    Russia
    Description

    Context

    This dataset was introduced at READi workshop (LREC-COLING 2024). This dataset was proposed to train a model for Russian legal text simplification. This dataset has already helped to train GPT and T5 models for this task, as well as in Dynamic Topic Modelling task to analyze the history of Russian law from 2009 to 2022 .

    We have collected our data from Rossiyskaya Gazeta website. It's a Russian newspaper published by the Government of Russia. The daily newspaper serves as the official government gazette of the Government of the Russian Federation, publishing government-related affairs such as official decrees, statements and documents of state bodies, the promulgation of newly approved laws, Presidential decrees, and government announcements. Rossiyskaya Gazeta provides legal text descriptions for common people called "comments". But these descriptions are made only for important documents, so while there are hundreds of thousands of legal documents in Russia, only a couple of thousands has a "comment". We used this "comment" as simplified version of the document.

    Content

    Overall there are 2963 pairs of original documents and simplified ones. Dataset contains documents from December 31, 2008 up to November 28, 2022 - thus it contains COVID-19-related laws too.

    Dataset has 5 columns: 1. Название документа (Document Title) 2. Ссылка (Link to the original document) 3. Текст (Original document text) 4. Комментарий РГ (Rossiyskaya Gazeta comment) 5. Дата (Publication date)

    Example of original document text (2nd article in row): https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F11047041%2F625daafe6319319700f754f5ac1d89e4%2Forig.png?generation=1676324147067453&alt=media" alt="">

    Example of Rossiyskaya Gazeta comment: https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F11047041%2F3230f4df60713b417826434ddacc8882%2Fcomm.png?generation=1676324185146892&alt=media" alt="">

    The photo on the headline was taken from Roscosmos official website. `

  3. d

    Ukraine and Russia Conflict Tweet IDs Release v1.3

    • dataone.org
    Updated Nov 8, 2023
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Chen, Emily; Ferrara, Emilio (2023). Ukraine and Russia Conflict Tweet IDs Release v1.3 [Dataset]. http://doi.org/10.7910/DVN/XZSYQO
    Explore at:
    Dataset updated
    Nov 8, 2023
    Dataset provided by
    Harvard Dataverse
    Authors
    Chen, Emily; Ferrara, Emilio
    Area covered
    Russia, Ukraine
    Description

    The repository contains an ongoing collection of tweets IDs associated with the current conflict in Ukraine and Russia, which we commenced collecting on Februrary 22, 2022. To comply with Twitter’s Terms of Service, we are only publicly releasing the Tweet IDs of the collected Tweets. The data is released for non-commercial research use. Note that the compressed files must be first uncompressed in order to use included scripts. This dataset is release v1.3 and is not actively maintained -- the actively maintained dataset can be found here: https://github.com/echen102/ukraine-russia. This release contains Tweet IDs collected from 2/22/22 - 1/08/23. Please refer to the README for more details regarding data, data organization and data usage agreement. This dataset is licensed under the Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International Public License . By using this dataset, you agree to abide by the stipulations in the license, remain in compliance with Twitter’s Terms of Service, and cite the following manuscript: Emily Chen and Emilio Ferrara. 2022. Tweets in Time of Conflict: A Public Dataset Tracking the Twitter Discourse on the War Between Ukraine and Russia. arXiv:cs.SI/2203.07488

  4. ⁠Audio Deepfake in Russian Language Dataset

    • kaggle.com
    zip
    Updated May 20, 2025
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Vladimir Keller (2025). ⁠Audio Deepfake in Russian Language Dataset [Dataset]. https://www.kaggle.com/datasets/vladimirkeller/generated-russian-phrases
    Explore at:
    zip(7880998891 bytes)Available download formats
    Dataset updated
    May 20, 2025
    Authors
    Vladimir Keller
    Description

    Dataset Overview

    This dataset is designed for research on audio deepfake detection, focusing specifically on generated speech in Russian. It contains TTS-generated audio, paired with transcriptions, and a mixed set for real vs fake classification tasks.

    Purpose

    The main goal is to support research on audio deepfake detection in underrepresented languages, especially Russian. The dataset simulates real-world scenarios using multiple state-of-the-art TTS systems to generate fakes and includes clean, real audio data.

    Links

    Generated Audio

    We used three high-quality TTS models to synthesize Russian speech:

    XTTS-v2: Cross-lingual, zero-shot voice cloning with multilingual support.

    Silero TTS: Lightweight, real-time Russian TTS model.

    VITS RU Multispeaker: VITS-based Russian model with speaker variability.

    Real Audio

    For real human speech, we used a part of SOVA dataset, which contains clean Russian utterances recorded by multiple speakers.

  5. H

    Russian Election Data

    • dataverse.harvard.edu
    Updated May 6, 2026
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Georgy Tarasenko; Konstantin Bogatyrev; Nikita Savin (2026). Russian Election Data [Dataset]. http://doi.org/10.7910/DVN/DFGNTP
    Explore at:
    CroissantCroissant is a format for machine-learning datasets. Learn more about this at mlcommons.org/croissant.
    Dataset updated
    May 6, 2026
    Dataset provided by
    Harvard Dataverse
    Authors
    Georgy Tarasenko; Konstantin Bogatyrev; Nikita Savin
    License

    CC0 1.0 Universal Public Domain Dedicationhttps://creativecommons.org/publicdomain/zero/1.0/
    License information was derived automatically

    Area covered
    Russia
    Description

    Russian Election Data: data (.csv, .rds) and codebook. Replication material are available at: https://github.com/georgytarasenko/RED-replication-package

  6. News dataset from Lenta.Ru 2019-2023

    • kaggle.com
    zip
    Updated Jan 8, 2024
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Maria Levchenko (2024). News dataset from Lenta.Ru 2019-2023 [Dataset]. https://www.kaggle.com/datasets/marialevchenko/news-dataset-from-lenta-ru-2019-2023
    Explore at:
    zip(358487534 bytes)Available download formats
    Dataset updated
    Jan 8, 2024
    Authors
    Maria Levchenko
    Description

    Corpus of news articles from Lenta.Ru / Корпус новостей Lenta.Ru

    Supplement to the Lenta.Ru news dataset (until December 2019)

    • Size: 1.23 Gb
    • Articles: 496257
    • Dates: Dec. 2019 - Dec. 2023

    Updated script to download data in current format: GitHub

  7. u

    Data from: The Russian Constructicon database

    • observatorio-cientifico.ua.es
    Updated 2022
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Endresen, Anna; Bast, Radovan; Janda, Laura A.; Zhukova, Valentina; Mordashova, Daria; Rakhilina, Ekaterina; Lyashevskaya, Olga; Lund, Marianne; McDonald, James D.; Tyers, Francis M.; Endresen, Anna; Bast, Radovan; Janda, Laura A.; Zhukova, Valentina; Mordashova, Daria; Rakhilina, Ekaterina; Lyashevskaya, Olga; Lund, Marianne; McDonald, James D.; Tyers, Francis M. (2022). The Russian Constructicon database [Dataset]. https://observatorio-cientifico.ua.es/documentos/67321cbdaea56d4af04839e9
    Explore at:
    Dataset updated
    2022
    Authors
    Endresen, Anna; Bast, Radovan; Janda, Laura A.; Zhukova, Valentina; Mordashova, Daria; Rakhilina, Ekaterina; Lyashevskaya, Olga; Lund, Marianne; McDonald, James D.; Tyers, Francis M.; Endresen, Anna; Bast, Radovan; Janda, Laura A.; Zhukova, Valentina; Mordashova, Daria; Rakhilina, Ekaterina; Lyashevskaya, Olga; Lund, Marianne; McDonald, James D.; Tyers, Francis M.
    Area covered
    Russia
    Description

    The set of over 2,250 files archived here comprises a database of the Russian Constructicon, an open-access electronic resource freely available at https://constructicon.github.io/russian/. The Russian Constructicon is a searchable database of constructions accompanied with thorough descriptions of their properties and annotated illustrative examples.

  8. h

    sova_balalaika

    • huggingface.co
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Lab260, sova_balalaika [Dataset]. https://huggingface.co/datasets/lab260/sova_balalaika
    Explore at:
    Dataset authored and provided by
    Lab260
    Description

    SOVA (Balalaika)

    [!IMPORTANT] Official dataset for our INTERSPEECH 2026 paper "A Data-Centric Framework for Addressing Phonetic and Prosodic Challenges in Russian Speech Generative Models" (arXiv:2507.13563). Part of the Balalaika Russian speech data-processing pipeline — code: https://github.com/lab260ru/balalaika. If you use this resource, please cite it.

    Part of the Balalaika Russian speech data-processing pipeline. See the code repository for details.… See the full description on the dataset page: https://huggingface.co/datasets/lab260/sova_balalaika.

  9. Data from: PoeTree. Poetry Treebanks in Czech, English, French, German,...

    • zenodo.org
    zip
    Updated May 2, 2024
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Petr Plecháč; Petr Plecháč; Robert Kolár; Robert Kolár; Silvie Cinková; Silvie Cinková; Artjoms Šeļa; Artjoms Šeļa; Mirella De Sisto; Mirella De Sisto; Lara Nugues; Lara Nugues; Thomas Haider; Thomas Haider; Benjamin Nagy; Benjamin Nagy; Éliane Delente; Richard Renault; Klemens Bobenhausen; Benjamin Hammerich; Adiel Mittmann; Gábor Palkó; Gábor Palkó; Péter Horváth; Péter Horváth; Borja Navarro Colorado; Borja Navarro Colorado; Pablo Ruiz Fabo; Pablo Ruiz Fabo; Helena Bermúdez Sabel; Helena Bermúdez Sabel; Kirill Korchagin; Kirill Korchagin; Vladimir Plungian; Vladimir Plungian; Dmitri Sitchinava; Dmitri Sitchinava; Éliane Delente; Richard Renault; Klemens Bobenhausen; Benjamin Hammerich; Adiel Mittmann (2024). PoeTree. Poetry Treebanks in Czech, English, French, German, Hungarian, Italian, Portuguese, Russian and Spanish [Dataset]. http://doi.org/10.5281/zenodo.10008459
    Explore at:
    zipAvailable download formats
    Dataset updated
    May 2, 2024
    Dataset provided by
    Zenodohttp://zenodo.org/
    Authors
    Petr Plecháč; Petr Plecháč; Robert Kolár; Robert Kolár; Silvie Cinková; Silvie Cinková; Artjoms Šeļa; Artjoms Šeļa; Mirella De Sisto; Mirella De Sisto; Lara Nugues; Lara Nugues; Thomas Haider; Thomas Haider; Benjamin Nagy; Benjamin Nagy; Éliane Delente; Richard Renault; Klemens Bobenhausen; Benjamin Hammerich; Adiel Mittmann; Gábor Palkó; Gábor Palkó; Péter Horváth; Péter Horváth; Borja Navarro Colorado; Borja Navarro Colorado; Pablo Ruiz Fabo; Pablo Ruiz Fabo; Helena Bermúdez Sabel; Helena Bermúdez Sabel; Kirill Korchagin; Kirill Korchagin; Vladimir Plungian; Vladimir Plungian; Dmitri Sitchinava; Dmitri Sitchinava; Éliane Delente; Richard Renault; Klemens Bobenhausen; Benjamin Hammerich; Adiel Mittmann
    Area covered
    French
    Description

    PoeTree (Poetry Treebanks) is a dataset comprising over 300,000 poems / 84,000,000 tokens in nine languages (Czech, English, French, German, Hungarian, Italian, Portuguese, Spanish, and Russian). Each corpus has been deduplicated, enriched with Universal Dependencies, provided with additional metadata and converted into a unified JSON structure (schema available at https://versologie.cz/poetree/json-schema).

  10. Z

    Tweets containing the keyword 'bucha' during the Russian invasion of Ukraine...

    • data.niaid.nih.gov
    Updated Nov 24, 2023
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Münch, Felix Victor (2023). Tweets containing the keyword 'bucha' during the Russian invasion of Ukraine [Dataset]. https://data.niaid.nih.gov/resources?id=zenodo_10204391
    Explore at:
    Dataset updated
    Nov 24, 2023
    Dataset provided by
    Kessling, Philipp
    Münch, Felix Victor
    Area covered
    Bucha, Russia, Ukraine
    Description

    Comprehensive dataset of Tweets containing the keyword 'bucha' around the Ukraine Invasion in February 2022. The user handle column has been excluded to protect deleted accounts that have not been retweeted or replied to. Tweets have been collected via the Academic API using the Search endpoint in four languages: Languages English search term: "bucha AND lang='en'" German search term: "(bucha OR butscha) AND lang='de'" Russian search term: (Бу́ча OR bucha) AND lang:ru Ukrainian search term: (Бу́ча OR bucha) AND lang:uk Timeframe Data starts 1. March 2022 and ends on 27. May 2023. Collection dates Details on collection dates per Tweet (e.g. to compare with creation dates) as well as the IDs of Tweets for consistency checks can be found here: https://github.com/Leibniz-HBI/ukraine_twitter_data (https://doi.org/10.17605/OSF.IO/RTQXN)

  11. Kyrgyz YouTube comments

    • kaggle.com
    zip
    Updated Apr 12, 2023
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Ruslan Isaev (2023). Kyrgyz YouTube comments [Dataset]. https://www.kaggle.com/datasets/pteacher/kyrgyz-youtube-comments
    Explore at:
    zip(3277443 bytes)Available download formats
    Dataset updated
    Apr 12, 2023
    Authors
    Ruslan Isaev
    License

    Attribution 4.0 (CC BY 4.0)https://creativecommons.org/licenses/by/4.0/
    License information was derived automatically

    Area covered
    YouTube, Kyrgyzstan
    Description

    Kyrgyz comments are still new things to examine. Defining whether comment is toxic or not is a big challenge. We used YouTube API to get comments. Code can be found here - https://github.com/abdulra7ma/ycc

  12. h

    mathematics_dataset

    • huggingface.co
    Updated Apr 2, 2019
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Dmitry Balobin (2019). mathematics_dataset [Dataset]. https://huggingface.co/datasets/d0rj/mathematics_dataset
    Explore at:
    Dataset updated
    Apr 2, 2019
    Authors
    Dmitry Balobin
    License

    Apache License, v2.0https://www.apache.org/licenses/LICENSE-2.0
    License information was derived automatically

    Description

    Mathematical Reasoning Dataset (English & Russian)

    A bilingual collection of synthetic school-level mathematics questions and answers, based on the DeepMind mathematics_dataset generator. This dataset contains two language splits:

    en — the original English data, taken as-is from the official mathematics_dataset-v1.0 release published by Google DeepMind (github.com/google-deepmind/mathematics_dataset). ru — a Russian version generated from scratch with a translated fork of the… See the full description on the dataset page: https://huggingface.co/datasets/d0rj/mathematics_dataset.

  13. H

    Supplementary Materials for "2D AMIS Normalization for Spatial Data"

    • dataverse.harvard.edu
    Updated Jun 9, 2026
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Gennady Kravtsov (2026). Supplementary Materials for "2D AMIS Normalization for Spatial Data" [Dataset]. http://doi.org/10.7910/DVN/J2EB8S
    Explore at:
    CroissantCroissant is a format for machine-learning datasets. Learn more about this at mlcommons.org/croissant.
    Dataset updated
    Jun 9, 2026
    Dataset provided by
    Harvard Dataverse
    Authors
    Gennady Kravtsov
    License

    Attribution 4.0 (CC BY 4.0)https://creativecommons.org/licenses/by/4.0/
    License information was derived automatically

    Description

    This dataset accompanies the article "2D ADAPTIVE MULTI-INTERVAL SCALE (AMIS): METHOD FOR NORMALIZATION AND VISUALIZATION OF SPATIAL DATA" (https://doi.org/10.5281/zenodo.20577673). The dataset contains: 1. Example data: - Reference dataset (32 teams, Russian Championship 2012-13, 25×40 grid, 32,000 observations) - Match data (France vs Croatia, 2018 World Cup final; England vs Germany, 1966 World Cup final) 2. Software: - Executable tool (2D AMIS normalizer for Windows, packaged as 2D_AMIS_Tool_eng.zip) - Python source code (2D_AMIS_Tool_eng.py) 3. Results (figures from the article): - AMIS-normalized heatmaps (cellwise, 25×40 grid) - AMIS transformation curve for central cell (row 13, column 20) - Spatial profile (row 13, Croatia) Method summary: The 2D AMIS method normalizes each spatial cell individually using adaptive multi-interval scaling, transforming raw data into a unified [0, 100] scale where 50 represents the reference norm. Unlike min-max or z-score normalization, AMIS accounts for the local density of data distribution. Software notes: - The executable file is provided as .zip due to archiving policies. Unzip before use. - First launch may take 10-30 seconds due to Python initialization. License: - Source code: MIT - Data: CC BY 4.0 Related links: - GitHub: https://github.com/Famimot/2D_AMIS_Tool - Related preprint (SSRN): https://ssrn.com/abstract=6792479

  14. h

    win32k-cot-dataset

    • huggingface.co
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Sys, win32k-cot-dataset [Dataset]. https://huggingface.co/datasets/CoomasX/win32k-cot-dataset
    Explore at:
    Authors
    Sys
    License

    MIT Licensehttps://opensource.org/licenses/MIT
    License information was derived automatically

    Description

    win32k-cot-dataset

    200 Chain-of-Thought reasoning examples for Windows kernel vulnerability analysis.Teach your LLM to think like a kernel security researcher, not just pattern-match.

    🔗 GitHub Repository: Cooma-sys/win32k-cot-dataset🔗 Hugging Face Dataset: CoomasX/win32k-cot-dataset

      What is this?
    

    A hand-curated dataset of PhD-level, Russian-language Chain-of-Thought (CoT) reasoning traces covering Windows kernel (win32k.sys / win32kfull.sys) vulnerability… See the full description on the dataset page: https://huggingface.co/datasets/CoomasX/win32k-cot-dataset.

  15. h

    IndustryBench

    • huggingface.co
    Updated May 11, 2026
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    1688 multimodal & industrial AI (2026). IndustryBench [Dataset]. https://huggingface.co/datasets/alibaba-multimodal-industrial-ai/IndustryBench
    Explore at:
    Dataset updated
    May 11, 2026
    Dataset authored and provided by
    1688 multimodal & industrial AI
    License

    MIT Licensehttps://opensource.org/licenses/MIT
    License information was derived automatically

    Description

    IndustryBench: Probing the Industrial Knowledge Boundaries of LLMs

    💻Github | 📝Paper IndustryBench is a multi-lingual benchmark for evaluating the industrial domain knowledge of large language models. It comprises 2,049 expert-curated QA pairs spanning 12 industrial sectors, with human-reviewed translations in Chinese, English, Russian, and Vietnamese.

      Overview
    

    Dimension Details

    Total questions 2,049

    Languages Chinese (zh), English (en), Russian (ru)… See the full description on the dataset page: https://huggingface.co/datasets/alibaba-multimodal-industrial-ai/IndustryBench.

  16. h

    ru_go_emotions

    • huggingface.co
    Updated Aug 26, 2023
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Vyacheslav Litvinov (2023). ru_go_emotions [Dataset]. https://huggingface.co/datasets/seara/ru_go_emotions
    Explore at:
    CroissantCroissant is a format for machine-learning datasets. Learn more about this at mlcommons.org/croissant.
    Dataset updated
    Aug 26, 2023
    Authors
    Vyacheslav Litvinov
    License

    MIT Licensehttps://opensource.org/licenses/MIT
    License information was derived automatically

    Description

    Description

    This dataset is a translation of the Google GoEmotions emotion classification dataset. All features remain unchanged, except for the addition of a new ru_text column containing the translated text in Russian. For the translation process, I used the Deep translator with the Google engine. You can find all the details about translation, raw .csv files and other stuff in this Github repository. For more information also check the official original dataset card.… See the full description on the dataset page: https://huggingface.co/datasets/seara/ru_go_emotions.

  17. Ophthalmology Russian/English Translations

    • kaggle.com
    zip
    Updated Jan 2, 2025
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    che_shr_cat (2025). Ophthalmology Russian/English Translations [Dataset]. https://www.kaggle.com/datasets/cheshrcat/ru-medical-texts-ophtalmology
    Explore at:
    zip(620314 bytes)Available download formats
    Dataset updated
    Jan 2, 2025
    Authors
    che_shr_cat
    License

    Attribution 4.0 (CC BY 4.0)https://creativecommons.org/licenses/by/4.0/
    License information was derived automatically

    Description

    Overview

    This dataset consists of high-quality parallel English-Russian sentences from medical and scientific literature, specifically curated for training language models in professional medical translation in ophthalmology. The corpus focuses on medical research abstracts, ensuring domain specificity and professional-level translation quality.

    Dataset Statistics

    • Total sentence pairs: 3304 in train, 169 in test
    • Glossary size: 1211 unique Russian terms
    • Domain coverage: Medical research abstracts, clinical observations, treatment methodologies
    • Quality threshold: COMET QE score > 0.75 for the test set and 0.73 for the train set.

    Data Sources

    The dataset was compiled from medical research abstracts published in peer-reviewed journals, ensuring high-quality source material. Each abstract was professionally translated, making this dataset particularly valuable for training medical translation models.

    The dataset is created from the Russian Journal of Clinical Ophthalmology: https://clinopht.com/en/

    "Russian Journal of Clinical Ophthalmology" is a peer-reviewed journal publishing clinical and basic science research and other relevant manuscripts that relate to epidemiology, etiology, pathogenesis, diagnosis and treatment of eye diseases.

    Data Processing Pipeline

    We implemented a rigorous multi-stage processing pipeline to ensure the highest possible quality of parallel sentences:

    Initial Data Collection and Cleaning

    • Extracted titles and abstracts from all the available issues of the "Russian Journal of Clinical Ophthalmology, separately from Russian and English pages
    • Cleaned Russian and English abstracts by removing:
      • Author information
      • Institutional affiliations
      • Publication metadata
      • References
    • Preserved core scientific content while eliminating auxiliary information
    • Extracted keyword sections with subsequence term splitting, alignment, manual quality check, and filtering

    Text Normalization and Preprocessing

    • Implemented specialized text normalization:
      • Handling hyphenation at line breaks
      • Standardizing formatting

    Advanced Sentence Alignment

    Utilized BERTAlign (https://github.com/bfsujason/bertalign) for accurate sentence alignment

    • Benefits over traditional length-based alignment:
      • Semantic awareness
      • Better handling of content reordering
      • Robust to translation variations

    Manually validated alignment results for accuracy.

    Quality Filtering

    Applied multiple layers of quality control: * Basic Filtering * Length ratio validation * Minimum/maximum length thresholds * Special character consistency * COMET QE Scoring: * Implemented quality estimation using COMET (https://unbabel.github.io/COMET/html/index.html) with the Unbabel/wmt22-cometkiwi-da metric * Manually estimated the quality thresholds for train and test splits * Removed pairs below quality threshold

    Quality Assurance Insights

    Lessons Learned in Dataset Creation

    • The importance of domain-specific preprocessing:
      • Medical texts require specialized handling
      • Standard NLP tools often need adaptation
    • Quality vs. Quantity trade-off:
      • Strict filtering significantly reduces dataset size
      • Higher quality data leads to better model performance
    • The value of semantic alignment:
      • Traditional length-based alignment often fails for medical texts
      • Semantic alignment tools better handle domain-specific variations

    Best Practices for Dataset Creation

    1. Start with high-quality sources
    2. Implement domain-specific preprocessing
    3. Use advanced alignment techniques
    4. Apply multiple layers of quality filtering
    5. Validate with domain experts when possible
    6. Document all processing decisions

    Usage Recommendations

    For Model Training

    • Consider domain adaptation techniques
    • Implement terminology-aware evaluation metrics
    • Use domain-specific augmentation if needed

    For Dataset Extension

    • Follow the documented cleaning pipeline
    • Maintain consistent quality thresholds
    • Validate new additions against existing quality metrics

    Limitations and Considerations

    • Dataset focuses on formal medical literature in ophthalmology domain
    • Does not cover other medical domains
    • May not cover informal medical communication
    • Limited to research abstract style and format
    • Specific to professional medical translation context
    • Some sentences in parallel pairs still use different styles (e.g. more or less verbose, etc), may require additional style normalization
  18. h

    golos_balalaika

    • huggingface.co
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Lab260, golos_balalaika [Dataset]. https://huggingface.co/datasets/lab260/golos_balalaika
    Explore at:
    Dataset authored and provided by
    Lab260
    License

    Attribution-ShareAlike 4.0 (CC BY-SA 4.0)https://creativecommons.org/licenses/by-sa/4.0/
    License information was derived automatically

    Description

    GOLOS Annotated by Balalaika

    [!IMPORTANT] Official dataset for our INTERSPEECH 2026 paper "A Data-Centric Framework for Addressing Phonetic and Prosodic Challenges in Russian Speech Generative Models" (arXiv:2507.13563). Part of the Balalaika Russian speech data-processing pipeline — code: https://github.com/lab260ru/balalaika. If you use this resource, please cite it.

    A curated Russian speech dataset for advanced speech generative tasks.

      Overview
    

    GOLOS… See the full description on the dataset page: https://huggingface.co/datasets/lab260/golos_balalaika.

  19. News dataset from Lenta.Ru

    • kaggle.com
    zip
    Updated Dec 14, 2019
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    DmitryYutkin (2019). News dataset from Lenta.Ru [Dataset]. https://www.kaggle.com/yutkin/corpus-of-russian-news-articles-from-lenta
    Explore at:
    zip(611887272 bytes)Available download formats
    Dataset updated
    Dec 14, 2019
    Authors
    DmitryYutkin
    Description

    Корпус новостей с Lenta.Ru

    • Размер: 2 Гб
    • Количество новостей: 800K+
    • Период: Сентябрь 1999 - декабрь 2019

    • Скрипт для скачивания новостей.

    (Eng) Corpus of news articles from Lenta.Ru

    • Size: 2 Gb
    • News articles: 800K+
    • Dates: Sept. 1999 - Dec 2019

    • Script for news downloading.

    Скачать / Download

  20. h

    alpaca-cleaned-ru

    • huggingface.co
    Updated Sep 16, 2023
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Pinzhen Chen (2023). alpaca-cleaned-ru [Dataset]. https://huggingface.co/datasets/pinzhenchen/alpaca-cleaned-ru
    Explore at:
    CroissantCroissant is a format for machine-learning datasets. Learn more about this at mlcommons.org/croissant.
    Dataset updated
    Sep 16, 2023
    Authors
    Pinzhen Chen
    License

    Attribution-NonCommercial 4.0 (CC BY-NC 4.0)https://creativecommons.org/licenses/by-nc/4.0/
    License information was derived automatically

    Description

    Data Description

    This HF data repository contains the Russian Alpaca dataset used in our study of monolingual versus multilingual instruction tuning.

    GitHub Paper

      Creation
    

    Machine-translated from yahma/alpaca-cleaned into Russian.

      Usage
    

    This data is intended to be used for Russian instruction tuning. The dataset has roughly 52K instances in the JSON format. Each instance has an instruction, an output, and an optional input. An example is shown below:

    {… See the full description on the dataset page: https://huggingface.co/datasets/pinzhenchen/alpaca-cleaned-ru.

Share
FacebookFacebook
TwitterTwitter
Email
Click to copy link
Link copied
Close
Cite
Ilya Gusev (2021). Gazeta Summaries [Dataset]. https://www.kaggle.com/phoenix120/gazeta-summaries
Organization logo

Gazeta Summaries

Russian News Summarization Dataset

Explore at:
zip(193749591 bytes)Available download formats
Dataset updated
Sep 5, 2021
Authors
Ilya Gusev
Description

Context

This is the first Russian news summarization dataset. A paper about this dataset: https://arxiv.org/pdf/2006.11063.pdf Additional files and notebooks: https://github.com/IlyaGusev/gazeta/ Previous datasets for headline generation: https://github.com/RossiyaSegodnya/ria_news_dataset https://www.kaggle.com/yutkin/corpus-of-russian-news-articles-from-lenta

Content

This is the second version of the dataset. The data structure is pretty straightforward. Every line of a file is a JSON object with 5 fields: URL, title, text, summary, and date. The dataset consists of 74126 examples. The first 60964 examples by date are in the training dataset, the proceeding 6369 examples are in the validation dataset, and the remaining 6793 pairs are in the test dataset.

Legal issues

Legal basis for distribution of the dataset: https://www.gazeta.ru/credits.shtml, paragraph 2.1.2. All rights belong to "www.gazeta.ru". This dataset can be removed at the request of the copyright holder. Usage of this dataset is possible only for personal purposes on a non-commercial basis.

Search
Clear search
Close search
Google apps
Main menu