63 datasets found
  1. g

    Datasets for evaluation of keyword extraction in Russian

    • github.com
    Updated Jun 11, 2018
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Mikhail Nefedov (2018). Datasets for evaluation of keyword extraction in Russian [Dataset]. https://github.com/mannefedov/ru_kw_eval_datasets
    Explore at:
    Dataset updated
    Jun 11, 2018
    Authors
    Mikhail Nefedov
    Description

    Datasets for evaluation of keyword extraction in Russian

  2. d

    Ukraine and Russia Conflict Tweet IDs Release v1.3

    • dataone.org
    Updated Nov 8, 2023
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Chen, Emily; Ferrara, Emilio (2023). Ukraine and Russia Conflict Tweet IDs Release v1.3 [Dataset]. http://doi.org/10.7910/DVN/XZSYQO
    Explore at:
    Dataset updated
    Nov 8, 2023
    Dataset provided by
    Harvard Dataverse
    Authors
    Chen, Emily; Ferrara, Emilio
    Description

    The repository contains an ongoing collection of tweets IDs associated with the current conflict in Ukraine and Russia, which we commenced collecting on Februrary 22, 2022. To comply with Twitter’s Terms of Service, we are only publicly releasing the Tweet IDs of the collected Tweets. The data is released for non-commercial research use. Note that the compressed files must be first uncompressed in order to use included scripts. This dataset is release v1.3 and is not actively maintained -- the actively maintained dataset can be found here: https://github.com/echen102/ukraine-russia. This release contains Tweet IDs collected from 2/22/22 - 1/08/23. Please refer to the README for more details regarding data, data organization and data usage agreement. This dataset is licensed under the Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International Public License . By using this dataset, you agree to abide by the stipulations in the license, remain in compliance with Twitter’s Terms of Service, and cite the following manuscript: Emily Chen and Emilio Ferrara. 2022. Tweets in Time of Conflict: A Public Dataset Tracking the Twitter Discourse on the War Between Ukraine and Russia. arXiv:cs.SI/2203.07488

  3. Gazeta Summaries

    • kaggle.com
    zip
    Updated Sep 5, 2021
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Ilya Gusev (2021). Gazeta Summaries [Dataset]. https://www.kaggle.com/phoenix120/gazeta-summaries
    Explore at:
    zip(193749591 bytes)Available download formats
    Dataset updated
    Sep 5, 2021
    Authors
    Ilya Gusev
    Description

    Context

    This is the first Russian news summarization dataset. A paper about this dataset: https://arxiv.org/pdf/2006.11063.pdf Additional files and notebooks: https://github.com/IlyaGusev/gazeta/ Previous datasets for headline generation: https://github.com/RossiyaSegodnya/ria_news_dataset https://www.kaggle.com/yutkin/corpus-of-russian-news-articles-from-lenta

    Content

    This is the second version of the dataset. The data structure is pretty straightforward. Every line of a file is a JSON object with 5 fields: URL, title, text, summary, and date. The dataset consists of 74126 examples. The first 60964 examples by date are in the training dataset, the proceeding 6369 examples are in the validation dataset, and the remaining 6793 pairs are in the test dataset.

    Legal issues

    Legal basis for distribution of the dataset: https://www.gazeta.ru/credits.shtml, paragraph 2.1.2. All rights belong to "www.gazeta.ru". This dataset can be removed at the request of the copyright holder. Usage of this dataset is possible only for personal purposes on a non-commercial basis.

  4. h

    russian_super_glue

    • huggingface.co
    • opendatalab.com
    Updated Oct 29, 2020
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Natural Language Processing in Russian (2020). russian_super_glue [Dataset]. https://huggingface.co/datasets/RussianNLP/russian_super_glue
    Explore at:
    Dataset updated
    Oct 29, 2020
    Dataset authored and provided by
    Natural Language Processing in Russian
    License

    MIT Licensehttps://opensource.org/licenses/MIT
    License information was derived automatically

    Description

    Recent advances in the field of universal language models and transformers require the development of a methodology for their broad diagnostics and testing for general intellectual skills - detection of natural language inference, commonsense reasoning, ability to perform simple logical operations regardless of text subject or lexicon. For the first time, a benchmark of nine tasks, collected and organized analogically to the SuperGLUE methodology, was developed from scratch for the Russian language. We provide baselines, human level evaluation, an open-source framework for evaluating models and an overall leaderboard of transformer models for the Russian language.

  5. Automated WarSpotting Equipment Losses

    • kaggle.com
    zip
    Updated Jun 2, 2026
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Zsolt Lazar (2026). Automated WarSpotting Equipment Losses [Dataset]. https://www.kaggle.com/datasets/zsoltlazar/automated-warspotting-equipment-losses
    Explore at:
    zip(543026 bytes)Available download formats
    Dataset updated
    Jun 2, 2026
    Authors
    Zsolt Lazar
    License

    MIT Licensehttps://opensource.org/licenses/MIT
    License information was derived automatically

    Description

    This dataset contains Russian military equipment losses collected from the open-source WarSpotting API. It is automatically updated multiple times per week using a Python scraper running on GitHub Actions.

    The data covers: Full historical scans updated weekly Incremental 30-day scans updated thrice weekly Precise geographic coordinates for equipment loss Equipment type and category details Dates of loss and related metadata

    This dataset is designed for researchers, analysts, and developers interested in: Open-source intelligence (OSINT) Conflict monitoring and analysis Machine learning model training Geospatial visualization of battlefield losses

    The scraper and automation tools powering this dataset are fully open-source and available on GitHub: https://github.com/lazar-bit/automated-warspotting-scraper

  6. h

    Data from: Balalaika

    • huggingface.co
    Updated Jul 17, 2025
    + more versions
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Lab260 (2025). Balalaika [Dataset]. https://huggingface.co/datasets/lab260/Balalaika
    Explore at:
    Dataset updated
    Jul 17, 2025
    Dataset authored and provided by
    Lab260
    License

    Attribution-NonCommercial-NoDerivs 4.0 (CC BY-NC-ND 4.0)https://creativecommons.org/licenses/by-nc-nd/4.0/
    License information was derived automatically

    Description

    A Data-Centric Framework for Addressing Phonetic and Prosodic Challenges in Russian Speech Generative Models

    [!IMPORTANT] Official dataset for our INTERSPEECH 2026 paper "A Data-Centric Framework for Addressing Phonetic and Prosodic Challenges in Russian Speech Generative Models" (arXiv:2507.13563). Part of the Balalaika Russian speech data-processing pipeline — code: https://github.com/lab260ru/balalaika. If you use this resource, please cite it.

    Paper | Code Russian… See the full description on the dataset page: https://huggingface.co/datasets/lab260/Balalaika.

  7. Z

    Database of Russian names, surnames and midnames for gender identification

    • data.niaid.nih.gov
    Updated Jan 24, 2020
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Ivan Begtin (2020). Database of Russian names, surnames and midnames for gender identification [Dataset]. https://data.niaid.nih.gov/resources?id=zenodo_2747010
    Explore at:
    Dataset updated
    Jan 24, 2020
    Dataset authored and provided by
    Ivan Begtin
    License

    Attribution 4.0 (CC BY 4.0)https://creativecommons.org/licenses/by/4.0/
    License information was derived automatically

    Description

    Database of names, surnames and midnames across the Russian federation used as source to teach algorithms for gender identification by fullname. Dataset prepared for MongoDB database. It has MongoDB dump and dump of tables as JSON lines files. Used in gender identification and fullname parsing software https://github.com/datacoon/russiannames Available under Creative Commons CC-BY SA by default.

  8. ⁠Audio Deepfake in Russian Language Dataset

    • kaggle.com
    zip
    Updated May 20, 2025
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Vladimir Keller (2025). ⁠Audio Deepfake in Russian Language Dataset [Dataset]. https://www.kaggle.com/datasets/vladimirkeller/generated-russian-phrases/code
    Explore at:
    zip(7880998891 bytes)Available download formats
    Dataset updated
    May 20, 2025
    Authors
    Vladimir Keller
    Description

    Dataset Overview

    This dataset is designed for research on audio deepfake detection, focusing specifically on generated speech in Russian. It contains TTS-generated audio, paired with transcriptions, and a mixed set for real vs fake classification tasks.

    Purpose

    The main goal is to support research on audio deepfake detection in underrepresented languages, especially Russian. The dataset simulates real-world scenarios using multiple state-of-the-art TTS systems to generate fakes and includes clean, real audio data.

    Links

    Generated Audio

    We used three high-quality TTS models to synthesize Russian speech:

    XTTS-v2: Cross-lingual, zero-shot voice cloning with multilingual support.

    Silero TTS: Lightweight, real-time Russian TTS model.

    VITS RU Multispeaker: VITS-based Russian model with speaker variability.

    Real Audio

    For real human speech, we used a part of SOVA dataset, which contains clean Russian utterances recorded by multiple speakers.

  9. Data from: MiDe22: An Annotated Multi-Event Tweet Dataset for Misinformation...

    • zenodo.org
    • data.niaid.nih.gov
    zip
    Updated Jun 14, 2023
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Cagri Toraman; Cagri Toraman; Oguzhan Ozcelik; Furkan Şahinuç; Fazli Can; Oguzhan Ozcelik; Furkan Şahinuç; Fazli Can (2023). MiDe22: An Annotated Multi-Event Tweet Dataset for Misinformation Detection [Dataset]. http://doi.org/10.5281/zenodo.8032136
    Explore at:
    zipAvailable download formats
    Dataset updated
    Jun 14, 2023
    Dataset provided by
    Zenodohttp://zenodo.org/
    Authors
    Cagri Toraman; Cagri Toraman; Oguzhan Ozcelik; Furkan Şahinuç; Fazli Can; Oguzhan Ozcelik; Furkan Şahinuç; Fazli Can
    License

    Attribution 4.0 (CC BY 4.0)https://creativecommons.org/licenses/by/4.0/
    License information was derived automatically

    Description

    The dataset is composed of 10,348 tweets: 5,284 for English and 5,064 for Turkish. Tweets in the dataset are human-annotated in terms of "false", "true", or "other". The dataset covers multiple topics: the Russia-Ukraine war, COVID-19 pandemic, Refugees, and additional miscellaneous events. The details can be found at https://github.com/avaapm/mide22

  10. H

    Russian Election Data

    • dataverse.harvard.edu
    • search.dataone.org
    Updated May 6, 2026
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Georgy Tarasenko; Konstantin Bogatyrev; Nikita Savin (2026). Russian Election Data [Dataset]. http://doi.org/10.7910/DVN/DFGNTP
    Explore at:
    CroissantCroissant is a format for machine-learning datasets. Learn more about this at mlcommons.org/croissant.
    Dataset updated
    May 6, 2026
    Dataset provided by
    Harvard Dataverse
    Authors
    Georgy Tarasenko; Konstantin Bogatyrev; Nikita Savin
    License

    CC0 1.0 Universal Public Domain Dedicationhttps://creativecommons.org/publicdomain/zero/1.0/
    License information was derived automatically

    Area covered
    Russia
    Description

    Russian Election Data: data (.csv, .rds) and codebook. Replication material are available at: https://github.com/georgytarasenko/RED-replication-package

  11. u

    Data from: The Russian Constructicon database

    • observatorio-cientifico.ua.es
    • dataverse.no
    • +1more
    Updated 2022
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Endresen, Anna; Bast, Radovan; Janda, Laura A.; Zhukova, Valentina; Mordashova, Daria; Rakhilina, Ekaterina; Lyashevskaya, Olga; Lund, Marianne; McDonald, James D.; Tyers, Francis M.; Endresen, Anna; Bast, Radovan; Janda, Laura A.; Zhukova, Valentina; Mordashova, Daria; Rakhilina, Ekaterina; Lyashevskaya, Olga; Lund, Marianne; McDonald, James D.; Tyers, Francis M. (2022). The Russian Constructicon database [Dataset]. https://observatorio-cientifico.ua.es/documentos/67321cbdaea56d4af04839e9
    Explore at:
    Dataset updated
    2022
    Authors
    Endresen, Anna; Bast, Radovan; Janda, Laura A.; Zhukova, Valentina; Mordashova, Daria; Rakhilina, Ekaterina; Lyashevskaya, Olga; Lund, Marianne; McDonald, James D.; Tyers, Francis M.; Endresen, Anna; Bast, Radovan; Janda, Laura A.; Zhukova, Valentina; Mordashova, Daria; Rakhilina, Ekaterina; Lyashevskaya, Olga; Lund, Marianne; McDonald, James D.; Tyers, Francis M.
    Area covered
    Russia
    Description

    The set of over 2,250 files archived here comprises a database of the Russian Constructicon, an open-access electronic resource freely available at https://constructicon.github.io/russian/. The Russian Constructicon is a searchable database of constructions accompanied with thorough descriptions of their properties and annotated illustrative examples.

  12. Z

    Emoji Gestures in Russian Tweets: Moscow

    • data.niaid.nih.gov
    Updated May 18, 2022
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Marina Zhukova (2022). Emoji Gestures in Russian Tweets: Moscow [Dataset]. https://data.niaid.nih.gov/resources?id=zenodo_5800199
    Explore at:
    Dataset updated
    May 18, 2022
    Dataset authored and provided by
    Marina Zhukova
    License

    Attribution 4.0 (CC BY 4.0)https://creativecommons.org/licenses/by/4.0/
    License information was derived automatically

    Area covered
    Moscow, Russia
    Description

    The dataset consists of 48 838 tweets each of them contains one of the 31 gesture emoji (different hand configurations) and its skin tone modifier options (e.g. 🙏🙏🏿🙏🏾🙏🏽🙏🏼🙏🏻), and posted within 50km from Moscow, Russia, in Russian, during May-August 2021. The dataset can be used to investigate the use of gesture emoji by Russian users of the Twitter platform. Python libraries used for collecting tweets and preprocessing: tweepy, re, preprocessor, emoji, regex, string, nltk. The dataset contains 11 columns: preprocessed preprocessed text of the tweet (4 steps) all_emoji lists all emoji in a given tweet hashtags lists all hashtags in a given tweet user_encoded encoded Twitter user name: the first 3 characters of the user name and the first 3 characters of the user's location location_encoded location of the user: "moscow", "moscow_region", or "other" mention_present checks whether each tweet contains mentions url_present checks whether each tweet contains url preprocess_tweet preprocessing step 1: tokenizing mentions, urls, and hashtags lowercase_tweet preprocessing step 2: lowercasing remove_punct_tweet preprocessing step 3: removing punctuation tokenize_tweet preprocessing step 4: tokenizing The further information on the research project can be found here: https://github.com/mzhukovaucsb/emoji_gestures/

  13. h

    alpaca-cleaned-ru

    • huggingface.co
    Updated Sep 16, 2023
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Pinzhen Chen (2023). alpaca-cleaned-ru [Dataset]. https://huggingface.co/datasets/pinzhenchen/alpaca-cleaned-ru
    Explore at:
    Dataset updated
    Sep 16, 2023
    Authors
    Pinzhen Chen
    License

    Attribution-NonCommercial 4.0 (CC BY-NC 4.0)https://creativecommons.org/licenses/by-nc/4.0/
    License information was derived automatically

    Description

    Data Description

    This HF data repository contains the Russian Alpaca dataset used in our study of monolingual versus multilingual instruction tuning.

    GitHub Paper

      Creation
    

    Machine-translated from yahma/alpaca-cleaned into Russian.

      Usage
    

    This data is intended to be used for Russian instruction tuning. The dataset has roughly 52K instances in the JSON format. Each instance has an instruction, an output, and an optional input. An example is shown below:

    {… See the full description on the dataset page: https://huggingface.co/datasets/pinzhenchen/alpaca-cleaned-ru.

  14. Russia - Ukraine War Tweets

    • kaggle.com
    zip
    Updated Nov 29, 2022
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    The Devastator (2022). Russia - Ukraine War Tweets [Dataset]. https://www.kaggle.com/datasets/thedevastator/invasion-of-ukraine-tweets-and-user-features
    Explore at:
    zip(19340125 bytes)Available download formats
    Dataset updated
    Nov 29, 2022
    Authors
    The Devastator
    License

    https://creativecommons.org/publicdomain/zero/1.0/https://creativecommons.org/publicdomain/zero/1.0/

    Area covered
    Ukraine, Russia
    Description

    Russia - Ukraine War Tweets

    Tweets about ongoing Russia - Ukraine war

    By [source]

    About this dataset

    This dataset consists of tweets relating to the Russian invasion of Ukraine that were scraped for this study. Only tweets of which user features were available are included in the dataset. The tweets and corresponding user features can be rehydrated using the Twitter API. However, it could be that some tweets or users might be deleted or put on private and are therefore no longer available. Moreover, user and tweet features might change over time

    More Datasets

    For more datasets, click here.

    Featured Notebooks

    • 🚨 Your notebook can be here! 🚨!

    How to use the dataset

    The dataset consists of tweets relating to the Russian invasion of Ukraine that were scraped for this study. Only tweets of which user features were available are included in the dataset. The tweets and corresponding user features can be rehydrated using the Twitter API. However, it could be that some tweets or users might be deleted or put on private and are therefore no longer available. Moreover, user and tweet features might change over time This dataset can be used to study the change in sentiment, and topics over time as the war continues

    Research Ideas

    • Find out which tweets are most popular among people interested in the Russian invasion of Ukraine
    • Identify which user attributes are associated with tweets about the Russian invasion of Ukraine
    • Study the change in sentiment and public opinion on the war as events unfold.

    Acknowledgements

    If you use this dataset in your research, please credit the original authors. Data Source

    License

    License: CC0 1.0 Universal (CC0 1.0) - Public Domain Dedication No Copyright - You can copy, modify, distribute and perform the work, even for commercial purposes, all without asking permission. See Other Information.

    Columns

    File: after_invasion_tweetids.csv | Column name | Description | |:--------------|:-----------------------| | id | The tweet id. (String) |

    File: before_invasion_tweetids.csv | Column name | Description | |:--------------|:-----------------------| | id | The tweet id. (String) |

    Acknowledgements

    If you use this dataset in your research, please credit the original authors. If you use this dataset in your research, please credit .

  15. Z

    Supplementary code and data for the paper: 'The fall of genres that did not...

    • data.niaid.nih.gov
    Updated Dec 7, 2023
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Šeļa, Artjoms (2023). Supplementary code and data for the paper: 'The fall of genres that did not happen: formalising history of the "universal" semantics of Russian iambic tetrameter' [Dataset]. https://data.niaid.nih.gov/resources?id=zenodo_7958273
    Explore at:
    Dataset updated
    Dec 7, 2023
    Dataset provided by
    Šeļa, Artjoms
    Martynenko, Antonina
    Description

    The dataset provides preprocessed data and the full code used in the paper 'The fall of genres that did not happen: formalising history of the "universal" semantics of Russian iambic tetrameter'. The code can be also be accessed as rendered notebooks on Github. The dataset is structured as follows: data/ : This folder contains preprocessed data,including a sampled corpus of periodicals and a document-term matrix used for topic modelling; scr/ : The code used for the analysis, with separate scripts for figures; plots/ : The figures used in the paper, which correspond to the aforementioned code.

  16. h

    Multilingual_Speech_Dataset

    • huggingface.co
    Updated Feb 13, 2025
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Institute of Smart Systems and Artificial Intelligence, Nazarbayev University (2025). Multilingual_Speech_Dataset [Dataset]. https://huggingface.co/datasets/issai/Multilingual_Speech_Dataset
    Explore at:
    Dataset updated
    Feb 13, 2025
    Dataset authored and provided by
    Institute of Smart Systems and Artificial Intelligence, Nazarbayev University
    License

    MIT Licensehttps://opensource.org/licenses/MIT
    License information was derived automatically

    Description

    Multilingual Speech Dataset

    Paper: A Study of Multilingual End-to-End Speech Recognition for Kazakh, Russian, and English Repository: https://github.com/IS2AI/MultilingualASR Description: This repository provides the dataset used in the paper "A Study of Multilingual End-to-End Speech Recognition for Kazakh, Russian, and English". The paper focuses on training a single end-to-end (E2E) ASR model for Kazakh, Russian, and English, comparing monolingual and multilingual approaches… See the full description on the dataset page: https://huggingface.co/datasets/issai/Multilingual_Speech_Dataset.

  17. Z

    RuBQ

    • data-staging.niaid.nih.gov
    Updated May 21, 2020
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Vladislav Korablinov (2020). RuBQ [Dataset]. https://data-staging.niaid.nih.gov/resources?id=zenodo_3835912
    Explore at:
    Dataset updated
    May 21, 2020
    Dataset provided by
    Pavel Braslavski
    Vladislav Korablinov
    License

    Attribution-ShareAlike 4.0 (CC BY-SA 4.0)https://creativecommons.org/licenses/by-sa/4.0/
    License information was derived automatically

    Description

    We present RuBQ (pronounced [`rubik]) -- Russian Knowledge Base Questions, a KBQA dataset that consists of 1,500 Russian questions of varying complexity along with their English machine translations, corresponding SPARQL queries, answers, as well as a subset of Wikidata covering entities with Russian labels. To the best of our knowledge, this is the first Russian KBQA and semantic parsing dataset.

    The dataset is thought to be used as a development and test sets in cross-lingual transfer, few-shot learning, or learning with synthetic data scenarios. Detailed information about RuBQ can be found on the Github page.

  18. Data from: CTLA4 gene polymorphisms are associated with, and linked to,...

    • healthdata.gov
    • data.ko.virginia.gov
    • +7more
    csv, xlsx, xml
    Updated Jul 13, 2025
    + more versions
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    (2025). CTLA4 gene polymorphisms are associated with, and linked to, insulin-dependent diabetes mellitus in a Russian population [Dataset]. https://healthdata.gov/NIH/CTLA4-gene-polymorphisms-are-associated-with-and-l/kzz4-sxw3
    Explore at:
    csv, xlsx, xmlAvailable download formats
    Dataset updated
    Jul 13, 2025
    Description

    Background The association between the human cytotoxic T lymphocyte-associated antigen-4 (CTLA4) gene and insulin-dependent diabetes mellitus (IDDM) is unclear in populations. We therefore investigated whether the gene conferred susceptibility to IDDM in a Russian population. We studied two polymorphic regions of the CTLA4 gene, the codon 17 dimorphism and the (AT)n microsatellite marker in the 3' untranslated region in 56 discordant sibling pairs and in 33 identical by descent (IBD) affected sibships.

       Results
       The Ala17 allele of the CTLA4 gene was preferentially transmitted from parents to diabetic offspring (p < 0.0001) as shown by the combined transmission/disequlibrium test (TDT) and sib TDT (S-TDT) analysis. A significant difference between diabetic and non-diabetic offspring was also observed for the transmission of alleles 17, 20, and 26 of the dinucleotide microsatellite. Allele 17 was transmitted significantly more frequently to affected offspring than to other children (p = 0.0112) whereas alleles 20 and 26 were transmitted preferentially to non-diabetic sibs (p = 0.045 and 0.00068 respectively). A nonrandom excess of the Ala17 CTLA4 molecular variant (maximum logarithm of odds score (MLS) of 3.26) and allele 17 of the dinucleotide marker (MLS = 3.14) was observed in IBD-affected sibling pairs.
    
    
       Conclusion
       The CTLA4 gene is strongly associated with, and linked to IDDM in a Russian population.
    
  19. Z

    Comprehensive Collection of English, German, Russian and Ukrainian Tweets...

    • data.niaid.nih.gov
    Updated Apr 11, 2024
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Münch, Felix Victor (2024). Comprehensive Collection of English, German, Russian and Ukrainian Tweets Containing the Word or Hashtag Ukraine During the Russian Invasion, February 2022 until May 2023 [Dataset]. https://data.niaid.nih.gov/resources?id=zenodo_10930821
    Explore at:
    Dataset updated
    Apr 11, 2024
    Dataset provided by
    Münch, Felix Victor
    Kessling, Philipp
    Area covered
    Ukraine, Russia
    Description

    Comprehensive dataset of Tweets containing the keyword 'ukraine' (in German, Russian and Ukrainian) as well as '#ukraine' (in English) since the Russian Ukraine Invasion in February 2022. The user handle column has been excluded to protect deleted accounts that have not been retweeted or replied to. Tweets have been collected via the Academic API using the Search endpoint in four languages: Language Query Name in Dataset (column: event) Number of Tweets English #ukraine AND lang='en' ukraine-en-hashtag 45.8 million German ukraine AND lang='de' ukraine 19.6 million Russian Украина AND lang:ru ukraine-ru 5.02 million Ukrainian Україна AND lang:uk ukraine-uk 4.1 million Collection dates Details on collection dates per Tweet (e.g. to compare with creation dates) as well as the IDs of Tweets for consistency checks can be found here: https://github.com/Leibniz-HBI/ukraine_twitter_data (https://doi.org/10.17605/OSF.IO/RTQXN) File Naming Scheme To enable downloads of selected timeframes and languages, the files are named by language, start and end date of the tweet creation timestamp. Columns The following columns are available: event: tag for query and language used for the query id: Tweet ID inserted_at: collection date last_updated_at: last update date (relevant for metrics such as follower count) text: Tweet text lang: language as determined by Twitter created_at: creation date of the Tweet conversation_id: Tweets with the same ID are part of the same reply tree to a tweet (provided by Twitter) author_follower_count: follower count of the Tweet's author account at the creation or last update time of the tweet replied_to: account the Tweet replies to replied_to_follower_count: follower count of the Tweet's replied to account at the creation or last update time of the tweet quoted: if quote tweet, ID of quoted tweet quoted_follower_count: analog to replied_to_follower_count retweeted: analog to quoted retweeted_follower_count: analog to replied_to_follower_count hashtags: hashtags of the Tweet urls: URLs in the tweet, shortened/unshortened, including links to media place_id: alphanumeric place ID provided by the Twitter API, mostly empty

  20. H

    Supplementary Materials for "2D AMIS Normalization for Spatial Data"

    • dataverse.harvard.edu
    Updated Jun 9, 2026
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Gennady Kravtsov (2026). Supplementary Materials for "2D AMIS Normalization for Spatial Data" [Dataset]. http://doi.org/10.7910/DVN/J2EB8S
    Explore at:
    CroissantCroissant is a format for machine-learning datasets. Learn more about this at mlcommons.org/croissant.
    Dataset updated
    Jun 9, 2026
    Dataset provided by
    Harvard Dataverse
    Authors
    Gennady Kravtsov
    License

    Attribution 4.0 (CC BY 4.0)https://creativecommons.org/licenses/by/4.0/
    License information was derived automatically

    Description

    This dataset accompanies the article "2D ADAPTIVE MULTI-INTERVAL SCALE (AMIS): METHOD FOR NORMALIZATION AND VISUALIZATION OF SPATIAL DATA" (https://doi.org/10.5281/zenodo.20577673). The dataset contains: 1. Example data: - Reference dataset (32 teams, Russian Championship 2012-13, 25×40 grid, 32,000 observations) - Match data (France vs Croatia, 2018 World Cup final; England vs Germany, 1966 World Cup final) 2. Software: - Executable tool (2D AMIS normalizer for Windows, packaged as 2D_AMIS_Tool_eng.zip) - Python source code (2D_AMIS_Tool_eng.py) 3. Results (figures from the article): - AMIS-normalized heatmaps (cellwise, 25×40 grid) - AMIS transformation curve for central cell (row 13, column 20) - Spatial profile (row 13, Croatia) Method summary: The 2D AMIS method normalizes each spatial cell individually using adaptive multi-interval scaling, transforming raw data into a unified [0, 100] scale where 50 represents the reference norm. Unlike min-max or z-score normalization, AMIS accounts for the local density of data distribution. Software notes: - The executable file is provided as .zip due to archiving policies. Unzip before use. - First launch may take 10-30 seconds due to Python initialization. License: - Source code: MIT - Data: CC BY 4.0 Related links: - GitHub: https://github.com/Famimot/2D_AMIS_Tool - Related preprint (SSRN): https://ssrn.com/abstract=6792479

Share
FacebookFacebook
TwitterTwitter
Email
Click to copy link
Link copied
Close
Cite
Mikhail Nefedov (2018). Datasets for evaluation of keyword extraction in Russian [Dataset]. https://github.com/mannefedov/ru_kw_eval_datasets

Datasets for evaluation of keyword extraction in Russian

Explore at:
Dataset updated
Jun 11, 2018
Authors
Mikhail Nefedov
Description

Datasets for evaluation of keyword extraction in Russian

Search
Clear search
Close search
Google apps
Main menu