100+ datasets found
  1. h

    Turkish-Alpaca

    • huggingface.co
    Updated Aug 15, 2023
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    TÜBİTAK Science High School AI Club (2023). Turkish-Alpaca [Dataset]. https://huggingface.co/datasets/TFLai/Turkish-Alpaca
    Explore at:
    CroissantCroissant is a format for machine-learning datasets. Learn more about this at mlcommons.org/croissant.
    Dataset updated
    Aug 15, 2023
    Dataset authored and provided by
    TÜBİTAK Science High School AI Club
    License

    Apache License, v2.0https://www.apache.org/licenses/LICENSE-2.0
    License information was derived automatically

    Description

    Stanford alpaca turkish: Stanford Alpaca

  2. h

    turkish-sentiment-analysis-dataset

    • huggingface.co
    Updated Jun 22, 2022
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Batuhan (2022). turkish-sentiment-analysis-dataset [Dataset]. https://huggingface.co/datasets/winvoker/turkish-sentiment-analysis-dataset
    Explore at:
    CroissantCroissant is a format for machine-learning datasets. Learn more about this at mlcommons.org/croissant.
    Dataset updated
    Jun 22, 2022
    Authors
    Batuhan
    License

    Attribution-ShareAlike 4.0 (CC BY-SA 4.0)https://creativecommons.org/licenses/by-sa/4.0/
    License information was derived automatically

    Description

    Dataset

    This dataset contains positive , negative and notr sentences from several data sources given in the references. In the most sentiment models , there are only two labels; positive and negative. However , user input can be totally notr sentence. For such cases there were no data I could find. Therefore I created this dataset with 3 class. Positive and negative sentences are listed below. Notr examples are extraced from turkish wiki dump. In addition, added some random text… See the full description on the dataset page: https://huggingface.co/datasets/winvoker/turkish-sentiment-analysis-dataset.

  3. Turkish Call Center Conversations

    • kaggle.com
    zip
    Updated Jul 1, 2025
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Anıl Sevinc (2025). Turkish Call Center Conversations [Dataset]. https://www.kaggle.com/datasets/anills/turkish-call-center-conversations
    Explore at:
    zip(442962 bytes)Available download formats
    Dataset updated
    Jul 1, 2025
    Authors
    Anıl Sevinc
    Description

    Turkish Multi-Domain Customer Service Conversations Dataset

    This dataset includes Turkish-language dialogues between customers and representatives across various service domains, such as finance, e-commerce, technical support, and general inquiries.

    Each conversation contains multiple turns and is structured with clearly labeled speaker roles (customer or representative), making it suitable for Natural Language Processing (NLP) tasks related to dialogue systems, intent detection, and chatbot development.

    🔹 Dataset Structure

    • conversation_id: Unique identifier for each conversation
    • category: Service domain label (e.g. Finance, Technical Support)
    • speaker: Role of the speaker (customer or representative)
    • text: Utterance in Turkish

    🔍 Use Cases

    • Chatbot training
    • Intent classification
    • Dialogue summarization
    • Speaker role detection
    • Turkish NLP pretraining/finetuning

    🧾 License

    CC BY 4.0 — Free to use with attribution

    🙋‍♂️ Creator

    Prepared by Anıl as part of a research and educational NLP project.

  4. NER Dataset(Turkish)

    • kaggle.com
    zip
    Updated Apr 14, 2024
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Akay (2024). NER Dataset(Turkish) [Dataset]. https://www.kaggle.com/datasets/akay16/ner-datasetturkish
    Explore at:
    zip(9149708 bytes)Available download formats
    Dataset updated
    Apr 14, 2024
    Authors
    Akay
    License

    MIT Licensehttps://opensource.org/licenses/MIT
    License information was derived automatically

    Description

    Data split:

    • 18.000 train
    • 1000 test
    • 1000 dev

    Labels:

    • CARDINAL
    • DATE
    • EVENT
    • FAC
    • GPE
    • LANGUAGE
    • LAW
    • LOC
    • MONEY
    • NORP
    • ORDINAL
    • ORG
    • PERCENT
    • PERSON
    • PRODUCT
    • QUANTITY
    • TIME
    • TITLE
    • WORK_OF_ART

    I do not **own **this dataset. I changed the original format for easier use and turned it into a **csv **and .spacy file.

    you reach the original version of the dataset from the **link **below

    https://github.com/turkish-nlp-suite/Turkish-Wiki-NER-Dataset

  5. F

    Turkish General Domain Scripted Monologue Speech Data

    • futurebeeai.com
    wav
    Updated Aug 1, 2022
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    FutureBee AI (2022). Turkish General Domain Scripted Monologue Speech Data [Dataset]. https://www.futurebeeai.com/dataset/monologue-speech-dataset/general-scripted-speech-monologues-turkish-turkey
    Explore at:
    wavAvailable download formats
    Dataset updated
    Aug 1, 2022
    Dataset provided by
    FutureBeeAI
    Authors
    FutureBee AI
    License

    https://www.futurebeeai.com/policies/ai-data-license-agreementhttps://www.futurebeeai.com/policies/ai-data-license-agreement

    Dataset funded by
    FutureBeeAI
    Description

    Introduction

    The Turkish Scripted Monologue Speech Dataset for the General Domain is a carefully curated resource designed to support the development of Turkish language speech recognition systems. This dataset focuses on general-purpose conversational topics and is ideal for a wide range of AI applications requiring natural, domain-agnostic Turkish speech data.

    Speech Data

    This dataset features over 6,000 high-quality scripted monologue recordings in Turkish. The prompts span diverse real-life topics commonly encountered in general conversations and are intended to help train robust and accurate speech-enabled technologies.

    Participant Diversity

    - Speakers: 60 native Turkish speakers

    - Regions: Broad regional coverage ensures diverse accents and dialects

    - Demographics: Participants aged 18 to 70, with a 60:40 male-to-female ratio

    Recording Specifications

    - Recording Type: Scripted monologues and prompt-based recordings

    - Audio Duration: 5 to 30 seconds per file

    - Format: WAV, mono channel, 16-bit, 8 kHz & 16 kHz sample rates

    - Environment: Clean, noise-free conditions to ensure clarity and usability

    Topic Coverage

    The dataset covers a wide variety of general conversation scenarios, including:

    Daily Conversations

    Topic-Specific Discussions

    General Knowledge and Advice

    Idioms and Sayings

    Contextual Features

    To enhance authenticity, the prompts include:

    Names: Male and female names specific to different Turkey regions

    Addresses: Commonly used address formats in daily Turkish speech

    Dates & Times: References used in general scheduling and time expressions

    Organization Names: Names of businesses, institutions, and other entities

    Numbers & Currencies: Mentions of quantities, prices, and monetary values

    Each prompt is designed to reflect everyday use cases, making it suitable for developing generalized NLP and ASR solutions.

    Transcription

    Every audio file in the dataset is accompanied by a verbatim text transcription, ensuring accurate training and evaluation of speech models.

    Content: Exact match to the spoken audio

    Format: Plain text (.TXT), named identically to the corresponding audio file

    Quality Control: All transcripts are validated by native Turkish transcribers

    Metadata

    Rich metadata is included for detailed filtering and analysis:

    Speaker Metadata: Unique speaker ID, age, gender, region, and dialect

    Audio Metadata: Prompt transcript, recording setup, device specs, sample rate, bit depth, and format.

    License

    This dataset is developed and owned by FutureBeeAI and is available for commercial use, offering high-value resources for enterprises and research organizations developing Turkish speech technologies.

  6. h

    turkish-offensive-language-detection

    • huggingface.co
    Updated Mar 12, 2024
    + more versions
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Toygar Tanyel (2024). turkish-offensive-language-detection [Dataset]. https://huggingface.co/datasets/Toygar/turkish-offensive-language-detection
    Explore at:
    Dataset updated
    Mar 12, 2024
    Authors
    Toygar Tanyel
    License

    Attribution 2.0 (CC BY 2.0)https://creativecommons.org/licenses/by/2.0/
    License information was derived automatically

    Description

    Dataset Summary

    This dataset is enhanced version of existing offensive language studies. Existing studies are highly imbalanced, and solving this problem is too costly. To solve this, we proposed contextual data mining method for dataset augmentation. Our method is basically prevent us from retrieving random tweets and label individually. We can directly access almost exact hate related tweets and label them directly without any further human interaction in order to solve imbalanced… See the full description on the dataset page: https://huggingface.co/datasets/Toygar/turkish-offensive-language-detection.

  7. NLI-TR (Turkish NLI Research)

    • kaggle.com
    zip
    Updated Dec 6, 2022
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    The Devastator (2022). NLI-TR (Turkish NLI Research) [Dataset]. https://www.kaggle.com/datasets/thedevastator/unlocking-turkish-nli-research-with-the-nli-tr-d
    Explore at:
    zip(43062479 bytes)Available download formats
    Dataset updated
    Dec 6, 2022
    Authors
    The Devastator
    License

    https://creativecommons.org/publicdomain/zero/1.0/https://creativecommons.org/publicdomain/zero/1.0/

    Description

    NLI-TR (Turkish NLI Research)

    Unleash Your NLI Research in Turkish Language!

    By Huggingface Hub [source]

    About this dataset

    NLI-TR is a revolutionary set of two datasets that provide an unparalleled opportunity for the natural language processing and machine learning community to conduct inference research in the Turkish Language. The datasets - SNLI-TR and MNLI-TR - contain carefully curated natural language inference data that have been translated into Turkish. With NLI-TR, researchers can explore the exciting prospects of developing automated models tailored to make inferences on texts produced in this vibrant language. Moreover, they can also investigate how models trained on data from one language fare when applied in another, a valuable insight into cross-lingual generalization capabilities. NLI-TR offers both seasoned and budding researchers an unprecedented platform to further our understanding of natural language inferencing capability

    More Datasets

    For more datasets, click here.

    Featured Notebooks

    • 🚨 Your notebook can be here! 🚨!

    How to use the dataset

    How To Use The NLI-TR Dataset to Unlock Turkish NLI Research

    Welcome to the exciting world of natural language inference (NLI) research! If you’re looking for a great dataset to use for your research in this field, the NLI-TR dataset is a perfect starting point. This guide will provide an overview of how you can use the data from this dataset to uncover new insights about NLI tasks in Turkish.

    The NLI-TR dataset contains two large scale datasets intended for natural language inference tasks – SNLI-TR and MNLI- TR. Both datasets offer researchers an opportunity to explore Natural Language Inference (NLI) research in the Turkish language, with examples ranging from sentence paraphrasing task and classification tasks to question answering scenarios using various NLP techniques.

    Using the Data:

    The data provided in this dataset includes both training and validation sets, making it easy for researchers who are just getting started with their projects. The SNLI_tr_train.csv file is used as input for training your models, while slni_tr_validation can be used as input for testing or validating model accuracy on unseen data. Additionally, multinli_tr_validation_{matched / mismatched}.csv files offer additional validation on how well your trained models perform on more complex scenarios such as sentence paraphrasing or question answering tasks using various NLP techniques.

    Each record includes four columns – premise ,hypothesis ,label , (and domain). The premise column specifies what information is provided before asking a question or making an inference; think of it as context clues that explain why one statement implies another statement more directly than others might do without them . The hypothesis column provides what lies at the heart of inference --the conclusion reached after introducing facts given before it . Last but not least we have label column which denotes whether two sentences entail each other (ENTAILMENT), contradict each other(CONTRADICTION) or are unrelated(NEUTRAL). A domain label has also been assigned by some authors when necessary; this mostly applies when inferring between sentences across different semantic domains such as weather vs sports vs finance etc .

    Research Ideas

    • Developing an NLI-based Turkish language question answering system.
    • Training a sentiment analysis algorithm to identify sentiment in text written in Turkish.
    • Building a Machine Learning Chatbot that uses NLI to understand conversational context and respond accordingly for users intending to converse in the Turkish language

    Acknowledgements

    If you use this dataset in your research, please credit the original authors. Data Source

    License

    License: CC0 1.0 Universal (CC0 1.0) - Public Domain Dedication No Copyright - You can copy, modify, distribute and perform the work, even for commercial purposes, all without asking permission. See Other Information.

    Columns

    File: snli_tr_train.csv | Column name | Description | |:---------------|:------------------------------------------------------------------------------------------------------------------------------------------...

  8. Turkish Book Data Set

    • kaggle.com
    zip
    Updated Jan 12, 2024
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Muhammed İbrahim Top (2024). Turkish Book Data Set [Dataset]. https://www.kaggle.com/datasets/muhammedbrahimtop/turkish-book-data-set
    Explore at:
    zip(17318668 bytes)Available download formats
    Dataset updated
    Jan 12, 2024
    Authors
    Muhammed İbrahim Top
    License

    Apache License, v2.0https://www.apache.org/licenses/LICENSE-2.0
    License information was derived automatically

    Description

    Turkish Book Data Set

    This dataset is a comprehensive compilation of Turkish books obtained through web scraping from the internet. Each record in the dataset contains essential information such as the book's title, author, publisher, publication year, page count, category, description, and image URL.

    This rich dataset can be utilized in various applications, particularly in the analysis through methods such as classification, content-based recommendation algorithms, and natural language processing (NLP). For researchers, students, and data scientists, this dataset serves as a valuable resource for exploring Turkish literature, generating book recommendations, or developing machine learning models.

  9. h

    Turkish_Speech_Corpus

    • huggingface.co
    • kaggle.com
    Updated Dec 17, 2025
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Institute of Smart Systems and Artificial Intelligence, Nazarbayev University (2025). Turkish_Speech_Corpus [Dataset]. https://huggingface.co/datasets/issai/Turkish_Speech_Corpus
    Explore at:
    CroissantCroissant is a format for machine-learning datasets. Learn more about this at mlcommons.org/croissant.
    Dataset updated
    Dec 17, 2025
    Dataset authored and provided by
    Institute of Smart Systems and Artificial Intelligence, Nazarbayev University
    License

    MIT Licensehttps://opensource.org/licenses/MIT
    License information was derived automatically

    Description

    Turkish Speech Corpus (TSC)

    This repository presents an open-source Turkish Speech Corpus, introduced in "Multilingual Speech Recognition for Turkic Languages". The corpus contains 218.2 hours of transcribed speech with 186,171 utterances and is the largest publicly available Turkish dataset of its kind at that time. Paper: Multilingual Speech Recognition for Turkic Languages.
    GitHub Repository: https://github.com/IS2AI/TurkicASR

      Citation
    

    @Article{info14020074… See the full description on the dataset page: https://huggingface.co/datasets/issai/Turkish_Speech_Corpus.

  10. Turkish Polite Dataset

    • kaggle.com
    zip
    Updated Apr 17, 2025
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Yunus Emre Akca (2025). Turkish Polite Dataset [Dataset]. https://www.kaggle.com/datasets/yunusemreakca/turkish-polite-dataset
    Explore at:
    zip(98871 bytes)Available download formats
    Dataset updated
    Apr 17, 2025
    Authors
    Yunus Emre Akca
    License

    MIT Licensehttps://opensource.org/licenses/MIT
    License information was derived automatically

    Description

    This dataset contains chat-based dialogs in the Turkish language. The dialogs are written in a particularly natural, polite and supportive style. The interactions between the user and the chatbot aim to provide information and support on different topics. This dataset is suitable for Turkish language processing (NLP) projects and can be used in areas such as chatbots, language modeling and text analysis.

  11. m

    Turkish Offensive Language Dataset

    • megatek.ai
    bin
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Turkish Offensive Language Dataset [Dataset]. https://megatek.ai/en/dataset/turkish-offensive-language-dataset/
    Explore at:
    binAvailable download formats
    License

    https://www.apache.org/licenses/LICENSE-2.0https://www.apache.org/licenses/LICENSE-2.0

    Description

    The Turkish Offensive Language Dataset is a Turkish-language dataset collected from Twitter, designed for training models in offensive language detection, hate speech detection, and text classification tasks. Created by Gülzade Evni and Zeynep Baydemir, the dataset covers multiple subcategories of harmful content including racism, profanity, insult, and sexism.

    It consists of seven files and is distributed under the Apache-2.0 license, making it openly available for research and development purposes. The dataset is intended for practitioners and researchers working on natural language processing for Turkish social media content, addressing a recognized gap in low-resource language resources for content moderation applications.

  12. F

    Turkish Call Center Data for Delivery & Logistics AI

    • futurebeeai.com
    wav
    Updated Aug 1, 2022
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    FutureBee AI (2022). Turkish Call Center Data for Delivery & Logistics AI [Dataset]. https://www.futurebeeai.com/dataset/speech-dataset/delivery-call-center-conversation-turkish-turkey
    Explore at:
    wavAvailable download formats
    Dataset updated
    Aug 1, 2022
    Dataset provided by
    FutureBeeAI
    Authors
    FutureBee AI
    License

    https://www.futurebeeai.com/policies/ai-data-license-agreementhttps://www.futurebeeai.com/policies/ai-data-license-agreement

    Dataset funded by
    FutureBeeAI
    Description

    Introduction

    This Turkish Call Center Speech Dataset for the Delivery and Logistics industry is purpose-built to accelerate the development of speech recognition, spoken language understanding, and conversational AI systems tailored for Turkish-speaking customers. With over 30 hours of real-world, unscripted call center audio, this dataset captures authentic delivery-related conversations essential for training high-performance ASR models.

    Curated by FutureBeeAI, this dataset empowers AI teams, logistics tech providers, and NLP researchers to build accurate, production-ready models for customer support automation in delivery and logistics.

    Speech Data

    The dataset contains 30 hours of dual-channel call center recordings between native Turkish speakers. Captured across various delivery and logistics service scenarios, these conversations cover everything from order tracking to missed delivery resolutions offering a rich, real-world training base for AI models.

    Participant Diversity:

    - Speakers: 60 native Turkish speakers from our verified contributor pool.

    - Regions: Multiple provinces of Turkey for accent and dialect diversity.

    - Participant Profile: Balanced gender distribution (60% male, 40% female) with ages ranging from 18 to 70.

    Recording Details:

    - Conversation Nature: Naturally flowing, unscripted customer-agent dialogues.

    - Call Duration: 5 to 15 minutes on average.

    - Audio Format: Stereo WAV, 16-bit depth, recorded at 8kHz and 16kHz.

    - Recording Environment: Captured in clean, noise-free, echo-free conditions.

    Topic Diversity

    This speech corpus includes both inbound and outbound delivery-related conversations, covering varied outcomes (positive, negative, neutral) to train adaptable voice models.

    Inbound Calls:

    - Order Tracking

    - Delivery Complaints

    - Undeliverable Addresses

    - Return Process Enquiries

    - Delivery Method Selection

    - Order Modifications, and more

    Outbound Calls:

    - Delivery Confirmations

    - Subscription Offer Calls

    - Incorrect Address Follow-ups

    - Missed Delivery Notifications

    - Delivery Feedback Surveys

    - Out-of-Stock Alerts, and others

    This comprehensive coverage reflects real-world logistics workflows, helping voice AI systems interpret context and intent with precision.

    Transcription

    All recordings come with high-quality, human-generated verbatim transcriptions in JSON format.

    Transcription Includes:

    - Speaker-Segmented Dialogues

    - Time-coded Segments

    - Non-speech Tags (e.g., pauses, noise)

    - High transcription accuracy with word error rate under 5% via dual-layer quality checks.

    These transcriptions support fast, reliable model development for Turkish voice AI applications in the delivery sector.

    Metadata

    Detailed metadata is included for each participant and conversation:

    Participant Metadata: ID, age, gender, region, accent, dialect.

    Conversation Metadata: Topic, call type, sentiment, sample rate, and technical attributes.

    This metadata aids in training specialized models, filtering demographics, and running advanced analytics.

    License

    This Delivery and Logistics domain dataset is commercially licensed and ready for use in ASR, NLP, and voice automation projects in Turkish.

  13. Genius-Turkish-Dataset

    • kaggle.com
    zip
    Updated Nov 2, 2025
    + more versions
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Mustafa Kemal Çıngıl (2025). Genius-Turkish-Dataset [Dataset]. https://www.kaggle.com/datasets/mustafakemal0146/genius-turkish-dataset
    Explore at:
    zip(22818735 bytes)Available download formats
    Dataset updated
    Nov 2, 2025
    Authors
    Mustafa Kemal Çıngıl
    License

    Apache License, v2.0https://www.apache.org/licenses/LICENSE-2.0
    License information was derived automatically

    Description

    TURKISH SONG LYRICS FROM GENIUS DATASET

    DATASET DESCRIPTION

    This dataset contains a comprehensive collection of 44,692 Turkish song lyrics, extracted from the larger "Genius Song Lyrics with Language Information" dataset available on Kaggle. The original 9.07 GB dataset was filtered to include only songs identified with the language code 'tr' (Turkish), making it a clean and focused resource for Turkish Natural Language Processing (NLP) tasks.

    [TR] Bu veri seti, Kaggle'da bulunan "Genius Song Lyrics with Language Information" adlı büyük veri setinden ayıklanmış 44,692 adet Türkçe şarkı sözü içermektedir. Orijinal 9.07 GB'lık veri seti, dil kodu 'tr' (Türkçe) olarak tanımlanmış şarkıları içerecek şekilde filtrelenmiştir. Bu, Türkçe Doğal Dil İşleme (DDİ) görevleri için temiz ve odaklanmış bir kaynak oluşturmaktadır.

    HOW TO USE

    You can easily load this dataset using the Hugging Face datasets library.

    [TR] Bu veri setini Hugging Face datasets kütüphanesini kullanarak kolayca yükleyebilirsiniz.

    Örnek Python Kodu: ```python from datasets import load_dataset

    Load the dataset from the Hugging Face Hub

    Veri setini Hugging Face Hub'dan yükleyin

    dataset = load_dataset("mustafakemal0146/Genius-Turkish-Dataset")

    Example: Access the lyrics of the first song in the training split

    Örnek: Eğitim setindeki ilk şarkının sözlerine erişim

    print(dataset['train'][0]['lyrics']) ```

    DATASET STRUCTURE

    The dataset consists of a single CSV file, loaded as the train split, with the following columns:

    [TR] Veri seti, train bölünmüşü olarak yüklenen ve aşağıdaki sütunları içeren tek bir CSV dosyasından oluşur:

    • title (string): The title of the song / Şarkının başlığı.
    • tag (string): The genre tag associated with the song (e.g., 'rap', 'pop') / Şarkıyla ilişkilendirilen tür etiketi.
    • artist (string): The name of the primary artist / Ana sanatçının adı.
    • year (int): The release year of the song / Şarkının çıkış yılı.
    • views (int): The number of views on Genius.com / Genius.com'daki görüntülenme sayısı.
    • features (string): A string representation of featuring artists / Düet yapılan sanatçıların metin formatı.
    • lyrics (string): The full lyrics of the song / Şarkının tam sözleri.
    • id (int): A unique identifier from the original dataset / Orijinal veri setinden gelen benzersiz ID.
    • language_cld3 (string): Language code detected by CLD3 model (all 'tr') / CLD3 ile tespit edilen dil kodu (tümü 'tr').
    • language_ft (string): Language code detected by FastText model (all 'tr') / FastText ile tespit edilen dil (tümü 'tr').
    • language (string): Final aggregated language code (all 'tr') / Nihai dil kodu (tümü 'tr').

    DATA SOURCE AND CURATION

    This dataset is a curated subset of the "Genius Song Lyrics with Language Information" (Link: https://www.kaggle.com/datasets/pavanelisetty/genius-song-lyrics-with-language-information) dataset on Kaggle, originally collected from Genius.com. The filtering process involved reading the main song_lyrics.csv file and selecting all rows where the language column was equal to 'tr'.

    [TR] Bu veri seti, orijinal olarak Genius.com'dan toplanmış olan Kaggle'daki "Genius Song Lyrics with Language Information" veri setinin düzenlenmiş bir alt kümesidir. Filtreleme işlemi, ana song_lyrics.csv dosyasını okuyarak language sütununun 'tr' olduğu tüm satırların seçilmesiyle yapılmıştır.

    LICENSE

    The original dataset on Kaggle does not specify a license. This curated version is shared under the "Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International (CC BY-NC-SA 4.0)" (Link: https://creativecommons.org/licenses/by-nc-sa/4.0/) license, assuming it will be used for non-commercial research and educational purposes. Please refer to the original data source for any commercial use inquiries.

    Created by MustafaKemal0146 (Hugging Face Profile: https://huggingface.co/MustafaKemal0146)

  14. s

    Turkish Language Training Dataset - 2000H Spontaneous Pair Conversational...

    • spirelight.ai
    16bit 48khz wav, json +1
    Updated Jun 21, 2026
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Spirelight (2026). Turkish Language Training Dataset - 2000H Spontaneous Pair Conversational Audio and Video [Dataset]. https://www.spirelight.ai/datasets/turkish-language-training-dataset-2000h-spontaneus-pair-conversational-audio-and-video
    Explore at:
    16bit 48khz wav, json, mkvAvailable download formats
    Dataset updated
    Jun 21, 2026
    Dataset authored and provided by
    Spirelight
    Description

    2000 hours of Turkish spontaneous conversations with metadata and transcripts.

  15. s

    Turkish Language Speech Datasets | NLP, Conversational AI & Machine Learning...

    • shaip.com
    Updated Feb 10, 2023
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Shaip (2023). Turkish Language Speech Datasets | NLP, Conversational AI & Machine Learning [Dataset]. https://www.shaip.com/offerings/speech-data-catalog/turkish-turkey-dataset/
    Explore at:
    Dataset updated
    Feb 10, 2023
    Dataset authored and provided by
    Shaip
    License

    CC0 1.0 Universal Public Domain Dedicationhttps://creativecommons.org/publicdomain/zero/1.0/
    License information was derived automatically

    Description

    Enhance your Conversational AI model with our Off-the-Shelf Turkish Language Dataset (Turkish Language Speech Datasets). Shaip high-quality audio datasets are a quick

  16. R

    Turkish Tiel Dataset

    • universe.roboflow.com
    zip
    Updated Dec 13, 2023
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    edoski (2023). Turkish Tiel Dataset [Dataset]. https://universe.roboflow.com/edoski/turkish-tiel
    Explore at:
    zipAvailable download formats
    Dataset updated
    Dec 13, 2023
    Dataset authored and provided by
    edoski
    License

    Attribution 4.0 (CC BY 4.0)https://creativecommons.org/licenses/by/4.0/
    License information was derived automatically

    Variables measured
    Money Bounding Boxes
    Description

    Turkish Tiel

    ## Overview
    
    Turkish Tiel is a dataset for object detection tasks - it contains Money annotations for 604 images.
    
    ## Getting Started
    
    You can download this dataset for use within your own projects, or fork it into a workspace on Roboflow to create your own model.
    
      ## License
    
      This dataset is available under the [CC BY 4.0 license](https://creativecommons.org/licenses/CC BY 4.0).
    
  17. Turkish Sentiment Analysis Dataset

    • humirapps.cs.hacettepe.edu.tr
    zip
    Updated Apr 12, 2017
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Hacettepe University Multimedia Information Retrieval Laboratory (2017). Turkish Sentiment Analysis Dataset [Dataset]. http://doi.org/10.1109/SITIS.2016.57
    Explore at:
    zipAvailable download formats
    Dataset updated
    Apr 12, 2017
    Dataset provided by
    Hacettepe Universityhttp://hacettepe.edu.tr/
    Authors
    Hacettepe University Multimedia Information Retrieval Laboratory
    License

    CC0 1.0 Universal Public Domain Dedicationhttps://creativecommons.org/publicdomain/zero/1.0/
    License information was derived automatically

    Description

    We have selected two most popular movie and hotel recommendation websites from those which attain a high rate in the Alexa website. We selected “beyazperde.com” and “otelpuan.com” for movie and hotel reviews, respectively. The reviews of 5,660 movies were investigated. The all 220,000 extracted reviews had been already rated by own authors using stars 1 to 5. As most of the reviews were positive, we selected the positive reviews as much as the negative ones to provide a balanced situation. The total of negative reviews rated by 1 or 2 stars were 26,700, thus, we randomly selected 26,700 out of 130,210 positive reviews rated by 4 or 5 stars. Overall, 53,400 movie reviews by the average length of 33 words were selected. The similar manner was used to hotel reviews with the difference that the hotel reviews had been rated by the numbers between 0 and 100 instead of stars. From 18,478 reviews extracted from 550 hotels, a balanced set of positive and negative reviews was selected. As there were only 5,802 negative hotel reviews using 0 to 40 rating, we selected 5800 out of 6499 positive reviews rated from 80 to 100. The average length of all 11,600 selected positive and negative hotel reviews were 74 which is more than two times of the movie reviews.

  18. i

    TRNEWS-2025: A Multiclass Turkish News Text Dataset

    • ieee-dataport.org
    Updated Oct 9, 2025
    + more versions
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    sengul HAYTA (2025). TRNEWS-2025: A Multiclass Turkish News Text Dataset [Dataset]. https://ieee-dataport.org/documents/trnews-2025-multiclass-turkish-news-text-dataset
    Explore at:
    Dataset updated
    Oct 9, 2025
    Authors
    sengul HAYTA
    Description

    Health

  19. F

    Turkish Shopping List OCR Image Dataset

    • futurebeeai.com
    wav
    Updated Aug 1, 2022
    + more versions
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    FutureBee AI (2022). Turkish Shopping List OCR Image Dataset [Dataset]. https://www.futurebeeai.com/dataset/ocr-dataset/turkish-shopping-list-ocr-image-dataset
    Explore at:
    wavAvailable download formats
    Dataset updated
    Aug 1, 2022
    Dataset provided by
    FutureBeeAI
    Authors
    FutureBee AI
    License

    https://www.futurebeeai.com/policies/ai-data-license-agreementhttps://www.futurebeeai.com/policies/ai-data-license-agreement

    Dataset funded by
    FutureBeeAI
    Description

    Introducing the Turkish Shopping List Image Dataset - a diverse and comprehensive collection of handwritten text images carefully curated to propel the advancement of text recognition and optical character recognition (OCR) models designed specifically for the Turkish language.

    Dataset Contain & Diversity:

    Containing more than 2000 images, this Turkish OCR dataset offers a wide distribution of different types of shopping list images. Within this dataset, you'll discover a variety of handwritten text, including sentences, and individual item name words, quantity, comments, etc on shopping lists. The images in this dataset showcase distinct handwriting styles, fonts, font sizes, and writing variations.

    To ensure diversity and robustness in training your OCR model, we allow limited (less than three) unique images in a single handwriting. This ensures we have diverse types of handwriting to train your OCR model on. Stringent measures have been taken to exclude any personally identifiable information (PII) and to ensure that in each image a minimum of 80% of space contains visible Turkish text.

    The images have been captured under varying lighting conditions, including day and night, as well as different capture angles and backgrounds. This diversity helps build a balanced OCR dataset, featuring images in both portrait and landscape modes.

    All these shopping lists were written and images were captured by native Turkish people to ensure text quality, prevent toxic content, and exclude PII text. We utilized the latest iOS and Android mobile devices with cameras above 5MP to maintain image quality. Images in this training dataset are available in both JPEG and HEIC formats.

    Metadata:

    In addition to the image data, you will receive structured metadata in CSV format. For each image, this metadata includes information on image orientation, country, language, and device details. Each image is correctly named to correspond with the metadata.

    This metadata serves as a valuable resource for understanding and characterizing the data, aiding informed decision-making in the development of Turkish text recognition models.

    Update & Custom Collection:

    We are committed to continually expanding this dataset by adding more images with the help of our native Turkish crowd community.

    If you require a customized OCR dataset containing shopping list images tailored to your specific guidelines or device distribution, please don't hesitate to contact us. We have the capability to curate specialized data to meet your unique requirements.

    Additionally, we can annotate or label the images with bounding boxes or transcribe the text in the images to align with your project's specific needs using our crowd community.

    License:

    This image dataset, created by FutureBeeAI, is now available for commercial use.

    Conclusion:

    Leverage this shopping list image OCR dataset to enhance the training and performance of text recognition, text detection, and optical character recognition models for the Turkish language. Your journey to improved language understanding and processing begins here.

  20. m

    AlicanKiraz0/Turkish-SFT-Dataset-v1.0

    • metatext.io
    parquet
    Updated Jul 4, 2026
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    AlicanKiraz0 (2026). AlicanKiraz0/Turkish-SFT-Dataset-v1.0 [Dataset]. https://metatext.io/datasets/alicankiraz0/turkish-sft-dataset-v1.0
    Explore at:
    parquetAvailable download formats
    Dataset updated
    Jul 4, 2026
    Dataset authored and provided by
    AlicanKiraz0
    License

    MIT Licensehttps://opensource.org/licenses/MIT
    License information was derived automatically

    Variables measured
    Text Generation, Question Answering, Text Classification
    Description

    Turkish-SFT-Dataset-v1.01

    Repo: AlicanKiraz0/Turkish-SFT-Dataset-v1.0Sürüm: v1.01Lisans: MITBiçim: jsonl (kolonlar: system, user, assistant)Boyut: ~5500 satır ve satır başına 3.000–4.500 token/satır (≈ 20M+ token)Dil: Türkçe (tr)Görevler: talim...

Share
FacebookFacebook
TwitterTwitter
Email
Click to copy link
Link copied
Close
Cite
TÜBİTAK Science High School AI Club (2023). Turkish-Alpaca [Dataset]. https://huggingface.co/datasets/TFLai/Turkish-Alpaca

Turkish-Alpaca

TFLai/Turkish-Alpaca

Explore at:
CroissantCroissant is a format for machine-learning datasets. Learn more about this at mlcommons.org/croissant.
Dataset updated
Aug 15, 2023
Dataset authored and provided by
TÜBİTAK Science High School AI Club
License

Apache License, v2.0https://www.apache.org/licenses/LICENSE-2.0
License information was derived automatically

Description

Stanford alpaca turkish: Stanford Alpaca

Search
Clear search
Close search
Google apps
Main menu