100+ datasets found
  1. French Conversations (from movie subtitles)

    • kaggle.com
    zip
    Updated Aug 3, 2023
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Dali Selmi (2023). French Conversations (from movie subtitles) [Dataset]. https://www.kaggle.com/datasets/daliselmi/french-conversational-dataset
    Explore at:
    zip(2880370702 bytes)Available download formats
    Dataset updated
    Aug 3, 2023
    Authors
    Dali Selmi
    License

    https://creativecommons.org/publicdomain/zero/1.0/https://creativecommons.org/publicdomain/zero/1.0/

    Area covered
    French
    Description

    French Movie Subtitle Conversations Dataset

    Description

    Dive into the world of French dialogue with the French Movie Subtitle Conversations dataset โ€“ a comprehensive collection of over 127,000 movie subtitle conversations. This dataset offers a deep exploration of authentic and diverse conversational contexts spanning various genres, eras, and scenarios. It is thoughtfully organized into three distinct sets: training, testing, and validation.

    Content Overview

    Each conversation in this dataset is structured as a JSON object, featuring three key attributes:

    1. Context: Get a holistic view of the conversation's flow with the preceding 9 lines of dialogue. This context provides invaluable insights into the conversation's dynamics and contextual cues.
    2. Knowledge: Immerse yourself in a wide range of thematic knowledge. This dataset covers an array of topics, ensuring that your models receive exposure to diverse information sources for generating well-informed responses.
    3. Response: Explore how characters react and respond across various scenarios. From casual conversations to intense emotional exchanges, this dataset encapsulates the authenticity of genuine human interaction.

    Data Sample

    Here's a snippet from the dataset to give you an idea of its structure:

    [
     {
      "context": [
       "Tu as attendu longtemps?",
       "Oui en effet.",
       "Je pense que c' est grossier pour un premier rencard.",
       // ... (6 more lines of context)
      ],
      "knowledge": "",
      "response": "On n' avait pas dit 9h?"
     },
     // ... (more data samples)
    ]
    

    Use Cases

    The French Movie Subtitle Conversations dataset serves as a valuable resource for several applications:

    • Conversational AI: Train advanced chatbots and dialogue systems in French that can engage users in fluid, contextually aware conversations.
    • Language Modeling: Enhance your language models by leveraging diverse dialogue patterns, colloquialisms, and contextual dependencies present in real-world conversations.
    • Sentiment Analysis: Investigate the emotional tones of conversations across different movie genres and periods, contributing to a better understanding of sentiment variation.

    Why This Dataset

    • Size and Diversity: With a vast collection of over 127,000 conversations spanning diverse genres and tones, this dataset offers an unparalleled breadth and depth in French dialogue data.
    • Contextual Richness: The inclusion of context empowers researchers and practitioners to explore the dynamics of conversation flow, leading to more accurate and contextually relevant responses.
    • Real-world Relevance: Originating from movie subtitles, this dataset mirrors real-world interactions, making it a valuable asset for training models that understand and generate human-like dialogue.

    Acknowledgments

    We extend our gratitude to the movie subtitle community for their contributions, which have enabled the creation of this diverse and comprehensive French dialogue dataset.

    Unlock the potential of authentic French conversations today with the French Movie Subtitle Conversations dataset. Engage in state-of-the-art research, enhance language models, and create applications that resonate with the nuances of real dialogue.

  2. R

    Subtitles Dataset

    • universe.roboflow.com
    zip
    Updated Oct 9, 2022
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    subtitles (2022). Subtitles Dataset [Dataset]. https://universe.roboflow.com/subtitles-jtdc8/subtitles-xmseb/dataset/3
    Explore at:
    zipAvailable download formats
    Dataset updated
    Oct 9, 2022
    Dataset authored and provided by
    subtitles
    License

    Attribution 4.0 (CC BY 4.0)https://creativecommons.org/licenses/by/4.0/
    License information was derived automatically

    Variables measured
    Letters Bounding Boxes
    Description

    Subtitles

    ## Overview
    
    Subtitles is a dataset for object detection tasks - it contains Letters annotations for 500 images.
    
    ## Getting Started
    
    You can download this dataset for use within your own projects, or fork it into a workspace on Roboflow to create your own model.
    
      ## License
    
      This dataset is available under the [CC BY 4.0 license](https://creativecommons.org/licenses/CC BY 4.0).
    
  3. h

    yyets-subtitles

    • huggingface.co
    Updated Dec 4, 2025
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    chenrm (2025). yyets-subtitles [Dataset]. https://huggingface.co/datasets/chenrm/yyets-subtitles
    Explore at:
    CroissantCroissant is a format for machine-learning datasets. Learn more about this at mlcommons.org/croissant.
    Dataset updated
    Dec 4, 2025
    Authors
    chenrm
    Description

    chenrm/yyets-subtitles dataset hosted on Hugging Face and contributed by the HF Datasets community

  4. YouTube Video Subtitles

    • kaggle.com
    zip
    Updated Feb 5, 2025
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Daniel Herman (2025). YouTube Video Subtitles [Dataset]. https://www.kaggle.com/datasets/jetakow/youtube-videos-subtitles
    Explore at:
    zip(42191918 bytes)Available download formats
    Dataset updated
    Feb 5, 2025
    Authors
    Daniel Herman
    License

    MIT Licensehttps://opensource.org/licenses/MIT
    License information was derived automatically

    Area covered
    YouTube
    Description

    Over 12k scraped YouTube EN subtitles for videos on GitHub topics.

    How? Based on the topics https://github.com/topics I searched YouTube with the phrase "What is {topic}?" and downloaded up to 100 video subtitles for a given topic. The extracted text can be found in the dataset together with the topic name, video title and video URL.

    Why? I wan to know if we can rate videos based on their information value, especially when we use YouTube as an information source.

    You can find the source code here: https://github.com/detrin/text-info-value

  5. h

    survivor-subtitles-cleaned

    • huggingface.co
    Updated Feb 7, 2025
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Paul Lambert (2025). survivor-subtitles-cleaned [Dataset]. https://huggingface.co/datasets/hipml/survivor-subtitles-cleaned
    Explore at:
    CroissantCroissant is a format for machine-learning datasets. Learn more about this at mlcommons.org/croissant.
    Dataset updated
    Feb 7, 2025
    Authors
    Paul Lambert
    Description

    Survivor Subtitles Dataset (cleaned)

      Dataset Description
    

    A collection of subtitles from the American reality television show "Survivor", spanning seasons 1 through 47. The dataset contains subtitle text extracted from episode broadcasts. This dataset is a modification of the original Survivor Subtitles dataset after cleaning up and joining subtitle fragments. This dataset is a work in progress and any contributions are welcome.

      Source
    

    The subtitles wereโ€ฆ See the full description on the dataset page: https://huggingface.co/datasets/hipml/survivor-subtitles-cleaned.

  6. Movie Parallel Subtitles (EN-IT-RU)

    • kaggle.com
    zip
    Updated Jun 26, 2025
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Timur Sharifullin (2025). Movie Parallel Subtitles (EN-IT-RU) [Dataset]. https://www.kaggle.com/datasets/timursharifullindata/movie-parallel-subtitles-small-sentiment-dataset
    Explore at:
    zip(6056 bytes)Available download formats
    Dataset updated
    Jun 26, 2025
    Authors
    Timur Sharifullin
    License

    MIT Licensehttps://opensource.org/licenses/MIT
    License information was derived automatically

    Description

    This dataset contains 25 aligned movie subtitle segments in English, Russian, and Italian, extracted from the ParTree corpus. Each row provides a short, context-rich movie line with its translations in all three languages, making it ideal for research and development in machine translation, multilingual NLP, and cross-lingual transfer learning.

    Key features: - Parallel triplets: English, Russian, Italian - Sourced from authentic movie subtitles for natural, conversational language - Suitable for training, validation, and benchmarking of translation and multilingual models

    Data originally from the ParTree corpus, available via Swiss-AL

  7. Reasons why adults use subtitles when watching TV in known language in the...

    • statista.com
    Updated Jun 4, 2024
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Statista (2024). Reasons why adults use subtitles when watching TV in known language in the U.S. 2023 [Dataset]. https://www.statista.com/statistics/1459167/reasons-use-subtitles-watching-tv-known-language-us/
    Explore at:
    Dataset updated
    Jun 4, 2024
    Dataset authored and provided by
    Statistahttps://statista.com/
    Time period covered
    Jun 29, 2023 - Jul 5, 2023
    Area covered
    United States
    Description

    Enhancement of comprehension and more profound understanding of accents were the most common reasons why American adults use subtitles while watching TV in a known language, according to a survey conducted between June and July 2023. Another ** percent of the respondents stated that they did so because they were in a noisy environment.

  8. d

    ParTree - Parallel Treebanks: A multilingual corpus of movie subtitles.

    • doi.org
    • swissubase.ch
    Updated Mar 21, 2023
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    (2023). ParTree - Parallel Treebanks: A multilingual corpus of movie subtitles. [Dataset]. http://doi.org/10.48656/5mz4-x435
    Explore at:
    Dataset updated
    Mar 21, 2023
    Description

    A multilingual corpus of movie subtitles aligned on the sentence-level. Contains data on more than 50 languages with a focus on the Indo-European language family. Morphosyntactic annotation (part-of-speech, features, dependencies) in Universal Dependency-style is available for 47 languages.

  9. D

    AI Subtitle Generation Market Research Report 2034

    • dataintelo.com
    csv, pdf, pptx
    Updated Mar 21, 2026
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Dataintelo (2026). AI Subtitle Generation Market Research Report 2034 [Dataset]. https://dataintelo.com/report/ai-subtitle-generation-market
    Explore at:
    csv, pdf, pptxAvailable download formats
    Dataset updated
    Mar 21, 2026
    Dataset authored and provided by
    Dataintelo
    License

    https://dataintelo.com/privacy-and-policyhttps://dataintelo.com/privacy-and-policy

    Time period covered
    2025 - 2034
    Area covered
    United States, United Kingdom, France, Germany, Worldwide, South Korea, Japan, China
    Description

    The AI Subtitle Generation market was valued at $3.8 billion in 2025 and is projected to reach $18.6 billion by 2034, growing at a CAGR of 19.3%.

  10. s

    Machine Translation on Subtitles (test)

    • sota2.com
    Updated Jun 10, 2024
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    (2024). Machine Translation on Subtitles (test) [Dataset]. https://www.sota2.com/research/sota/machine-translation-on-subtitles-test
    Explore at:
    Dataset updated
    Jun 10, 2024
    Variables measured
    BLEU, COMET
    Description

    Evaluation of Machine Translation performance on the test split of the Subtitles dataset, measured by BLEU and COMET scores.

  11. SubIMDB: A Structured Corpus of Subtitles

    • zenodo.org
    • data.europa.eu
    tar
    Updated Jan 24, 2020
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Paetzold; Specia; Paetzold; Specia (2020). SubIMDB: A Structured Corpus of Subtitles [Dataset]. http://doi.org/10.5281/zenodo.2552407
    Explore at:
    tarAvailable download formats
    Dataset updated
    Jan 24, 2020
    Dataset provided by
    Zenodohttp://zenodo.org/
    Authors
    Paetzold; Specia; Paetzold; Specia
    License

    Attribution 4.0 (CC BY 4.0)https://creativecommons.org/licenses/by/4.0/
    License information was derived automatically

    Description

    Exploring language usage through frequency analysis in large corpora is a defining feature in most recent work in corpus and computational linguistics. From a psycholinguistic perspective, however, the corpora used in these contributions are often not representative of language usage: they are either domain-specific, limited in size, or extracted from unreliable sources. In an effort to address this limitation, we introduce SubIMDB, a corpus of everyday language spoken text we created which contains over 225 million words. The corpus was extracted from 38,102 subtitles of family, comedy and children movies and series, and is the first sizeable structured corpus of subtitles made available. Our experiments show that word frequency norms extracted from this corpus are more effective than those from well-known norms such as Kucera-Francis, HAL and SUBTLEXus in predicting various psycholinguistic properties of words, such as lexical decision times, familiarity, age of acquisition and simplicity. We also provide evidence that contradict the long-standing assumption that the ideal size for a corpus can be determined solely based on how well its word frequencies correlate with lexical decision times.

  12. h

    Subtitles

    • huggingface.co
    Updated Apr 4, 2009
    + more versions
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Peanut Jar Mixers Development (2009). Subtitles [Dataset]. https://huggingface.co/datasets/PJMixers-Dev/Subtitles
    Explore at:
    CroissantCroissant is a format for machine-learning datasets. Learn more about this at mlcommons.org/croissant.
    Dataset updated
    Apr 4, 2009
    Dataset authored and provided by
    Peanut Jar Mixers Development
    Description

    PJMixers-Dev/Subtitles dataset hosted on Hugging Face and contributed by the HF Datasets community

  13. subtitles

    • kaggle.com
    zip
    Updated Aug 1, 2023
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Sergey Sohackiy (2023). subtitles [Dataset]. https://www.kaggle.com/datasets/sergeysohackiy/subtitles
    Explore at:
    zip(9239834 bytes)Available download formats
    Dataset updated
    Aug 1, 2023
    Authors
    Sergey Sohackiy
    Description

    Dataset

    This dataset was created by Sergey Sohackiy

    Contents

  14. s

    Video Understanding on Video-MME without subtitles

    • sota2.com
    Updated Aug 1, 2026
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    (2026). Video Understanding on Video-MME without subtitles [Dataset]. https://www.sota2.com/research/sota/video-understanding-on-video-mme-without-subtitles
    Explore at:
    Dataset updated
    Aug 1, 2026
    Variables measured
    Accuracy, Score (Long), Overall Score, Score (Short), Relative Score, Score (Medium), Score (Overall)
    Description

    Evaluation of video understanding capabilities using the Video-MME dataset, specifically excluding subtitle information, reporting scores segmented by video length (Short, Medium, Long) and overall.

  15. a

    Video Mme Long No Subtitles Leaderboard 2026

    • anotherwrapper.com
    Updated Sep 5, 2026
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    AnotherWrapper (2026). Video Mme Long No Subtitles Leaderboard 2026 [Dataset]. https://anotherwrapper.com/tools/llm-pricing/evals/video-mme-long-no-subtitles
    Explore at:
    Dataset updated
    Sep 5, 2026
    Dataset authored and provided by
    AnotherWrapper
    Variables measured
    Video Mme Long No Subtitles
    Description

    As of September 5, 2026, GPT-4.1 is #1 for Video Mme Long No Subtitles at 72%. Ranked by the Video Mme Long No Subtitles score Video Mme Long No Subtitles leaderboard: rank models by Video Mme Long No Subtitles next to live API token prices.

  16. D

    Real Time Subtitles Market Research Report 2034

    • dataintelo.com
    csv, pdf, pptx
    Updated May 14, 2026
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Dataintelo (2026). Real Time Subtitles Market Research Report 2034 [Dataset]. https://dataintelo.com/report/real-time-subtitles-market
    Explore at:
    csv, pdf, pptxAvailable download formats
    Dataset updated
    May 14, 2026
    Dataset authored and provided by
    Dataintelo
    License

    https://dataintelo.com/privacy-and-policyhttps://dataintelo.com/privacy-and-policy

    Time period covered
    2025 - 2034
    Area covered
    United Kingdom, Germany, South Korea, Worldwide, China, Japan, France, United States
    Description

    Real time subtitles market valued at $4.2 billion in 2025, projected to reach $9.8 billion by 2034 at 10.5% CAGR, driven by accessibility demands and media expansion.

  17. movie-subtitles

    • kaggle.com
    zip
    Updated May 29, 2021
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    The citation is currently not available for this dataset.
    Explore at:
    zip(3526841 bytes)Available download formats
    Dataset updated
    May 29, 2021
    Authors
    Peter Bajko
    Description

    Dataset

    This dataset was created by Peter Bajko

    Contents

  18. m

    Helsinki-NLP/open_subtitles

    • metatext.io
    parquet
    Updated Sep 15, 2026
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Helsinki-NLP (2026). Helsinki-NLP/open_subtitles [Dataset]. https://metatext.io/datasets/helsinki-nlp/open_subtitles
    Explore at:
    parquetAvailable download formats
    Dataset updated
    Sep 15, 2026
    Dataset authored and provided by
    Helsinki-NLP
    Area covered
    Helsinki
    Variables measured
    Translation
    Description

    This is a new collection of translated movie subtitles from http://www.opensubtitles.org/.

    IMPORTANT: If you use the OpenSubtitle corpus: Please, add a link to http://www.opensubtitles.org/ to your website and to your reports and publications pro...

  19. s

    Subtitle Translation on ko-zh subtitle dataset

    • sota2.com
    Updated Jan 30, 2026
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    (2026). Subtitle Translation on ko-zh subtitle dataset [Dataset]. https://www.sota2.com/research/sota/subtitle-translation-on-ko-zh-subtitle-dataset
    Explore at:
    Dataset updated
    Jan 30, 2026
    Variables measured
    PA Score (%), TC Score (%), Vivacity Score, Naturalness Score, Translation Score
    Description

    Evaluation of Subtitle Translation performance on the 'ko-zh subtitle dataset', measured using PA (%), TC (%), Trans., Nat, and Vivi scores.

  20. FILMS Dataset

    • osf.io
    Updated Mar 16, 2026
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Elizaveta Sineva; Sara Chilson; Xenia Schmalz (2026). FILMS Dataset [Dataset]. https://osf.io/rd7p6
    Explore at:
    Dataset updated
    Mar 16, 2026
    Dataset provided by
    Center for Open Sciencehttps://cos.io/
    Authors
    Elizaveta Sineva; Sara Chilson; Xenia Schmalz
    Description

    Word Frequency IPA MultiLingual Subtitles Dataset (FILMS Dataset) is a frequency list dataset based on the movie subtitles data taken from OpenSubtitles corpus (v2024).

Share
FacebookFacebook
TwitterTwitter
Email
Click to copy link
Link copied
Close
Cite
Dali Selmi (2023). French Conversations (from movie subtitles) [Dataset]. https://www.kaggle.com/datasets/daliselmi/french-conversational-dataset
Organization logo

French Conversations (from movie subtitles)

French Conversational Dataset, extracted from OpenSubtitles movie subtitles

Explore at:
zip(2880370702 bytes)Available download formats
Dataset updated
Aug 3, 2023
Authors
Dali Selmi
License

https://creativecommons.org/publicdomain/zero/1.0/https://creativecommons.org/publicdomain/zero/1.0/

Area covered
French
Description

French Movie Subtitle Conversations Dataset

Description

Dive into the world of French dialogue with the French Movie Subtitle Conversations dataset โ€“ a comprehensive collection of over 127,000 movie subtitle conversations. This dataset offers a deep exploration of authentic and diverse conversational contexts spanning various genres, eras, and scenarios. It is thoughtfully organized into three distinct sets: training, testing, and validation.

Content Overview

Each conversation in this dataset is structured as a JSON object, featuring three key attributes:

  1. Context: Get a holistic view of the conversation's flow with the preceding 9 lines of dialogue. This context provides invaluable insights into the conversation's dynamics and contextual cues.
  2. Knowledge: Immerse yourself in a wide range of thematic knowledge. This dataset covers an array of topics, ensuring that your models receive exposure to diverse information sources for generating well-informed responses.
  3. Response: Explore how characters react and respond across various scenarios. From casual conversations to intense emotional exchanges, this dataset encapsulates the authenticity of genuine human interaction.

Data Sample

Here's a snippet from the dataset to give you an idea of its structure:

[
 {
  "context": [
   "Tu as attendu longtemps?",
   "Oui en effet.",
   "Je pense que c' est grossier pour un premier rencard.",
   // ... (6 more lines of context)
  ],
  "knowledge": "",
  "response": "On n' avait pas dit 9h?"
 },
 // ... (more data samples)
]

Use Cases

The French Movie Subtitle Conversations dataset serves as a valuable resource for several applications:

  • Conversational AI: Train advanced chatbots and dialogue systems in French that can engage users in fluid, contextually aware conversations.
  • Language Modeling: Enhance your language models by leveraging diverse dialogue patterns, colloquialisms, and contextual dependencies present in real-world conversations.
  • Sentiment Analysis: Investigate the emotional tones of conversations across different movie genres and periods, contributing to a better understanding of sentiment variation.

Why This Dataset

  • Size and Diversity: With a vast collection of over 127,000 conversations spanning diverse genres and tones, this dataset offers an unparalleled breadth and depth in French dialogue data.
  • Contextual Richness: The inclusion of context empowers researchers and practitioners to explore the dynamics of conversation flow, leading to more accurate and contextually relevant responses.
  • Real-world Relevance: Originating from movie subtitles, this dataset mirrors real-world interactions, making it a valuable asset for training models that understand and generate human-like dialogue.

Acknowledgments

We extend our gratitude to the movie subtitle community for their contributions, which have enabled the creation of this diverse and comprehensive French dialogue dataset.

Unlock the potential of authentic French conversations today with the French Movie Subtitle Conversations dataset. Engage in state-of-the-art research, enhance language models, and create applications that resonate with the nuances of real dialogue.

Search
Clear search
Close search
Google apps
Main menu