100+ datasets found
  1. French Conversations (from movie subtitles)

    • kaggle.com
    zip
    Updated Aug 3, 2023
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Dali Selmi (2023). French Conversations (from movie subtitles) [Dataset]. https://www.kaggle.com/datasets/daliselmi/french-conversational-dataset
    Explore at:
    zip(2880370702 bytes)Available download formats
    Dataset updated
    Aug 3, 2023
    Authors
    Dali Selmi
    License

    https://creativecommons.org/publicdomain/zero/1.0/https://creativecommons.org/publicdomain/zero/1.0/

    Area covered
    French
    Description

    French Movie Subtitle Conversations Dataset

    Description

    Dive into the world of French dialogue with the French Movie Subtitle Conversations dataset – a comprehensive collection of over 127,000 movie subtitle conversations. This dataset offers a deep exploration of authentic and diverse conversational contexts spanning various genres, eras, and scenarios. It is thoughtfully organized into three distinct sets: training, testing, and validation.

    Content Overview

    Each conversation in this dataset is structured as a JSON object, featuring three key attributes:

    1. Context: Get a holistic view of the conversation's flow with the preceding 9 lines of dialogue. This context provides invaluable insights into the conversation's dynamics and contextual cues.
    2. Knowledge: Immerse yourself in a wide range of thematic knowledge. This dataset covers an array of topics, ensuring that your models receive exposure to diverse information sources for generating well-informed responses.
    3. Response: Explore how characters react and respond across various scenarios. From casual conversations to intense emotional exchanges, this dataset encapsulates the authenticity of genuine human interaction.

    Data Sample

    Here's a snippet from the dataset to give you an idea of its structure:

    [
     {
      "context": [
       "Tu as attendu longtemps?",
       "Oui en effet.",
       "Je pense que c' est grossier pour un premier rencard.",
       // ... (6 more lines of context)
      ],
      "knowledge": "",
      "response": "On n' avait pas dit 9h?"
     },
     // ... (more data samples)
    ]
    

    Use Cases

    The French Movie Subtitle Conversations dataset serves as a valuable resource for several applications:

    • Conversational AI: Train advanced chatbots and dialogue systems in French that can engage users in fluid, contextually aware conversations.
    • Language Modeling: Enhance your language models by leveraging diverse dialogue patterns, colloquialisms, and contextual dependencies present in real-world conversations.
    • Sentiment Analysis: Investigate the emotional tones of conversations across different movie genres and periods, contributing to a better understanding of sentiment variation.

    Why This Dataset

    • Size and Diversity: With a vast collection of over 127,000 conversations spanning diverse genres and tones, this dataset offers an unparalleled breadth and depth in French dialogue data.
    • Contextual Richness: The inclusion of context empowers researchers and practitioners to explore the dynamics of conversation flow, leading to more accurate and contextually relevant responses.
    • Real-world Relevance: Originating from movie subtitles, this dataset mirrors real-world interactions, making it a valuable asset for training models that understand and generate human-like dialogue.

    Acknowledgments

    We extend our gratitude to the movie subtitle community for their contributions, which have enabled the creation of this diverse and comprehensive French dialogue dataset.

    Unlock the potential of authentic French conversations today with the French Movie Subtitle Conversations dataset. Engage in state-of-the-art research, enhance language models, and create applications that resonate with the nuances of real dialogue.

  2. Movie Subtitle Dataset

    • kaggle.com
    zip
    Updated Aug 8, 2021
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Adiamaan (2021). Movie Subtitle Dataset [Dataset]. https://www.kaggle.com/adiamaan/movie-subtitle-dataset
    Explore at:
    zip(254871718 bytes)Available download formats
    Dataset updated
    Aug 8, 2021
    Authors
    Adiamaan
    License

    https://creativecommons.org/publicdomain/zero/1.0/https://creativecommons.org/publicdomain/zero/1.0/

    Description

    💡 Motive

    I was thinking about movie sentiments and wanted to see if there is any strong pattern behind how sentiment fluctuates across the movie to how that movie is received or performed.

    🍎 Lowest hanging fruit

    To track movie sentiments across the run time, the easy way is to get the movie subtitles and identify the sentiment for each text in the subtitle. The advantage of this approach is that movie subtitles are easy to get, parse, and process and NLP frameworks can easily help with the task. This approach is scalable since irrespective of language, english subtitles are available for almost all movies albeit translation errors.

  3. R

    Subtitles Dataset

    • universe.roboflow.com
    zip
    Updated Oct 9, 2022
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    subtitles (2022). Subtitles Dataset [Dataset]. https://universe.roboflow.com/subtitles-jtdc8/subtitles-xmseb/dataset/3
    Explore at:
    zipAvailable download formats
    Dataset updated
    Oct 9, 2022
    Dataset authored and provided by
    subtitles
    License

    Attribution 4.0 (CC BY 4.0)https://creativecommons.org/licenses/by/4.0/
    License information was derived automatically

    Variables measured
    Letters Bounding Boxes
    Description

    Subtitles

    ## Overview
    
    Subtitles is a dataset for object detection tasks - it contains Letters annotations for 500 images.
    
    ## Getting Started
    
    You can download this dataset for use within your own projects, or fork it into a workspace on Roboflow to create your own model.
    
      ## License
    
      This dataset is available under the [CC BY 4.0 license](https://creativecommons.org/licenses/CC BY 4.0).
    
  4. h

    Subtitles

    • huggingface.co
    Updated Apr 4, 2009
    + more versions
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Peanut Jar Mixers Development (2009). Subtitles [Dataset]. https://huggingface.co/datasets/PJMixers-Dev/Subtitles
    Explore at:
    CroissantCroissant is a format for machine-learning datasets. Learn more about this at mlcommons.org/croissant.
    Dataset updated
    Apr 4, 2009
    Dataset authored and provided by
    Peanut Jar Mixers Development
    Description

    PJMixers-Dev/Subtitles dataset hosted on Hugging Face and contributed by the HF Datasets community

  5. Movie Subtitles

    • kaggle.com
    zip
    Updated Nov 26, 2024
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Ahwar (2024). Movie Subtitles [Dataset]. https://www.kaggle.com/datasets/ahwardev/movie-subtitles
    Explore at:
    zip(133455 bytes)Available download formats
    Dataset updated
    Nov 26, 2024
    Authors
    Ahwar
    Description

    This dataset includes subtitle files in the SRT (SubRip Subtitle) format for several popular movies, such as Oppenheimer and Tenet. SRT files are plain-text files widely used for subtitles, containing a series of structured entries to synchronize text with video content. Each entry in an SRT file comprises:

    1. A sequential index number to indicate the order of the subtitles.
    2. Timestamps that specify when a subtitle should appear and disappear, formatted as hours:minutes:seconds,milliseconds.
    3. The subtitle text, which is displayed during the designated time interval.

    For example:
    ``` 1 00:00:01,500 --> 00:00:04,000 This is a sample subtitle.

    2 00:00:04,500 --> 00:00:07,000 Here is another subtitle to demonstrate multiple entries.

    
    This straightforward format is highly compatible with media players and easy to edit. SRT files enhance accessibility by providing subtitles for different languages or accommodating viewers with hearing impairments, enriching the experience of enjoying these popular movies.
    
  6. Movie Parallel Subtitles (EN-IT-RU)

    • kaggle.com
    zip
    Updated Jun 26, 2025
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Timur Sharifullin (2025). Movie Parallel Subtitles (EN-IT-RU) [Dataset]. https://www.kaggle.com/datasets/timursharifullindata/movie-parallel-subtitles-small-sentiment-dataset
    Explore at:
    zip(6056 bytes)Available download formats
    Dataset updated
    Jun 26, 2025
    Authors
    Timur Sharifullin
    License

    MIT Licensehttps://opensource.org/licenses/MIT
    License information was derived automatically

    Description

    This dataset contains 25 aligned movie subtitle segments in English, Russian, and Italian, extracted from the ParTree corpus. Each row provides a short, context-rich movie line with its translations in all three languages, making it ideal for research and development in machine translation, multilingual NLP, and cross-lingual transfer learning.

    Key features: - Parallel triplets: English, Russian, Italian - Sourced from authentic movie subtitles for natural, conversational language - Suitable for training, validation, and benchmarking of translation and multilingual models

    Data originally from the ParTree corpus, available via Swiss-AL

  7. h

    survivor-subtitles-cleaned

    • huggingface.co
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Paul Lambert, survivor-subtitles-cleaned [Dataset]. https://huggingface.co/datasets/hipml/survivor-subtitles-cleaned
    Explore at:
    CroissantCroissant is a format for machine-learning datasets. Learn more about this at mlcommons.org/croissant.
    Authors
    Paul Lambert
    Description

    Survivor Subtitles Dataset (cleaned)

      Dataset Description
    

    A collection of subtitles from the American reality television show "Survivor", spanning seasons 1 through 47. The dataset contains subtitle text extracted from episode broadcasts. This dataset is a modification of the original Survivor Subtitles dataset after cleaning up and joining subtitle fragments. This dataset is a work in progress and any contributions are welcome.

      Source
    

    The subtitles were… See the full description on the dataset page: https://huggingface.co/datasets/hipml/survivor-subtitles-cleaned.

  8. D

    AI Subtitle Generation Market Research Report 2034

    • dataintelo.com
    csv, pdf, pptx
    Updated Mar 21, 2026
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Dataintelo (2026). AI Subtitle Generation Market Research Report 2034 [Dataset]. https://dataintelo.com/report/ai-subtitle-generation-market
    Explore at:
    csv, pdf, pptxAvailable download formats
    Dataset updated
    Mar 21, 2026
    Dataset authored and provided by
    Dataintelo
    License

    https://dataintelo.com/privacy-and-policyhttps://dataintelo.com/privacy-and-policy

    Time period covered
    2025 - 2034
    Area covered
    United States, China, France, Germany, Worldwide, South Korea, Japan, United Kingdom
    Description

    The AI Subtitle Generation market was valued at $3.8 billion in 2025 and is projected to reach $18.6 billion by 2034, growing at a CAGR of 19.3%.

  9. SubIMDB: A Structured Corpus of Subtitles

    • zenodo.org
    • live.european-language-grid.eu
    tar
    Updated Jan 24, 2020
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Paetzold; Specia; Paetzold; Specia (2020). SubIMDB: A Structured Corpus of Subtitles [Dataset]. http://doi.org/10.5281/zenodo.2552407
    Explore at:
    tarAvailable download formats
    Dataset updated
    Jan 24, 2020
    Dataset provided by
    Zenodohttp://zenodo.org/
    Authors
    Paetzold; Specia; Paetzold; Specia
    License

    Attribution 4.0 (CC BY 4.0)https://creativecommons.org/licenses/by/4.0/
    License information was derived automatically

    Description

    Exploring language usage through frequency analysis in large corpora is a defining feature in most recent work in corpus and computational linguistics. From a psycholinguistic perspective, however, the corpora used in these contributions are often not representative of language usage: they are either domain-specific, limited in size, or extracted from unreliable sources. In an effort to address this limitation, we introduce SubIMDB, a corpus of everyday language spoken text we created which contains over 225 million words. The corpus was extracted from 38,102 subtitles of family, comedy and children movies and series, and is the first sizeable structured corpus of subtitles made available. Our experiments show that word frequency norms extracted from this corpus are more effective than those from well-known norms such as Kucera-Francis, HAL and SUBTLEXus in predicting various psycholinguistic properties of words, such as lexical decision times, familiarity, age of acquisition and simplicity. We also provide evidence that contradict the long-standing assumption that the ideal size for a corpus can be determined solely based on how well its word frequencies correlate with lexical decision times.

  10. h

    Subtitles-rag-questions-qwq-all-aphrodite

    • huggingface.co
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Peanut Jar Mixers Development, Subtitles-rag-questions-qwq-all-aphrodite [Dataset]. https://huggingface.co/datasets/PJMixers-Dev/Subtitles-rag-questions-qwq-all-aphrodite
    Explore at:
    Dataset authored and provided by
    Peanut Jar Mixers Development
    Description

    PJMixers-Dev/Subtitles-rag-questions-qwq-all-aphrodite dataset hosted on Hugging Face and contributed by the HF Datasets community

  11. h

    survivor-subtitles

    • huggingface.co
    Updated Jan 3, 2025
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Paul Lambert (2025). survivor-subtitles [Dataset]. https://huggingface.co/datasets/hipml/survivor-subtitles
    Explore at:
    CroissantCroissant is a format for machine-learning datasets. Learn more about this at mlcommons.org/croissant.
    Dataset updated
    Jan 3, 2025
    Authors
    Paul Lambert
    License

    Attribution-ShareAlike 4.0 (CC BY-SA 4.0)https://creativecommons.org/licenses/by-sa/4.0/
    License information was derived automatically

    Description

    Survivor Subtitles Dataset

      Dataset Description
    

    A collection of subtitles from the American reality television show "Survivor", spanning seasons 1 through 47. The dataset contains subtitle text extracted from episode broadcasts.

      Source
    

    The subtitles were obtained from OpenSubtitles.com.

      Dataset Details
    

    Coverage:

    Seasons: 1-47 Episodes per season: ~13-14 Total episodes: ~600

    Format:

    Text files containing timestamped subtitle data Character… See the full description on the dataset page: https://huggingface.co/datasets/hipml/survivor-subtitles.

  12. FILMS Dataset

    • osf.io
    Updated Mar 16, 2026
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Elizaveta Sineva; Sara Chilson; Xenia Schmalz (2026). FILMS Dataset [Dataset]. https://osf.io/rd7p6
    Explore at:
    Dataset updated
    Mar 16, 2026
    Dataset provided by
    Center for Open Sciencehttps://cos.io/
    Authors
    Elizaveta Sineva; Sara Chilson; Xenia Schmalz
    Description

    Word Frequency IPA MultiLingual Subtitles Dataset (FILMS Dataset) is a frequency list dataset based on the movie subtitles data taken from OpenSubtitles corpus (v2024).

  13. Movie Parallel Subtitles

    • kaggle.com
    zip
    Updated Nov 12, 2023
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    August murr (2023). Movie Parallel Subtitles [Dataset]. https://www.kaggle.com/datasets/augustmurr/movie-parallel-dataset
    Explore at:
    zip(144446662 bytes)Available download formats
    Dataset updated
    Nov 12, 2023
    Authors
    August murr
    License

    https://creativecommons.org/publicdomain/zero/1.0/https://creativecommons.org/publicdomain/zero/1.0/

    Description

    The dataset comprises parallel subtitle files sourced from Subscene.com, encompassing four language pairs: "English to Arabic," "English to French," "English to Indonesian," and "English to Thai."

    Each folder corresponds to a specific movie title and contains two CSV files: "parallel_line_by_line" and "parallel_time_based."

    • Parallel_line_by_line: Each row in this file features two columns, one in English and the other presenting the translation in another language.

    • Parallel_time_based: Each row encapsulates all the subtitle text occurring within one minute of the movie, alongside its translation.

    Additionally, the notebook utilized for web scraping, collection, and alignment of these files is accessible via this GitHub link: GitHub - Parallel Subtitle Collection and Alignment. Utilize the code available in the notebook to download parallel subtitles for your preferred movie and language pair.

    The collection script operated across a selection of the highest grossing movies, ensuring inclusion of a comprehensive range of popular movies.

  14. r

    Opus Subtitles Corpus

    • resodate.org
    Updated Jan 28, 2016
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Jörg Tiedemann (2016). Opus Subtitles Corpus [Dataset]. https://resodate.org/resources/aHR0cHM6Ly91cm4uZmkvdXJuOm5ibjpmaTpsYi0yMDE2MDEyODA0
    Explore at:
    Dataset updated
    Jan 28, 2016
    Dataset provided by
    Research.fi
    Kielipankki
    University of Helsinki
    Authors
    Jörg Tiedemann
    Description

    The corpus, containing the OpenSubtitles sub-corpora of the Opus open parallel corpus (http://opus.lingfil.uu.se/), will be made available for download at https://korp.csc.fi/download/

  15. Reasons why adults use subtitles when watching TV in known language in the...

    • statista.com
    Updated Nov 27, 2025
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Statista (2025). Reasons why adults use subtitles when watching TV in known language in the U.S. 2023 [Dataset]. https://www.statista.com/statistics/1459167/reasons-use-subtitles-watching-tv-known-language-us/
    Explore at:
    Dataset updated
    Nov 27, 2025
    Dataset authored and provided by
    Statistahttps://statista.com/
    Time period covered
    Jun 29, 2023 - Jul 5, 2023
    Area covered
    United States
    Description

    Enhancement of comprehension and more profound understanding of accents were the most common reasons why American adults use subtitles while watching TV in a known language, according to a survey conducted between June and July 2023. Another ** percent of the respondents stated that they did so because they were in a noisy environment.

  16. youtube subtitles

    • kaggle.com
    zip
    Updated Sep 25, 2022
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    wb-08 (2022). youtube subtitles [Dataset]. https://www.kaggle.com/datasets/wadzim/youtube-subtitles/data
    Explore at:
    zip(2436353180 bytes)Available download formats
    Dataset updated
    Sep 25, 2022
    Authors
    wb-08
    Area covered
    YouTube
    Description

    Dataset

    This dataset was created by wb-08

    Contents

    Labeled subtitles from YouTube in YOLO format

  17. subtitles

    • kaggle.com
    zip
    Updated Aug 1, 2023
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Sergey Sohackiy (2023). subtitles [Dataset]. https://www.kaggle.com/sergeysohackiy/subtitles
    Explore at:
    zip(9239834 bytes)Available download formats
    Dataset updated
    Aug 1, 2023
    Authors
    Sergey Sohackiy
    Description

    Dataset

    This dataset was created by Sergey Sohackiy

    Contents

  18. Data from: ChatSubs: A dataset of movie dialogues in Spanish, Catalan,...

    • zenodo.org
    • produccioncientifica.ugr.es
    • +1more
    gz, txt
    Updated Aug 7, 2023
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    Ksenia Kharitonova; Ksenia Kharitonova; Zoraida Callejas; Zoraida Callejas; David Pérez-Fernández; David Pérez-Fernández; Asier Gutiérrez-Fandiño; Asier Gutiérrez-Fandiño; David Griol; David Griol (2023). ChatSubs: A dataset of movie dialogues in Spanish, Catalan, Basque and Galician [Dataset]. http://doi.org/10.5281/zenodo.8192331
    Explore at:
    gz, txtAvailable download formats
    Dataset updated
    Aug 7, 2023
    Dataset provided by
    Zenodohttp://zenodo.org/
    Authors
    Ksenia Kharitonova; Ksenia Kharitonova; Zoraida Callejas; Zoraida Callejas; David Pérez-Fernández; David Pérez-Fernández; Asier Gutiérrez-Fandiño; Asier Gutiérrez-Fandiño; David Griol; David Griol
    Description

    Description: The ChatSubs dataset contains dialogues in Spanish and three co-official languages of Spain (Catalan, Basque, and Galician). It was obtained from OpenSubtitles and processed to generate clearly segmented dialogues and turns. The dataset consists of 206,706 JSON files, with over 20 million dialogues and 96 million turns, making it one of the largest dialogue corpora available. It serves as an excellent resource for research teams interested in training dialogue models in Spanish, Catalan, Basque, and Galician.

    License: CC BY-NC 4.0.

  19. v

    Film Subtitling Market Size By Service Type (Translation Services, Technical...

    • verifiedmarketresearch.com
    Updated Jun 15, 2026
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    VERIFIED MARKET RESEARCH (2026). Film Subtitling Market Size By Service Type (Translation Services, Technical Services, Quality Assurance Services), By End-User (Streaming Platforms, Film Studios, Television Networks, Educational Institutions), By Delivery Format (SRT Files, WebVTT Files, Closed Captions, Open Subtitles), By Geographic Scope and Forecast [Dataset]. https://www.verifiedmarketresearch.com/product/film-subtitling-market/
    Explore at:
    Dataset updated
    Jun 15, 2026
    Dataset authored and provided by
    VERIFIED MARKET RESEARCH
    License

    https://www.verifiedmarketresearch.com/privacy-policy/https://www.verifiedmarketresearch.com/privacy-policy/

    Time period covered
    2026 - 2032
    Area covered
    Global
    Description

    Film Subtitling Market size was valued at USD 1.2 Billion in 2024 and is expected to reach USD 2.29 Billion by 2032, growing at a CAGR of 8.5% during the forecast period 2026-2032.Greater service standardization is expected to be supported by disability inclusion mandates, hearing-impaired audience considerations, and government accessibility regulations encouraging comprehensive subtitle provision and inclusive media consumption.

  20. m

    Self-Authored Multilingual Subtitle Alignment Samples for AI Video...

    • data.mendeley.com
    Updated May 12, 2026
    Share
    FacebookFacebook
    TwitterTwitter
    Email
    Click to copy link
    Link copied
    Close
    Cite
    regi maz (2026). Self-Authored Multilingual Subtitle Alignment Samples for AI Video Translation [Dataset]. http://doi.org/10.17632/spzyr66zn3.1
    Explore at:
    Dataset updated
    May 12, 2026
    Authors
    regi maz
    License

    Attribution 4.0 (CC BY 4.0)https://creativecommons.org/licenses/by/4.0/
    License information was derived automatically

    Description

    This dataset contains a rights-cleared collection of self-authored multilingual subtitle alignment samples for evaluating video subtitle translation workflows. The release includes 180 short scripted clips represented as subtitle-like segments, 540 timestamped source segments, and 1,080 aligned translation rows across English, Spanish, and Chinese (Simplified). Supporting documentation includes a clip-level manifest, a machine-readable schema, a field-level data dictionary, methodology notes, a short abstract, and full SRT subtitle files for all clips. The package was designed to support research and workflow evaluation for multilingual video localization, subtitle alignment, translation quality review, and subtitle-aware ingestion pipelines. The material is synthetic in the sense that all source text was authored specifically for this release; however, the record structure reflects common subtitle segmentation patterns, including clip identifiers, segment identifiers, timestamps, language pairs, scenario labels, and aligned text fields. Only derived text annotations and subtitle files are distributed. No third-party videos, audio tracks, platform exports, scraped captions, or copyrighted transcripts are included. No personal data or sensitive information is present in the release. All content is distributed under CC BY 4.0. The package is intended for repository deposit, reproducible documentation, and evaluation of multilingual subtitle processing workflows. It is not intended as a representation of the full distribution of public web video subtitles.

Share
FacebookFacebook
TwitterTwitter
Email
Click to copy link
Link copied
Close
Cite
Dali Selmi (2023). French Conversations (from movie subtitles) [Dataset]. https://www.kaggle.com/datasets/daliselmi/french-conversational-dataset
Organization logo

French Conversations (from movie subtitles)

French Conversational Dataset, extracted from OpenSubtitles movie subtitles

Explore at:
zip(2880370702 bytes)Available download formats
Dataset updated
Aug 3, 2023
Authors
Dali Selmi
License

https://creativecommons.org/publicdomain/zero/1.0/https://creativecommons.org/publicdomain/zero/1.0/

Area covered
French
Description

French Movie Subtitle Conversations Dataset

Description

Dive into the world of French dialogue with the French Movie Subtitle Conversations dataset – a comprehensive collection of over 127,000 movie subtitle conversations. This dataset offers a deep exploration of authentic and diverse conversational contexts spanning various genres, eras, and scenarios. It is thoughtfully organized into three distinct sets: training, testing, and validation.

Content Overview

Each conversation in this dataset is structured as a JSON object, featuring three key attributes:

  1. Context: Get a holistic view of the conversation's flow with the preceding 9 lines of dialogue. This context provides invaluable insights into the conversation's dynamics and contextual cues.
  2. Knowledge: Immerse yourself in a wide range of thematic knowledge. This dataset covers an array of topics, ensuring that your models receive exposure to diverse information sources for generating well-informed responses.
  3. Response: Explore how characters react and respond across various scenarios. From casual conversations to intense emotional exchanges, this dataset encapsulates the authenticity of genuine human interaction.

Data Sample

Here's a snippet from the dataset to give you an idea of its structure:

[
 {
  "context": [
   "Tu as attendu longtemps?",
   "Oui en effet.",
   "Je pense que c' est grossier pour un premier rencard.",
   // ... (6 more lines of context)
  ],
  "knowledge": "",
  "response": "On n' avait pas dit 9h?"
 },
 // ... (more data samples)
]

Use Cases

The French Movie Subtitle Conversations dataset serves as a valuable resource for several applications:

  • Conversational AI: Train advanced chatbots and dialogue systems in French that can engage users in fluid, contextually aware conversations.
  • Language Modeling: Enhance your language models by leveraging diverse dialogue patterns, colloquialisms, and contextual dependencies present in real-world conversations.
  • Sentiment Analysis: Investigate the emotional tones of conversations across different movie genres and periods, contributing to a better understanding of sentiment variation.

Why This Dataset

  • Size and Diversity: With a vast collection of over 127,000 conversations spanning diverse genres and tones, this dataset offers an unparalleled breadth and depth in French dialogue data.
  • Contextual Richness: The inclusion of context empowers researchers and practitioners to explore the dynamics of conversation flow, leading to more accurate and contextually relevant responses.
  • Real-world Relevance: Originating from movie subtitles, this dataset mirrors real-world interactions, making it a valuable asset for training models that understand and generate human-like dialogue.

Acknowledgments

We extend our gratitude to the movie subtitle community for their contributions, which have enabled the creation of this diverse and comprehensive French dialogue dataset.

Unlock the potential of authentic French conversations today with the French Movie Subtitle Conversations dataset. Engage in state-of-the-art research, enhance language models, and create applications that resonate with the nuances of real dialogue.

Search
Clear search
Close search
Google apps
Main menu