Facebook
Twitterhttps://creativecommons.org/publicdomain/zero/1.0/https://creativecommons.org/publicdomain/zero/1.0/
Dive into the world of French dialogue with the French Movie Subtitle Conversations dataset โ a comprehensive collection of over 127,000 movie subtitle conversations. This dataset offers a deep exploration of authentic and diverse conversational contexts spanning various genres, eras, and scenarios. It is thoughtfully organized into three distinct sets: training, testing, and validation.
Each conversation in this dataset is structured as a JSON object, featuring three key attributes:
Here's a snippet from the dataset to give you an idea of its structure:
[
{
"context": [
"Tu as attendu longtemps?",
"Oui en effet.",
"Je pense que c' est grossier pour un premier rencard.",
// ... (6 more lines of context)
],
"knowledge": "",
"response": "On n' avait pas dit 9h?"
},
// ... (more data samples)
]
The French Movie Subtitle Conversations dataset serves as a valuable resource for several applications:
We extend our gratitude to the movie subtitle community for their contributions, which have enabled the creation of this diverse and comprehensive French dialogue dataset.
Unlock the potential of authentic French conversations today with the French Movie Subtitle Conversations dataset. Engage in state-of-the-art research, enhance language models, and create applications that resonate with the nuances of real dialogue.
Facebook
TwitterAttribution 4.0 (CC BY 4.0)https://creativecommons.org/licenses/by/4.0/
License information was derived automatically
## Overview
Subtitles is a dataset for object detection tasks - it contains Letters annotations for 500 images.
## Getting Started
You can download this dataset for use within your own projects, or fork it into a workspace on Roboflow to create your own model.
## License
This dataset is available under the [CC BY 4.0 license](https://creativecommons.org/licenses/CC BY 4.0).
Facebook
Twitterchenrm/yyets-subtitles dataset hosted on Hugging Face and contributed by the HF Datasets community
Facebook
TwitterMIT Licensehttps://opensource.org/licenses/MIT
License information was derived automatically
Over 12k scraped YouTube EN subtitles for videos on GitHub topics.
How? Based on the topics https://github.com/topics I searched YouTube with the phrase "What is {topic}?" and downloaded up to 100 video subtitles for a given topic. The extracted text can be found in the dataset together with the topic name, video title and video URL.
Why? I wan to know if we can rate videos based on their information value, especially when we use YouTube as an information source.
You can find the source code here: https://github.com/detrin/text-info-value
Facebook
TwitterSurvivor Subtitles Dataset (cleaned)
Dataset Description
A collection of subtitles from the American reality television show "Survivor", spanning seasons 1 through 47. The dataset contains subtitle text extracted from episode broadcasts. This dataset is a modification of the original Survivor Subtitles dataset after cleaning up and joining subtitle fragments. This dataset is a work in progress and any contributions are welcome.
Source
The subtitles wereโฆ See the full description on the dataset page: https://huggingface.co/datasets/hipml/survivor-subtitles-cleaned.
Facebook
TwitterMIT Licensehttps://opensource.org/licenses/MIT
License information was derived automatically
This dataset contains 25 aligned movie subtitle segments in English, Russian, and Italian, extracted from the ParTree corpus. Each row provides a short, context-rich movie line with its translations in all three languages, making it ideal for research and development in machine translation, multilingual NLP, and cross-lingual transfer learning.
Key features: - Parallel triplets: English, Russian, Italian - Sourced from authentic movie subtitles for natural, conversational language - Suitable for training, validation, and benchmarking of translation and multilingual models
Data originally from the ParTree corpus, available via Swiss-AL
Facebook
TwitterEnhancement of comprehension and more profound understanding of accents were the most common reasons why American adults use subtitles while watching TV in a known language, according to a survey conducted between June and July 2023. Another ** percent of the respondents stated that they did so because they were in a noisy environment.
Facebook
TwitterA multilingual corpus of movie subtitles aligned on the sentence-level. Contains data on more than 50 languages with a focus on the Indo-European language family. Morphosyntactic annotation (part-of-speech, features, dependencies) in Universal Dependency-style is available for 47 languages.
Facebook
Twitterhttps://dataintelo.com/privacy-and-policyhttps://dataintelo.com/privacy-and-policy
The AI Subtitle Generation market was valued at $3.8 billion in 2025 and is projected to reach $18.6 billion by 2034, growing at a CAGR of 19.3%.
Facebook
TwitterEvaluation of Machine Translation performance on the test split of the Subtitles dataset, measured by BLEU and COMET scores.
Facebook
TwitterAttribution 4.0 (CC BY 4.0)https://creativecommons.org/licenses/by/4.0/
License information was derived automatically
Exploring language usage through frequency analysis in large corpora is a defining feature in most recent work in corpus and computational linguistics. From a psycholinguistic perspective, however, the corpora used in these contributions are often not representative of language usage: they are either domain-specific, limited in size, or extracted from unreliable sources. In an effort to address this limitation, we introduce SubIMDB, a corpus of everyday language spoken text we created which contains over 225 million words. The corpus was extracted from 38,102 subtitles of family, comedy and children movies and series, and is the first sizeable structured corpus of subtitles made available. Our experiments show that word frequency norms extracted from this corpus are more effective than those from well-known norms such as Kucera-Francis, HAL and SUBTLEXus in predicting various psycholinguistic properties of words, such as lexical decision times, familiarity, age of acquisition and simplicity. We also provide evidence that contradict the long-standing assumption that the ideal size for a corpus can be determined solely based on how well its word frequencies correlate with lexical decision times.
Facebook
TwitterPJMixers-Dev/Subtitles dataset hosted on Hugging Face and contributed by the HF Datasets community
Facebook
TwitterThis dataset was created by Sergey Sohackiy
Facebook
TwitterEvaluation of video understanding capabilities using the Video-MME dataset, specifically excluding subtitle information, reporting scores segmented by video length (Short, Medium, Long) and overall.
Facebook
TwitterAs of September 5, 2026, GPT-4.1 is #1 for Video Mme Long No Subtitles at 72%. Ranked by the Video Mme Long No Subtitles score Video Mme Long No Subtitles leaderboard: rank models by Video Mme Long No Subtitles next to live API token prices.
Facebook
Twitterhttps://dataintelo.com/privacy-and-policyhttps://dataintelo.com/privacy-and-policy
Real time subtitles market valued at $4.2 billion in 2025, projected to reach $9.8 billion by 2034 at 10.5% CAGR, driven by accessibility demands and media expansion.
Facebook
TwitterThis is a new collection of translated movie subtitles from http://www.opensubtitles.org/.
IMPORTANT: If you use the OpenSubtitle corpus: Please, add a link to http://www.opensubtitles.org/ to your website and to your reports and publications pro...
Facebook
TwitterEvaluation of Subtitle Translation performance on the 'ko-zh subtitle dataset', measured using PA (%), TC (%), Trans., Nat, and Vivi scores.
Facebook
TwitterWord Frequency IPA MultiLingual Subtitles Dataset (FILMS Dataset) is a frequency list dataset based on the movie subtitles data taken from OpenSubtitles corpus (v2024).
Facebook
Twitterhttps://creativecommons.org/publicdomain/zero/1.0/https://creativecommons.org/publicdomain/zero/1.0/
Dive into the world of French dialogue with the French Movie Subtitle Conversations dataset โ a comprehensive collection of over 127,000 movie subtitle conversations. This dataset offers a deep exploration of authentic and diverse conversational contexts spanning various genres, eras, and scenarios. It is thoughtfully organized into three distinct sets: training, testing, and validation.
Each conversation in this dataset is structured as a JSON object, featuring three key attributes:
Here's a snippet from the dataset to give you an idea of its structure:
[
{
"context": [
"Tu as attendu longtemps?",
"Oui en effet.",
"Je pense que c' est grossier pour un premier rencard.",
// ... (6 more lines of context)
],
"knowledge": "",
"response": "On n' avait pas dit 9h?"
},
// ... (more data samples)
]
The French Movie Subtitle Conversations dataset serves as a valuable resource for several applications:
We extend our gratitude to the movie subtitle community for their contributions, which have enabled the creation of this diverse and comprehensive French dialogue dataset.
Unlock the potential of authentic French conversations today with the French Movie Subtitle Conversations dataset. Engage in state-of-the-art research, enhance language models, and create applications that resonate with the nuances of real dialogue.