Saved datasets
1 dataset found
  1. h

    YouTube-Commons

    • huggingface.co
    Updated Apr 17, 2024
    + more versions
  2. Not seeing a result you expected?
    Learn how you can add new datasets to our index.

Share
FacebookFacebook
TwitterTwitter
Email
Click to copy link
Link copied
Close
Cite
PleIAs (2024). YouTube-Commons [Dataset]. https://huggingface.co/datasets/PleIAs/YouTube-Commons

YouTube-Commons

PleIAs/YouTube-Commons

Youtube Commons Corpus

Explore at:
24 scholarly articles cite this dataset (View in Google Scholar)
Dataset updated
Apr 17, 2024
Dataset authored and provided by
PleIAs
License

Attribution 4.0 (CC BY 4.0)https://creativecommons.org/licenses/by/4.0/
License information was derived automatically

Area covered
YouTube
Description

📺 YouTube-Commons 📺

YouTube-Commons is a collection of audio transcripts of 2,063,066 videos shared on YouTube under a CC-By license.

  Content

The collection comprises 22,709,724 original and automatically translated transcripts from 3,156,703 videos (721,136 individual channels). In total, this represents nearly 45 billion words (44,811,518,375). All the videos where shared on YouTube with a CC-BY license: the dataset provide all the necessary provenance information… See the full description on the dataset page: https://huggingface.co/datasets/PleIAs/YouTube-Commons.

Search
Clear search
Close search
Google apps
Main menu