dm-petrov/youtube-commons-small
๐บ YouTube-Commons-Small ๐บ This is a smaller subset of the YouTube-Commons dataset, which is a collection of audio transcripts from videos shared on YouTube under a CC-By license. Dataset Description This smaller version contains a subset of the original dataset, maintaining the same structure and features. It's designed for easier experimentation and testing purposes. Features The dataset includes the following information for each video: Videoโฆ See the full description on the dataset page: https://huggingface.co/datasets/dm-petrov/youtube-commons-small.
๐บ YouTube-Commons-Small ๐บ
This is a smaller subset of the YouTube-Commons dataset, which is a collection of audio transcripts from videos shared on YouTube under a CC-By license.
Dataset Description
This smaller version contains a subset of the original dataset, maintaining the same structure and features. It's designed for easier experimentation and testing purposes.
Features
The dataset includes the following information for each video:
- Video ID and link
- Title and text transcript
- Channel information
- Upload date
- License information
- Language information (original, source, and transcription)
- Word and character counts
Original Dataset Attribution
This dataset is derived from the YouTube-Commons dataset created by PleIAs. The original dataset was built with the support of:
- Scaleway
- LANGU:IA (French Ministry of Culture and DINUM)
- Alliance for Language technologies EDIC (ALT-EDIC)
- Open science LLM community (Occiglot, Eleuther AI, Allen AI)
License
This dataset is licensed under CC-BY-4.0, the same license as the original dataset. All content must be properly attributed to the original creators as specified in the video metadata.
Content
The collection comprises 22,709,724 original and automatically translated transcripts from 3,156,703 videos (721,136 individual channels).
In total, this represents nearly 45 billion words (44,811,518,375).
All the videos where shared on YouTube with a CC-BY license: the dataset provide all the necessary provenance information including the title, link, channel name and upload date.
The corpus is multilingual with a majority of English-speaking content (71%) for original languages. Automated translations are provided for nearly all the videos in English, French, Spanish, German, Russian, Italian and Dutch.
Uses
The collection aims to expand the availability of conversational data for research in AI, computational social science and digital humanities.
Most of the available resources under free licenses are written texts such as public domain works or open science articles.
The text can be used for training model and republished with for reproducibility purposes.
License and ethics
All the transcripts are part of a video shared under a CC-By license. In accordance with the provision of the license, every YouTube channels is fully credited.
While content under a free license can be lawfully reproduced in any setting, there is currently a debate over the legitimacy and proper ethical use of free content for pre-training large language models.
In accordance with the philosophy of Creative Commons, we recommend that this set be preferably used for open research. Furthermore, the license requires that contribution of each individual author is properly credited. In a research context, the best way to achieve this aim would be to fully release the data sources used for training or, at the very least, provide an extensive open documentation.
Future developments
The collection is far from covering the total amount of available YouTube videos under a Creative Commons license. We will continue to expand it significantly.
Other additional release will also focus on transcripts from other video sources not available on YouTube (especially from public service/university websites).
Acknowledgements
The corpus was stored and processed with the generous support of Scaleway. It was built up with the support and concerted efforts of the state start-up LANGU:IA (start-up d'Etat), supported by the French Ministry of Culture and DINUM, as part of the prefiguration of the service offering of the Alliance for Language technologies EDIC (ALT-EDIC).
Pleias corpus collection projects have been also facilitated thanks to the open science LLM community support, insights and cooperation (Occiglot, Eleuther AI, Allen AI).
<div style="text-align: center;"> <img src="https://github.com/mch-dd/datasetlogo/blob/main/scaleway.jpeg?raw=true" style="width: 33%; margin: 0 auto; display: inline-block;"/> <img src="https://github.com/mch-dd/datasetlogo/blob/main/ministere.png?raw=true" style="width: 33%; margin: 0 auto; display: inline-block;"/> <img src="https://github.com/mch-dd/datasetlogo/blob/main/occiglot.jpg?raw=true" style="width: 33%; margin: 0 auto; display: inline-block;"/> </div>
