datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
YouTube-Commons-nl-audio
YouTube Commons NL Audio
This dataset has audio files for the Dutch-language videos in Rijgersberg/YouTube-Commons-nl-transcriptions,
all under a CC BY 4.0 license.
It contains 11,669 files for a total runtime of 2493h 43m 5s, coming in at approximately 130 GB.
Source
The original source of the dataset (minus the titles, descriptions and audio files) is YouTube Commons:
YouTube-Commons is a collection of audio transcripts of 2,063,066 videos shared on YouTube… See the full description on the dataset page: https://huggingface.co/datasets/Rijgersberg/YouTube-Commons-nl-audio.TTIC-commoncommon-craft-processedcommons-video-vqa
Wikimedia Commons Video VQA
88 short (20–90 s) public-domain videos from Wikimedia Commons annotated with Gemini for visual question answering. Each row has the video (embedded bytes) plus structured annotations.
Columns
Column
Type
Notes
video
Video
Embedded video bytes (mp4)
pageid
string
Wikimedia page id
file
string
Original local path
title, url, duration, mime
metadata
Commons metadata
license, license_url, artist, credit, date
licensing /… See the full description on the dataset page: https://huggingface.co/datasets/Reubencf/commons-video-vqa.common_murre_temporalImages and video provided by the Baltic Seabird Project (http://www.balticseabird.com/).
Annotations by Johannes Hägerlind.
so101-lift-cube-rl-extended-commonThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so101_follower",
"total_episodes": 180,
"total_frames": 16121,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:180"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/igor-saprygin/so101-lift-cube-rl-extended-common.planting_seeds_commonThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "lekiwi_client",
"total_episodes": 6,
"total_frames": 3849,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 120,
"splits": {
"train": "0:6"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/oscarz511/planting_seeds_common.
