CoolFace
Datasetpublic

retkowski/ytseg

YTSeg: A Benchmark for Audio Chaptering and Video Transcript Segmentation We present YTSeg, a topically and structurally diverse benchmark for the audio chaptering and transcript segmentation task based on YouTube videos. The dataset comprises 19,299 videos from 393 channels, amounting to 6,533 content hours. The topics are wide-ranging, covering domains such as science, lifestyle, politics, health, economy, and technology. The videos are from various types of content formats… See the full description on the dataset page: https://huggingface.co/datasets/retkowski/ytseg.

sourceHugging Facecc-by-nc-sa-4.0updated 2mo agoView on Hugging Face
8likes2.6kdownloads
../
filetest-00000-of-00001.parquet75.5 MBdownload
filetrain-00000-of-00004.parquet208.2 MBdownload
filetrain-00001-of-00004.parquet205.1 MBdownload
filetrain-00002-of-00004.parquet210.7 MBdownload
filetrain-00003-of-00004.parquet206.3 MBdownload
filevalidation-00000-of-00001.parquet74.8 MBdownload

retkowski/ytseg · main · files are served by the source, never re-hosted here

retkowski/ytseg · CoolFace