CoolFace
Datasetpublic

komats/mega-ssum

Mega-SSum A large-scale English sentence-wise speech summarization (Sen-SSum) dataset Consists of 3.8M+ synthesized speech, transcription, summary triplets Derived from the Gigaword dataset Rush+2015 Overview The dataset is divided into five splits: train/core/dev/eval/duc2003. (See below table) We added a new evaluation split "test" for in-domain evaluation. The train split is here: MegaSSum(train). orig. data split #samples #speakers total dur.… See the full description on the dataset page: https://huggingface.co/datasets/komats/mega-ssum.

sourceHugging Facecc-by-4.0updated 2y agoView on Hugging Face
3likes314downloads

komats/mega-ssum · main · files are served by the source, never re-hosted here