komats/mega-ssum
Mega-SSum A large-scale English sentence-wise speech summarization (Sen-SSum) dataset Consists of 3.8M+ synthesized speech, transcription, summary triplets Derived from the Gigaword dataset Rush+2015 Overview The dataset is divided into five splits: train/core/dev/eval/duc2003. (See below table) We added a new evaluation split "test" for in-domain evaluation. The train split is here: MegaSSum(train). orig. data split #samples #speakers total dur.… See the full description on the dataset page: https://huggingface.co/datasets/komats/mega-ssum.
Conversations for this repository live on Hugging Face.
CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.
Open discussions on Hugging Face