CoolFace
Datasetpublic

komats/mega-ssum

Mega-SSum A large-scale English sentence-wise speech summarization (Sen-SSum) dataset Consists of 3.8M+ synthesized speech, transcription, summary triplets Derived from the Gigaword dataset Rush+2015 Overview The dataset is divided into five splits: train/core/dev/eval/duc2003. (See below table) We added a new evaluation split "test" for in-domain evaluation. The train split is here: MegaSSum(train). orig. data split #samples #speakers total dur.… See the full description on the dataset page: https://huggingface.co/datasets/komats/mega-ssum.

sourceHugging Facecc-by-4.0updated 2y agoView on Hugging Face
3likes314downloads
Dataset Card

Mega-SSum

  • —A large-scale English sentence-wise speech summarization (Sen-SSum) dataset
  • —Consists of 3.8M+ synthesized speech, transcription, summary triplets
  • —Derived from the Gigaword dataset Rush+2015

Overview

  • —The dataset is divided into five splits: train/core/dev/eval/duc2003. (See below table)
  • —We added a new evaluation split "test" for in-domain evaluation.
  • —The train split is here: MegaSSum(train).
orig. datasplit#samples#speakerstotal dur. (hrs)ave. dur. (sec)CR* (%)
Gigawordtrain3,800,0002,55911,678.211.126.2
Gigawordcore50,0002,559154.611.125.8
Gigawordvalid1,000963.010.725.1
Gigawordtest4,0008012.511.224.1
DUC2003duc2003624802.112.227.5

CR (compression rate, %) = #words in summary / #words in transcription 100. Lower is shorter summary.

Notes

  • —The core set is identical to the first 50k samples of the train split.
  • —You may train your model and report the results only with the core set because the train split is very large.
  • —Using the entire train split is generally not recommended unless there are special reasons (e.g., to investigate the upper bound).
  • —The duc2003 split has four reference summaries for each speech. You can report the best score from 4 scores.
  • —Spoken sentences were generated using VITS Kim+2021 trained with LibriTTS-R Koizumi+2023.
  • —More details and some experiments on this dataset can be found here.

Citation

  @inproceedings{matsuura24_interspeech,
    title     = {{Sentence-wise Speech Summarization}: Task, Datasets, and End-to-End Modeling with LM Knowledge Distillation},
    author    = {Kohei Matsuura and Takanori Ashihara and Takafumi Moriya and Masato Mimura and Takatomo Kano and Atsunori Ogawa and Marc Delcroix},
    year      = {2024},
    booktitle = {Interspeech 2024},
    pages     = {1945--1949},
  }
  @article{Rush_2015,
     title={A Neural Attention Model for Abstractive Sentence Summarization},
     journal={Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing},
     author={Rush, Alexander M. and Chopra, Sumit and Weston, Jason},
     year={2015}
  }
  @InProceedings{pmlr-v139-kim21f,
    title = 	 {Conditional Variational Autoencoder with Adversarial Learning for End-to-End Text-to-Speech},
    author =       {Kim, Jaehyeon and Kong, Jungil and Son, Juhee},
    booktitle = 	 {Proceedings of the 38th International Conference on Machine Learning},
    pages = 	 {5530--5540},
    year = 	 {2021},
  }
  @inproceedings{koizumi23_interspeech,
    author={Yuma Koizumi and Heiga Zen and Shigeki Karita and Yifan Ding and Kohei Yatabe and Nobuyuki Morioka and Michiel Bacchiani and Yu Zhang and Wei Han and Ankur Bapna},
    title={{LibriTTS-R}: A Restored Multi-Speaker Text-to-Speech Corpus},
    year=2023,
    booktitle={Proc. INTERSPEECH 2023},
    pages={5496--5500},
  }