CoolFace
Datasetpublic

llm-lab/SpeechBrown

Models | Springer Link | arXiv Link | Proposed Dataset | ACM Digital Library | Website Dataset Summary Speech Brown is a comprehensive, synthetic, and diverse paired speech-text dataset in 15 categories, covering a wide range of topics from fiction to religion. This dataset consists of over 55,000 sentence-level samples. To train the CLASP model, we created this dataset based on the Brown Corpus. The synthetic speech was generated using the NVIDIA Tacotron 2 text-to-speech… See the full description on the dataset page: https://huggingface.co/datasets/llm-lab/SpeechBrown.

sourceHugging Facemitupdated 1y agoView on Hugging Face
2likes82downloads
filedataset_part1.zip3.07 GBdownload
filedataset_part10.zip1.98 GBdownload
filedataset_part2.zip2.82 GBdownload
filedataset_part3.zip2.80 GBdownload
filedataset_part4.zip3.06 GBdownload
filedataset_part5.zip3.05 GBdownload
filedataset_part6.zip3.34 GBdownload
filedataset_part7.zip2.70 GBdownload
filedataset_part8.zip1.84 GBdownload
filedataset_part9.zip1.88 GBdownload
filesamples-brown.zip33.0 MBdownload

llm-lab/SpeechBrown · main · files are served by the source, never re-hosted here