llm-lab/SpeechBrown
Models | Springer Link | arXiv Link | Proposed Dataset | ACM Digital Library | Website Dataset Summary Speech Brown is a comprehensive, synthetic, and diverse paired speech-text dataset in 15 categories, covering a wide range of topics from fiction to religion. This dataset consists of over 55,000 sentence-level samples. To train the CLASP model, we created this dataset based on the Brown Corpus. The synthetic speech was generated using the NVIDIA Tacotron 2 text-to-speech… See the full description on the dataset page: https://huggingface.co/datasets/llm-lab/SpeechBrown.
Update README.md
Upload dataset_part5.zip
Delete dataset_part5.zip
Upload dataset_part1.zip
Delete dataset_part1.zip
Upload 2 files
Delete localized_metadata.json
Delete global_metadata.json
Update README.md
Upload 3 files
Upload 3 files
Upload dataset_part7.zip
Upload dataset_part6.zip
Upload dataset_part5.zip
Upload dataset_part4.zip
Upload dataset_part3.zip
Update README.md
Update README.md
Upload dataset_part2.zip
Delete datasets
Upload dataset_part2.zip
Update README.md
Update README.md
Update README.md
Upload dataset_part1.zip
Update README.md
Update README.md
Update README.md
initial commit
