speechbrain/LoquaciousSet
LargeScaleASR: 25,000 hours of transcribed and heterogeneous English speech recognition data for research and commercial use. The full details are available in the paper. Made of 6 subsets: large contains 25,000 hours of read / spontaneous and clean / noisy transcribed speech. medium contains 2,500 hours of read / spontaneous and clean / noisy transcribed speech. small contains 250 hours of read / spontaneous and clean / noisy transcribed speech. clean contains 13,000 hours of… See the full description on the dataset page: https://huggingface.co/datasets/speechbrain/LoquaciousSet.
636.8k
1version https://git-lfs.github.com/spec/v12oid sha256:e70d87860437f3a77209e932eeab3dd9dd73b4150b8cba4d3379044d29be4dad3size 16906194 