CoolFace
Datasetpublic

speechbrain/LoquaciousSet

LargeScaleASR: 25,000 hours of transcribed and heterogeneous English speech recognition data for research and commercial use. The full details are available in the paper. Made of 6 subsets: large contains 25,000 hours of read / spontaneous and clean / noisy transcribed speech. medium contains 2,500 hours of read / spontaneous and clean / noisy transcribed speech. small contains 250 hours of read / spontaneous and clean / noisy transcribed speech. clean contains 13,000 hours of… See the full description on the dataset page: https://huggingface.co/datasets/speechbrain/LoquaciousSet.

sourceHugging Facecc-by-3.0updated 7mo agoView on Hugging Face
63likes6.8kdownloads
dirdev/
file3gram-pruned.arpa.bin720.4 MBdownload
file3gram-pruned.arpa.gz330.3 MBdownload
file4gram-pruned.arpa.bin1.12 GBdownload
file4gram-pruned.arpa.gz537.1 MBdownload
file4gram-unpruned.arpa.bin4.68 GBdownload
file4gram-unpruned.arpa.gz2.34 GBdownload

speechbrain/LoquaciousSet · main · files are served by the source, never re-hosted here