MLCommons/peoples_speech_v1.0
Dataset Card for People's Speech Dataset Summary The People's Speech Dataset is among the world's largest English speech recognition corpus today that is licensed for academic and commercial usage under CC-BY-SA and CC-BY 4.0. It includes 30,000+ hours of transcribed speech in English languages with a diverse set of speakers. This open dataset is large enough to train speech-to-text systems and crucially is available with a permissive license.… See the full description on the dataset page: https://huggingface.co/datasets/MLCommons/peoples_speech_v1.0.
Remove deprecated tasks (#3)
fix meta tags
add microset
fix n_files.txt
Update README.md
update metadata
add missing data
Fill out datasheet some more.
Create README.md
add splits support
move n_files
add n_files for dev and test splits
add index.json
add dev and test sets
specify paths to all archives
add txt files with number of archives
hotfix of datasets viewer
add json files for first tar of each split
move training data into /train
add some comments and todos
make urls relative
store local paths in non-streaming mode
remove comment
add loading script
add data
initial commit
