thanhnew2001/VietSuperSpeech
VietSuperSpeech Vietnamese Speech Recognition Dataset Dataset Information Total samples: 32,267 Train samples: 29,041 Dev samples: 3,226 Total duration: 103.18 hours Sample rate: 16000 Hz Average segment length: ~12 seconds Source Datasets asr_dataset_nguoivietdailynews asr_dataset_nguyenkhangofficial asr_dataset_trinhlieu Format The dataset follows Icefall format: train.json: Training samples dev.json: Development samples… See the full description on the dataset page: https://huggingface.co/datasets/thanhnew2001/VietSuperSpeech.
63.4k
VietSuperSpeech
Vietnamese Speech Recognition Dataset
Dataset Information
- Total samples: 32,267
- Train samples: 29,041
- Dev samples: 3,226
- Total duration: 103.18 hours
- Sample rate: 16000 Hz
- Average segment length: ~12 seconds
Source Datasets
- asrdatasetnguoivietdailynews
- asrdatasetnguyenkhangofficial
- asrdatasettrinhlieu
Format
The dataset follows Icefall format:
train.json: Training samplesdev.json: Development samplesmanifest.json: Dataset metadataaudio/: Audio files organized by source dataset
Each sample contains:
audio: Relative path to audio filetext: Transcribed textduration: Duration in secondssource: Source video/file name
Transcription Model
Transcribed using Zipformer-30M-RNNT-6000h model.
License
MIT
