CoolFace
Datasetpublic

thanhnew2001/VietSuperSpeech

VietSuperSpeech Vietnamese Speech Recognition Dataset Dataset Information Total samples: 32,267 Train samples: 29,041 Dev samples: 3,226 Total duration: 103.18 hours Sample rate: 16000 Hz Average segment length: ~12 seconds Source Datasets asr_dataset_nguoivietdailynews asr_dataset_nguyenkhangofficial asr_dataset_trinhlieu Format The dataset follows Icefall format: train.json: Training samples dev.json: Development samples… See the full description on the dataset page: https://huggingface.co/datasets/thanhnew2001/VietSuperSpeech.

sourceHugging Faceupdated 7mo agoView on Hugging Face
6likes3.4kdownloads
Dataset Card

VietSuperSpeech

Vietnamese Speech Recognition Dataset

Dataset Information

  • Total samples: 32,267
  • Train samples: 29,041
  • Dev samples: 3,226
  • Total duration: 103.18 hours
  • Sample rate: 16000 Hz
  • Average segment length: ~12 seconds

Source Datasets

  • asrdatasetnguoivietdailynews
  • asrdatasetnguyenkhangofficial
  • asrdatasettrinhlieu

Format

The dataset follows Icefall format:

  • train.json: Training samples
  • dev.json: Development samples
  • manifest.json: Dataset metadata
  • audio/: Audio files organized by source dataset

Each sample contains:

  • audio: Relative path to audio file
  • text: Transcribed text
  • duration: Duration in seconds
  • source: Source video/file name

Transcription Model

Transcribed using Zipformer-30M-RNNT-6000h model.

License

MIT

thanhnew2001/VietSuperSpeech · CoolFace