datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
VietSuperSpeech
VietSuperSpeech
Vietnamese Speech Recognition Dataset
Dataset Information
Total samples: 32,267
Train samples: 29,041
Dev samples: 3,226
Total duration: 103.18 hours
Sample rate: 16000 Hz
Average segment length: ~12 seconds
Source Datasets
asr_dataset_nguoivietdailynews
asr_dataset_nguyenkhangofficial
asr_dataset_trinhlieu
Format
The dataset follows Icefall format:
train.json: Training samples
dev.json: Development samples
manifest.json:… See the full description on the dataset page: https://huggingface.co/datasets/thanhnew2001/VietSuperSpeech.viet_vlsp
Dataset Card for "viet_vlsp"
More Information needed
dolly-audio-1000h-vietnamese
Dolly-Audio: Vietnamese Multi-Speaker High-Quality Speech Corpus
Dataset Summary
Dolly-Audio is a large-scale, high-quality Vietnamese speech corpus created by the Dolly AI Team.
Inspired by Dolly, the world’s first cloned mammal, the project aims to advance research in Vietnamese speech synthesis, speech recognition, and voice modeling.
This release provides nearly 1,000 hours of professionally cleaned audio, featuring 152 speakers across different Vietnamese regions and… See the full description on the dataset page: https://huggingface.co/datasets/dolly-vn/dolly-audio-1000h-vietnamese.VieNeu-TTS-140h
pnnbao-ump/VieNeu-TTS-140h
Mô tả Dataset
A high-quality Vietnamese Text-to-Speech (TTS) dataset containing 74,858 audio samples with phonemized transcripts. This benchmark dataset is designed for fine-tuning modern TTS models with maximum synthesis quality. The text corpus is completely phonemized using standard international phonetic alphabet (IPA) representations suitable for neural acoustic modeling.
Quick Facts
Language: Vietnamese 🇻🇳
Tasks:… See the full description on the dataset page: https://huggingface.co/datasets/pnnbao-ump/VieNeu-TTS-140h.VietSpeech
VietSpeech: Vietnamese social voice dataset
Introdution
This dataset includes over 1,100 hours of speech data. The voice samples were collected from a variety of social resources, ensuring a diverse representation of accents (north, central, south), dialects, and speaking styles. This diversity makes the dataset particularly valuable for training and evaluating ASR models, as it enhances their ability to accurately recognize and transcribe speech across different… See the full description on the dataset page: https://huggingface.co/datasets/NhutP/VietSpeech.Vietnam-Celeb
unofficial mirror of Vietnam-Celeb dataset
official announcement:
https://www.isca-archive.org/interspeech_2023/pham23b_interspeech.html
https://github.com/Vietnam-Celeb/Vietnam-Celeb
https://huggingface.co/datasets/hustep-lab/Vietnam-Celeb
official download:
Part 0: https://drive.google.com/file/d/1pMuT3DFzSwib7SVcRS8VkDwPuLTsemSG/view?usp=share_link
Part 1: https://drive.google.com/file/d/1xayHt2HRqE1aJ4HvtUT40_9XlgvfDfRY/view?usp=share_linkPart 2:… See the full description on the dataset page: https://huggingface.co/datasets/doof-ferb/Vietnam-Celeb.viet_bud500
Bud500: A Comprehensive Vietnamese ASR Dataset
Introducing Bud500, a diverse Vietnamese speech corpus designed to support ASR research community. With aprroximately 500 hours of audio, it covers a broad spectrum of topics including podcast, travel, book, food, and so on, while spanning accents from Vietnam's North, South, and Central regions. Derived from free public audio resources, this publicly accessible dataset is designed to significantly enhance the work of developers and… See the full description on the dataset page: https://huggingface.co/datasets/linhtran92/viet_bud500.youtube-center-vietnamese-asrvlsp-vie-speech2text
Dataset Card for "vlsp-vie-speech2text"
More Information needed
vietnamese-music-dataset
Vietnamese Music Dataset
A collection of 4,820 Vietnamese music tracks with matching cover thumbnails and per-track metadata collected from YouTube, packaged as an audiofolder dataset.
Repository structure
Path
Contents
Count
audio/
MP3 audio files, named by YouTube video ID
4,820
images/
PNG cover thumbnails, same IDs as audio/
4,820
data/
Parquet metadata files, one per collection session
31
Metadata schema
Each Parquet file in… See the full description on the dataset page: https://huggingface.co/datasets/Toan-Minh-Duong-Son/vietnamese-music-dataset.VietCasualSpeechAVietMed
VietMed: A Dataset and Benchmark for Automatic Speech Recognition of Vietnamese in the Medical Domain (LREC-COLING 2024, Oral)
Description:
We introduced a Vietnamese speech recognition dataset in the medical domain comprising 16h of labeled medical speech, 1000h of unlabeled medical speech and 1200h of unlabeled general-domain speech.
To our best knowledge, VietMed is by far the world’s largest public medical speech recognition dataset in 7 aspects:
total duration… See the full description on the dataset page: https://huggingface.co/datasets/leduckhai/VietMed.dolly-audio-1000h-vietnamese
Dolly-Audio: Vietnamese Multi-Speaker High-Quality Speech Corpus
Dataset Summary
Dolly-Audio is a large-scale, high-quality Vietnamese speech corpus created by the Dolly AI Team.
Inspired by Dolly, the world’s first cloned mammal, the project aims to advance research in Vietnamese speech synthesis, speech recognition, and voice modeling.
This release provides nearly 1,000 hours of professionally cleaned audio, featuring 152 speakers across different Vietnamese regions and… See the full description on the dataset page: https://huggingface.co/datasets/fnooub/dolly-audio-1000h-vietnamese.viet_vlsp
Dataset Card for "viet_vlsp"
More Information needed
dolly-audio-1000h-vietnamese
Dolly-Audio: Vietnamese Multi-Speaker High-Quality Speech Corpus
Dataset Summary
Dolly-Audio is a large-scale, high-quality Vietnamese speech corpus created by the Dolly AI Team.
Inspired by Dolly, the world’s first cloned mammal, the project aims to advance research in Vietnamese speech synthesis, speech recognition, and voice modeling.
This release provides nearly 1,000 hours of professionally cleaned audio, featuring 152 speakers across different Vietnamese regions and… See the full description on the dataset page: https://huggingface.co/datasets/quanghd96/dolly-audio-1000h-vietnamese.vieneu-tts-140h-dataset
pnnbao-ump/VieNeu-TTS-140h
Mô tả Dataset
Dataset tiếng Việt chất lượng cao cho Text-to-Speech (TTS) với 74,858 mẫu audio và transcript được phonemize.
Mục tiêu của mình là tạo bộ dataset chuẩn mực để finetune các model TTS hiện nay với chất lượng cao nhất. Mình thu thập audio chất lượng cao từ youtube, làm sạch nền, loại bỏ noise, dùng whisper-large-v3 để tạo transcription, sau đó cho Agent sửa lỗi chính tả và feedback lại cho con người. Bộ dữ liệu cũng được phonemize hóa… See the full description on the dataset page: https://huggingface.co/datasets/LanguaMan/vieneu-tts-140h-dataset.vietspeech500G of vietnamese speech corpus from Youtube
data-voice-vietnamese-restaurant-quan-oc
Vietnamese Restaurant Order Speech
This dataset contains Vietnamese spoken restaurant orders paired with text transcripts. Each utterance typically includes a table number, item quantities, dishes, drinks, and add-ons.
Dataset Structure
Files are split into subdirectories by filename-derived speaker_code to satisfy Hugging Face repository file-count limits:
metadata.csv: one row per audio sample.
audio/{speaker_code}/*.wav: mono WAV audio files.… See the full description on the dataset page: https://huggingface.co/datasets/EmilyNguyen235/data-voice-vietnamese-restaurant-quan-oc.VietBibleVox-aligned
VietBibleVox Dataset
The VietBibleVox Dataset is based on the data extracted from open.bible specifically for the Vietnamese language. As the original data is provided under the cc-by-sa-4.0 license, this derived dataset is also licensed under cc-by-sa-4.0.
The dataset comprises 29,185 pairs of (verse, audio clip), with each verse from the Bible read in Vietnamese by a male voice.
The verses are the original texts and may not be directly usable for training text-to-speech models.… See the full description on the dataset page: https://huggingface.co/datasets/ntt123/VietBibleVox-aligned.VietMed-NER
Medical Spoken Named Entity Recognition (NAACL 2025)
Description:
Spoken Named Entity Recognition (NER) aims to extract named entities from speech and categorise them into types like person, location, organization, etc. In this work, we present VietMed-NER - the first spoken NER dataset in the medical domain. To our knowledge, our Vietnamese real-world dataset is the largest spoken NER dataset in the world regarding the number of entity types, featuring 18 distinct types.… See the full description on the dataset page: https://huggingface.co/datasets/leduckhai/VietMed-NER.vie-speech-corpus-completeBahnar_Vietnamese
Bahnar Speech Translation Dataset
This dataset contains Bahnar speech audio aligned with Bahnar, Vietnamese, and English text. It was created from internet data sources and automatically aligned using the pipeline available at Bahnar-Vietnamese-S2TT.
The main purpose of this dataset is to support research on low-resource speech-to-text translation (S2TT), especially direct translation from Bahnar speech to Vietnamese text.
Data Statistics
Train: 113,830… See the full description on the dataset page: https://huggingface.co/datasets/cuong06/Bahnar_Vietnamese.VietCasualSpeechBvie-speech-corpus-1dolly-audio-1000h-vietnamese
Dolly-Audio: Vietnamese Multi-Speaker High-Quality Speech Corpus
Dataset Summary
Dolly-Audio is a large-scale, high-quality Vietnamese speech corpus created by the Dolly AI Team.
Inspired by Dolly, the world’s first cloned mammal, the project aims to advance research in Vietnamese speech synthesis, speech recognition, and voice modeling.
This release provides nearly 1,000 hours of professionally cleaned audio, featuring 152 speakers across different Vietnamese regions and… See the full description on the dataset page: https://huggingface.co/datasets/AnhTuan89/dolly-audio-1000h-vietnamese.vlsp-vie-speech2text1
Dataset Card for "vlsp-vie-speech2text1"
More Information needed
dolly-audio-1000h-vietnamese
Dolly-Audio: Vietnamese Multi-Speaker High-Quality Speech Corpus
Dataset Summary
Dolly-Audio is a large-scale, high-quality Vietnamese speech corpus created by the Dolly AI Team.
Inspired by Dolly, the world’s first cloned mammal, the project aims to advance research in Vietnamese speech synthesis, speech recognition, and voice modeling.
This release provides nearly 1,000 hours of professionally cleaned audio, featuring 152 speakers across different Vietnamese regions and… See the full description on the dataset page: https://huggingface.co/datasets/ThanhNV1999/dolly-audio-1000h-vietnamese.f5tts-vietnamese-datasetdataset-vietvoice_v2VietMed_unlabeled
unofficial mirror of VietMed (Vietnamese speech data in medical domain) unlabeled set
official announcement: https://arxiv.org/abs/2404.05659
official download: https://huggingface.co/datasets/leduckhai/VietMed
this repo contains the unlabeled set: 966h - 230k samples
i also gather the metadata: see info.csv
my extraction code: https://github.com/phineas-pta/fine-tune-whisper-vi/blob/main/misc/vietmed-unlabeled.py
need to do: check misspelling, restore foreign words phonetised to… See the full description on the dataset page: https://huggingface.co/datasets/doof-ferb/VietMed_unlabeled.
