datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
VieNeu-TTS-140h
pnnbao-ump/VieNeu-TTS-140h
Mô tả Dataset
A high-quality Vietnamese Text-to-Speech (TTS) dataset containing 74,858 audio samples with phonemized transcripts. This benchmark dataset is designed for fine-tuning modern TTS models with maximum synthesis quality. The text corpus is completely phonemized using standard international phonetic alphabet (IPA) representations suitable for neural acoustic modeling.
Quick Facts
Language: Vietnamese 🇻🇳
Tasks:… See the full description on the dataset page: https://huggingface.co/datasets/pnnbao-ump/VieNeu-TTS-140h.VietSpeech
VietSpeech: Vietnamese social voice dataset
Introdution
This dataset includes over 1,100 hours of speech data. The voice samples were collected from a variety of social resources, ensuring a diverse representation of accents (north, central, south), dialects, and speaking styles. This diversity makes the dataset particularly valuable for training and evaluating ASR models, as it enhances their ability to accurately recognize and transcribe speech across different… See the full description on the dataset page: https://huggingface.co/datasets/NhutP/VietSpeech.viet_bud500
Bud500: A Comprehensive Vietnamese ASR Dataset
Introducing Bud500, a diverse Vietnamese speech corpus designed to support ASR research community. With aprroximately 500 hours of audio, it covers a broad spectrum of topics including podcast, travel, book, food, and so on, while spanning accents from Vietnam's North, South, and Central regions. Derived from free public audio resources, this publicly accessible dataset is designed to significantly enhance the work of developers and… See the full description on the dataset page: https://huggingface.co/datasets/linhtran92/viet_bud500.VietMed
VietMed: A Dataset and Benchmark for Automatic Speech Recognition of Vietnamese in the Medical Domain (LREC-COLING 2024, Oral)
Description:
We introduced a Vietnamese speech recognition dataset in the medical domain comprising 16h of labeled medical speech, 1000h of unlabeled medical speech and 1200h of unlabeled general-domain speech.
To our best knowledge, VietMed is by far the world’s largest public medical speech recognition dataset in 7 aspects:
total duration… See the full description on the dataset page: https://huggingface.co/datasets/leduckhai/VietMed.vieneu-tts-140h-dataset
pnnbao-ump/VieNeu-TTS-140h
Mô tả Dataset
Dataset tiếng Việt chất lượng cao cho Text-to-Speech (TTS) với 74,858 mẫu audio và transcript được phonemize.
Mục tiêu của mình là tạo bộ dataset chuẩn mực để finetune các model TTS hiện nay với chất lượng cao nhất. Mình thu thập audio chất lượng cao từ youtube, làm sạch nền, loại bỏ noise, dùng whisper-large-v3 để tạo transcription, sau đó cho Agent sửa lỗi chính tả và feedback lại cho con người. Bộ dữ liệu cũng được phonemize hóa… See the full description on the dataset page: https://huggingface.co/datasets/LanguaMan/vieneu-tts-140h-dataset.vietspeech500G of vietnamese speech corpus from Youtube
data-voice-vietnamese-restaurant-quan-oc
Vietnamese Restaurant Order Speech
This dataset contains Vietnamese spoken restaurant orders paired with text transcripts. Each utterance typically includes a table number, item quantities, dishes, drinks, and add-ons.
Dataset Structure
Files are split into subdirectories by filename-derived speaker_code to satisfy Hugging Face repository file-count limits:
metadata.csv: one row per audio sample.
audio/{speaker_code}/*.wav: mono WAV audio files.… See the full description on the dataset page: https://huggingface.co/datasets/EmilyNguyen235/data-voice-vietnamese-restaurant-quan-oc.VietMed-NER
Medical Spoken Named Entity Recognition (NAACL 2025)
Description:
Spoken Named Entity Recognition (NER) aims to extract named entities from speech and categorise them into types like person, location, organization, etc. In this work, we present VietMed-NER - the first spoken NER dataset in the medical domain. To our knowledge, our Vietnamese real-world dataset is the largest spoken NER dataset in the world regarding the number of entity types, featuring 18 distinct types.… See the full description on the dataset page: https://huggingface.co/datasets/leduckhai/VietMed-NER.Bahnar_Vietnamese
Bahnar Speech Translation Dataset
This dataset contains Bahnar speech audio aligned with Bahnar, Vietnamese, and English text. It was created from internet data sources and automatically aligned using the pipeline available at Bahnar-Vietnamese-S2TT.
The main purpose of this dataset is to support research on low-resource speech-to-text translation (S2TT), especially direct translation from Bahnar speech to Vietnamese text.
Data Statistics
Train: 113,830… See the full description on the dataset page: https://huggingface.co/datasets/cuong06/Bahnar_Vietnamese.VietMed_unlabeled
unofficial mirror of VietMed (Vietnamese speech data in medical domain) unlabeled set
official announcement: https://arxiv.org/abs/2404.05659
official download: https://huggingface.co/datasets/leduckhai/VietMed
this repo contains the unlabeled set: 966h - 230k samples
i also gather the metadata: see info.csv
my extraction code: https://github.com/phineas-pta/fine-tune-whisper-vi/blob/main/misc/vietmed-unlabeled.py
need to do: check misspelling, restore foreign words phonetised to… See the full description on the dataset page: https://huggingface.co/datasets/doof-ferb/VietMed_unlabeled.voice-Viet-Nam
🇻🇳 Voice Dataset Viet Nam
📝 Giới thiệu (Dataset Description)
Đây là kho dữ liệu âm thanh tiếng Việt mã nguồn mở (Voice Dataset Viet Nam). Bộ dữ liệu này được thu thập và cấu trúc theo chuẩn AudioFolder của Hugging Face, phục vụ cho các bài toán:
Automatic Speech Recognition (ASR): Nhận dạng giọng nói.
Text-to-Speech (TTS): Tổng hợp tiếng nói.
Voice Cloning: Huấn luyện mô hình nhân bản giọng nói.
📊 Thống kê dữ liệu (Statistics)
Thuộc tính… See the full description on the dataset page: https://huggingface.co/datasets/dinhlam2210/voice-Viet-Nam.Vietnamese-asr-leaderboard
📊 Vietnamese Open ASR Evaluation Dataset Storage
Kho lưu trữ dữ liệu nhãn bảo mật (Ground Truth) phục vụ cho hệ thống Vietnamese Open ASR Leaderboard. Toàn bộ dữ liệu được tổng hợp từ 9 bộ dữ liệu tiếng Việt công khai lớn nhất hiện nay, sau đó trải qua quy trình chuẩn hóa văn bản nghiêm ngặt để làm thước đo chuẩn mực đánh giá hiệu năng các mô hình nhận dạng giọng nói (ASR).
[!TIP]
🚀 NỘP BÀI ĐÁNH GIÁ TẠI ĐÂY:
📈 1. Bảng Thống Kê Chi Tiết Hệ Dữ Liệu… See the full description on the dataset page: https://huggingface.co/datasets/VietAudio-team/Vietnamese-asr-leaderboard.VietMed_labeled
unofficial mirror of VietMed (Vietnamese speech data in medical domain) labeled set
official announcement: https://arxiv.org/abs/2404.05659
official download: https://huggingface.co/datasets/leduckhai/VietMed
this repo contains the labeled set: 9.2k samples
i also gather the metadata: see info.csv
my extraction code: https://github.com/phineas-pta/fine-tune-whisper-vi/blob/main/misc/vietmed-labeled.py
need to do: check misspelling, restore foreign words phonetised to vietnamese… See the full description on the dataset page: https://huggingface.co/datasets/doof-ferb/VietMed_labeled.vietnamese_asr
Vietnamese ASR (VIVOS Corpus)
Dataset Overview
The VIVOS corpus is a free Vietnamese speech dataset consisting of 15 hours of recorded speech. This corpus was meticulously prepared for Automatic Speech Recognition (ASR) tasks and is explicitly divided into training and testing sets.
Speech was recorded in a quiet environment using high-quality microphones, where native speakers were asked to read pre-prepared texts line by line.
Note: This repository… See the full description on the dataset page: https://huggingface.co/datasets/duymanh1606/vietnamese_asr.uts2025_vietipa
Vietnamese IPA Dataset
A comprehensive Vietnamese IPA (International Phonetic Alphabet) dataset with word pronunciations and MP3 audio files for text-to-speech and pronunciation learning applications.
Dataset Description
Dataset Summary
This dataset contains 50 common Vietnamese words with their IPA (International Phonetic Alphabet) transcriptions and corresponding audio files. It's designed for:
Text-to-speech systems development
Vietnamese pronunciation… See the full description on the dataset page: https://huggingface.co/datasets/undertheseanlp/uts2025_vietipa.vietnamese-speech-recognition
Vietnamese Speech Dataset
Dataset comprises 10+ hours of telephone dialogues in Vietnamese, collected from 20 native speakers across various topics and domains. It is designed for research in speech recognition, focusing on various recognition models, primarily aimed at meeting the requirements for automatic speech recognition (ASR) systems.
By utilizing this dataset, researchers and developers can advance their understanding and capabilities in transcribing audio, and natural… See the full description on the dataset page: https://huggingface.co/datasets/UniDataPro/vietnamese-speech-recognition.Speech-MASSIVE_vie
Vietnamse subset of the Speech-MASSIVE dataset
extracted from:
https://huggingface.co/datasets/FBK-MT/Speech-MASSIVE
https://huggingface.co/datasets/FBK-MT/Speech-MASSIVE-test
my extraction code: https://github.com/phineas-pta/fine-tune-whisper-vi/blob/main/misc/load-speechmassive.py
BibleMMS_vie
Vietnamse subset of the BibleMMS dataset
extracted from: https://huggingface.co/datasets/Flux9665/BibleMMS
my extraction code: https://github.com/phineas-pta/fine-tune-whisper-vi/blob/main/misc/load-biblemms.py
VietMDD
unofficial mirror of VietMMD (Mispronunciation Detection and Diagnosis)
official announcement: https://github.com/VietMDDDataset/VietMDD
official download: https://drive.google.com/drive/folders/1TjTluTxEB99QhGFTYFWb-vEdWXM-lyKJ?usp=sharing
DOI: 10.21437/Interspeech.2023-364
5h, 4.2k samples
pre-process: see my code: https://github.com/phineas-pta/fine-tune-whisper-vi/blob/main/misc/viet-mdd.py
custom split: orphan: speech without any transcription unlike in train/validation/test… See the full description on the dataset page: https://huggingface.co/datasets/doof-ferb/VietMDD.Vietnamese_Medical_Consultationgigaspeech2_vie
Vietnamse subset of the Gigaspeech2 dataset
extracted from: https://huggingface.co/datasets/speechcolab/gigaspeech2
VietENT-Text
VietENT-Text — Vietnamese ENT clinical sentence corpus
20,908 unique Vietnamese ear-nose-throat clinical sentences and a 235-term ENT
catalogue, written to train and evaluate speech recognition on ENT consultations.
from datasets import load_dataset
ds = load_dataset("diepduclai/VietENT-Text", split="train")
ds[0]["text"]
Why this exists
Vietnamese ASR handles general speech well and medical terminology badly, and the
failures are the dangerous kind. Measured on… See the full description on the dataset page: https://huggingface.co/datasets/diepduclai/VietENT-Text.Vietnamese_ASR_TestingDataVietENT-Speech
VietENT-Speech — synthetic Vietnamese ENT clinical speech
41,816 synthetic 16 kHz utterances of Vietnamese ear-nose-throat clinical speech:
20,908 unique sentences, each rendered once clean and once through a measured
recording-channel model, across 106 cloned voices.
from datasets import load_dataset
ds = load_dataset("diepduclai/VietENT-Speech", split="train", streaming=True) # ~7.5 GB, stream it
next(iter(ds))["audio"]
License is CC BY-NC-SA 4.0, and that was… See the full description on the dataset page: https://huggingface.co/datasets/diepduclai/VietENT-Speech.vietnamese-speech-recognition
Vietnamese Speech Dataset - 10+ hours
Dataset comprises 10+ hours of telephone dialogues in Vietnamese, collected from 20 native speakers across various topics and domains. It is designed for research in speech recognition, focusing on various recognition models, primarily aimed at meeting the requirements for automatic speech recognition (ASR) systems. - Get the data
Dataset characteristics:
Characteristic
Data
Description
Audio of telephone dialogues in… See the full description on the dataset page: https://huggingface.co/datasets/ud-nlp/vietnamese-speech-recognition.Vietnamese-Speech-Dataset
🎧 Vietnamese Speech Dataset
The Vietnamese Speech Dataset is a large-scale speech audio dataset designed to support advanced AI systems with diverse and high-quality audio data. It includes 179 hours of recorded speech data across 710 files, delivered in MP3 and WAV formats, with a total size of 280 MB. This well-structured audio dataset provides balanced and representative voice data, featuring 52% female and 48% male speakers, with age coverage from 18 to 50+ years. The dataset… See the full description on the dataset page: https://huggingface.co/datasets/Speech-data/Vietnamese-Speech-Dataset.Vietnamese_ASR_TestingData_Old
About
This dataset is for only ASR testing in Vietnamese.
We collect data from various sources.
This is the first version of the dataset.
VieNeu-TTS-1000h
pnnbao-ump/VieNeu-TTS-1000h
Mô tả Dataset
Dataset tiếng Việt chất lượng cao cho Text-to-Speech (TTS) với 443,641 mẫu audio và transcript được phonemize.
Bộ dữ liệu 1000 giờ này được thiết kế để train từ đầu (from scratch) hoặc finetune các mô hình TTS hoặc ASR với chất lượng cao nhất có thể đạt được hiện nay.
Tuy nhiên, dataset không được mở tải xuống công khai vì lý do bản quyền, kiểm soát chất lượng và yêu cầu về mục đích nghiên cứu.
This dataset is not publicly… See the full description on the dataset page: https://huggingface.co/datasets/pnnbao-ump/VieNeu-TTS-1000h.vietmed-reviewed
VietMed Reviewed
VietMed Reviewed is a reviewed Vietnamese medical speech dataset for automatic speech recognition.
This dataset is built from reviewed pseudo-labeled samples of the VietMed unlabeled subset. Each sample contains a WAV audio segment and a final reviewed transcript.
The dataset follows a simplified schema inspired by leduckhai/VietMed, with one split named reviewed.
Dataset Details
Task: Automatic Speech Recognition
Language: Vietnamese
Domain:… See the full description on the dataset page: https://huggingface.co/datasets/phuvo05/vietmed-reviewed.knihi-be-jeva_vieznaviec_pa_sto_idzies_vouca_all
AudioSet Pipeline Output
Мова / Language: Беларуская (Belarusian)
Аўдыё нарэзана з арыгінальнага запісу ў зыходнай частаце дыскрэтызацыі (native), мона, фрагменты да 30 секунд.
Частка калекцыі Belarusian Audiobooks (native).
Радкоў у датасеце
1,023
Працягласць
3 гадз 26 хв
Частата дыскрэтызацыі
44100 Hz
Каналы
мона
Даўжыня фрагмента
да 30 с
Структура
Кожны радок змяшчае:
audio — аўдыёфрагмент (native SR, мона, ≤30 с)
text — транскрыпцыя… See the full description on the dataset page: https://huggingface.co/datasets/fosters/knihi-be-jeva_vieznaviec_pa_sto_idzies_vouca_all.
