datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Vietnamese-Traditional-Musicdolly-audio-1000h-vietnamese
Dolly-Audio: Vietnamese Multi-Speaker High-Quality Speech Corpus
Dataset Summary
Dolly-Audio is a large-scale, high-quality Vietnamese speech corpus created by the Dolly AI Team.
Inspired by Dolly, the world’s first cloned mammal, the project aims to advance research in Vietnamese speech synthesis, speech recognition, and voice modeling.
This release provides nearly 1,000 hours of professionally cleaned audio, featuring 152 speakers across different Vietnamese regions and… See the full description on the dataset page: https://huggingface.co/datasets/dolly-vn/dolly-audio-1000h-vietnamese.youtube-center-vietnamese-asrvietnamese-music-dataset
Vietnamese Music Dataset
A collection of 4,820 Vietnamese music tracks with matching cover thumbnails and per-track metadata collected from YouTube, packaged as an audiofolder dataset.
Repository structure
Path
Contents
Count
audio/
MP3 audio files, named by YouTube video ID
4,820
images/
PNG cover thumbnails, same IDs as audio/
4,820
data/
Parquet metadata files, one per collection session
31
Metadata schema
Each Parquet file in… See the full description on the dataset page: https://huggingface.co/datasets/Toan-Minh-Duong-Son/vietnamese-music-dataset.dolly-audio-1000h-vietnamese
Dolly-Audio: Vietnamese Multi-Speaker High-Quality Speech Corpus
Dataset Summary
Dolly-Audio is a large-scale, high-quality Vietnamese speech corpus created by the Dolly AI Team.
Inspired by Dolly, the world’s first cloned mammal, the project aims to advance research in Vietnamese speech synthesis, speech recognition, and voice modeling.
This release provides nearly 1,000 hours of professionally cleaned audio, featuring 152 speakers across different Vietnamese regions and… See the full description on the dataset page: https://huggingface.co/datasets/fnooub/dolly-audio-1000h-vietnamese.dolly-audio-1000h-vietnamese
Dolly-Audio: Vietnamese Multi-Speaker High-Quality Speech Corpus
Dataset Summary
Dolly-Audio is a large-scale, high-quality Vietnamese speech corpus created by the Dolly AI Team.
Inspired by Dolly, the world’s first cloned mammal, the project aims to advance research in Vietnamese speech synthesis, speech recognition, and voice modeling.
This release provides nearly 1,000 hours of professionally cleaned audio, featuring 152 speakers across different Vietnamese regions and… See the full description on the dataset page: https://huggingface.co/datasets/quanghd96/dolly-audio-1000h-vietnamese.data-voice-vietnamese-restaurant-quan-oc
Vietnamese Restaurant Order Speech
This dataset contains Vietnamese spoken restaurant orders paired with text transcripts. Each utterance typically includes a table number, item quantities, dishes, drinks, and add-ons.
Dataset Structure
Files are split into subdirectories by filename-derived speaker_code to satisfy Hugging Face repository file-count limits:
metadata.csv: one row per audio sample.
audio/{speaker_code}/*.wav: mono WAV audio files.… See the full description on the dataset page: https://huggingface.co/datasets/EmilyNguyen235/data-voice-vietnamese-restaurant-quan-oc.dolly-audio-1000h-vietnamese
Dolly-Audio: Vietnamese Multi-Speaker High-Quality Speech Corpus
Dataset Summary
Dolly-Audio is a large-scale, high-quality Vietnamese speech corpus created by the Dolly AI Team.
Inspired by Dolly, the world’s first cloned mammal, the project aims to advance research in Vietnamese speech synthesis, speech recognition, and voice modeling.
This release provides nearly 1,000 hours of professionally cleaned audio, featuring 152 speakers across different Vietnamese regions and… See the full description on the dataset page: https://huggingface.co/datasets/AnhTuan89/dolly-audio-1000h-vietnamese.Bahnar_Vietnamese
Bahnar Speech Translation Dataset
This dataset contains Bahnar speech audio aligned with Bahnar, Vietnamese, and English text. It was created from internet data sources and automatically aligned using the pipeline available at Bahnar-Vietnamese-S2TT.
The main purpose of this dataset is to support research on low-resource speech-to-text translation (S2TT), especially direct translation from Bahnar speech to Vietnamese text.
Data Statistics
Train: 113,830… See the full description on the dataset page: https://huggingface.co/datasets/cuong06/Bahnar_Vietnamese.dolly-audio-1000h-vietnamese
Dolly-Audio: Vietnamese Multi-Speaker High-Quality Speech Corpus
Dataset Summary
Dolly-Audio is a large-scale, high-quality Vietnamese speech corpus created by the Dolly AI Team.
Inspired by Dolly, the world’s first cloned mammal, the project aims to advance research in Vietnamese speech synthesis, speech recognition, and voice modeling.
This release provides nearly 1,000 hours of professionally cleaned audio, featuring 152 speakers across different Vietnamese regions and… See the full description on the dataset page: https://huggingface.co/datasets/ThanhNV1999/dolly-audio-1000h-vietnamese.f5tts-vietnamese-datasetvietnamese-acoustic-boundary-verifier-data
Vietnamese Acoustic Boundary & Speaker Purity Dataset (Gemini 3.8 Flash Distilled)
This dataset contains 202 curated Vietnamese audio samples with fine-grained acoustic boundary annotations distilled from Google Gemini 3.8 Flash (thinkingLevel="LOW").
It is specifically designed to train and evaluate multimodal models (e.g., Gemma 4 E4B Audio) on acoustic quality control for speech synthesis and speaker diarization pipelines.
Dataset Structure
Each sample is… See the full description on the dataset page: https://huggingface.co/datasets/tungnguyenlam/vietnamese-acoustic-boundary-verifier-data.Vietnamese-asr-leaderboard
📊 Vietnamese Open ASR Evaluation Dataset Storage
Kho lưu trữ dữ liệu nhãn bảo mật (Ground Truth) phục vụ cho hệ thống Vietnamese Open ASR Leaderboard. Toàn bộ dữ liệu được tổng hợp từ 9 bộ dữ liệu tiếng Việt công khai lớn nhất hiện nay, sau đó trải qua quy trình chuẩn hóa văn bản nghiêm ngặt để làm thước đo chuẩn mực đánh giá hiệu năng các mô hình nhận dạng giọng nói (ASR).
[!TIP]
🚀 NỘP BÀI ĐÁNH GIÁ TẠI ĐÂY:
📈 1. Bảng Thống Kê Chi Tiết Hệ Dữ Liệu… See the full description on the dataset page: https://huggingface.co/datasets/VietAudio-team/Vietnamese-asr-leaderboard.vietnamese_asr
Vietnamese ASR (VIVOS Corpus)
Dataset Overview
The VIVOS corpus is a free Vietnamese speech dataset consisting of 15 hours of recorded speech. This corpus was meticulously prepared for Automatic Speech Recognition (ASR) tasks and is explicitly divided into training and testing sets.
Speech was recorded in a quiet environment using high-quality microphones, where native speakers were asked to read pre-prepared texts line by line.
Note: This repository… See the full description on the dataset page: https://huggingface.co/datasets/duymanh1606/vietnamese_asr.dolly-audio-1000h-vietnamese
Dolly-Audio: Vietnamese Multi-Speaker High-Quality Speech Corpus
Dataset Summary
Dolly-Audio is a large-scale, high-quality Vietnamese speech corpus created by the Dolly AI Team.
Inspired by Dolly, the world’s first cloned mammal, the project aims to advance research in Vietnamese speech synthesis, speech recognition, and voice modeling.
This release provides nearly 1,000 hours of professionally cleaned audio, featuring 152 speakers across different Vietnamese regions and… See the full description on the dataset page: https://huggingface.co/datasets/vuhoanhuy/dolly-audio-1000h-vietnamese.vietnamese-speech-recognition
Vietnamese Speech Dataset
Dataset comprises 10+ hours of telephone dialogues in Vietnamese, collected from 20 native speakers across various topics and domains. It is designed for research in speech recognition, focusing on various recognition models, primarily aimed at meeting the requirements for automatic speech recognition (ASR) systems.
By utilizing this dataset, researchers and developers can advance their understanding and capabilities in transcribing audio, and natural… See the full description on the dataset page: https://huggingface.co/datasets/UniDataPro/vietnamese-speech-recognition.test-vietnameseS2T_English_Vietnamese_AuVi_2vietnamese-girl-natural-voice-2h
🌸 Giọng Nữ Việt Nam Tự Nhiên - Dataset 2 Giờ (Phiên Bản Tốt Nhất 2026)
Giọng nữ trẻ, ngọt ngào, rõ ràng, tự nhiênSiêu phù hợp để train GPT-SoVITS • Fish Speech • F5-TTS • CosyVoice • XTTS và các model voice cloning khác.
🛠️ Tool dùng
Turn YouTube URLs or local media into a TTS-ready voice dataset: MinhxThanh/Voice-Dataset-Builder
✨ Thông tin Dataset
Thời lượng: Khoảng 2 giờ (đã xử lý sạch + phân đoạn tốt)
Giọng nói: Giọng nữ (girl voice) - ngọt… See the full description on the dataset page: https://huggingface.co/datasets/minhxthanh/vietnamese-girl-natural-voice-2h.28k_vietnamese_voice_augmented_of_VigBigData
Dataset Card for "28k_vietnamese_voice_augmented_of_VigBigData"
More Information needed
vietnamese-ttsVietnameseCarASR
Car Brand ASR Dataset
Description
This dataset is specifically designed for Automatic Speech Recognition (ASR) and Named Entity Recognition (NER) tasks focused on Vietnameese autom brands. It features audio recordings of spoken text containing various car brand names.
The primary objective of this dataset is to assist models in accurately identifying, transcribing, and extracting automotive brand entities from conversational or command-based audio.
Vietnamese_Medical_Consultationemotion-tts-vietnamese-sample18k_vietnamese_voice_augmented_of_VigBigData
Dataset Card for "18k_vietnamese_voice_augmented_of_VigBigData"
More Information needed
vlsp2023-vietnamese-asrVietnamese-audio-question-answeringvietnamese-tts-datasetvietnamese-speech-datasetvietnamese-voices
Vietnamese Voices Dataset
Đây là bộ dữ liệu giọng nói tiếng Việt...
Cách sử dụng
Bạn có thể dễ dàng tải và sử dụng bộ dữ liệu này với thư viện datasets của Hugging Face:
from datasets import load_dataset
ds = load_dataset("hungkieu/vietnamese-voices")
print(ds)
