datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
data-voice-vietnamese-restaurant-quan-oc
Vietnamese Restaurant Order Speech
This dataset contains Vietnamese spoken restaurant orders paired with text transcripts. Each utterance typically includes a table number, item quantities, dishes, drinks, and add-ons.
Dataset Structure
Files are split into subdirectories by filename-derived speaker_code to satisfy Hugging Face repository file-count limits:
metadata.csv: one row per audio sample.
audio/{speaker_code}/*.wav: mono WAV audio files.… See the full description on the dataset page: https://huggingface.co/datasets/EmilyNguyen235/data-voice-vietnamese-restaurant-quan-oc.Bahnar_Vietnamese
Bahnar Speech Translation Dataset
This dataset contains Bahnar speech audio aligned with Bahnar, Vietnamese, and English text. It was created from internet data sources and automatically aligned using the pipeline available at Bahnar-Vietnamese-S2TT.
The main purpose of this dataset is to support research on low-resource speech-to-text translation (S2TT), especially direct translation from Bahnar speech to Vietnamese text.
Data Statistics
Train: 113,830… See the full description on the dataset page: https://huggingface.co/datasets/cuong06/Bahnar_Vietnamese.Vietnamese-asr-leaderboard
📊 Vietnamese Open ASR Evaluation Dataset Storage
Kho lưu trữ dữ liệu nhãn bảo mật (Ground Truth) phục vụ cho hệ thống Vietnamese Open ASR Leaderboard. Toàn bộ dữ liệu được tổng hợp từ 9 bộ dữ liệu tiếng Việt công khai lớn nhất hiện nay, sau đó trải qua quy trình chuẩn hóa văn bản nghiêm ngặt để làm thước đo chuẩn mực đánh giá hiệu năng các mô hình nhận dạng giọng nói (ASR).
[!TIP]
🚀 NỘP BÀI ĐÁNH GIÁ TẠI ĐÂY:
📈 1. Bảng Thống Kê Chi Tiết Hệ Dữ Liệu… See the full description on the dataset page: https://huggingface.co/datasets/VietAudio-team/Vietnamese-asr-leaderboard.vietnamese_asr
Vietnamese ASR (VIVOS Corpus)
Dataset Overview
The VIVOS corpus is a free Vietnamese speech dataset consisting of 15 hours of recorded speech. This corpus was meticulously prepared for Automatic Speech Recognition (ASR) tasks and is explicitly divided into training and testing sets.
Speech was recorded in a quiet environment using high-quality microphones, where native speakers were asked to read pre-prepared texts line by line.
Note: This repository… See the full description on the dataset page: https://huggingface.co/datasets/duymanh1606/vietnamese_asr.vietnamese-speech-recognition
Vietnamese Speech Dataset
Dataset comprises 10+ hours of telephone dialogues in Vietnamese, collected from 20 native speakers across various topics and domains. It is designed for research in speech recognition, focusing on various recognition models, primarily aimed at meeting the requirements for automatic speech recognition (ASR) systems.
By utilizing this dataset, researchers and developers can advance their understanding and capabilities in transcribing audio, and natural… See the full description on the dataset page: https://huggingface.co/datasets/UniDataPro/vietnamese-speech-recognition.Vietnamese_Medical_ConsultationVietnamese_ASR_TestingDataVietnamese-Speech-Dataset
🎧 Vietnamese Speech Dataset
The Vietnamese Speech Dataset is a large-scale speech audio dataset designed to support advanced AI systems with diverse and high-quality audio data. It includes 179 hours of recorded speech data across 710 files, delivered in MP3 and WAV formats, with a total size of 280 MB. This well-structured audio dataset provides balanced and representative voice data, featuring 52% female and 48% male speakers, with age coverage from 18 to 50+ years. The dataset… See the full description on the dataset page: https://huggingface.co/datasets/Speech-data/Vietnamese-Speech-Dataset.vietnamese-speech-recognition
Vietnamese Speech Dataset - 10+ hours
Dataset comprises 10+ hours of telephone dialogues in Vietnamese, collected from 20 native speakers across various topics and domains. It is designed for research in speech recognition, focusing on various recognition models, primarily aimed at meeting the requirements for automatic speech recognition (ASR) systems. - Get the data
Dataset characteristics:
Characteristic
Data
Description
Audio of telephone dialogues in… See the full description on the dataset page: https://huggingface.co/datasets/ud-nlp/vietnamese-speech-recognition.Vietnamese_ASR_TestingData_Old
About
This dataset is for only ASR testing in Vietnamese.
We collect data from various sources.
This is the first version of the dataset.
YodaLingua-Vietnamese
YodaLingua-Vietnamese
YodaLingua is a high-quality speech dataset designed for training text-to-speech (TTS) systems, ASR models, and any application requiring clean, well-aligned audio–text pairs.This release contains the Vietnamese portion of the multilingual YodaLingua collection.
🧾 Dataset Overview
Property
Value
Total clips
47,223 audio–transcription pairs
Total duration
125 hours
Speakers
1,426 distinct speakers
Audio format
MP3 • mono • 24 kHz •… See the full description on the dataset page: https://huggingface.co/datasets/Thomcles/YodaLingua-Vietnamese.Vietnamese_Text2Speech_AB2
