datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
internal-ast-eval-cacheRAVDESShey-computer-speech-commands
Hey Computer: Speech Command Recognition
Dataset Summary
A public, viewer-ready educational challenge dataset. Host-only scoring data and hidden targets are excluded.
Splits
Split
Examples
Description
train
13,192
Labeled training data
test
3,295
Public inputs with withheld target labels or annotations
Data Fields
Field
Type
audio
Audio
id
string
label
string (test sentinel: unlabeled)… See the full description on the dataset page: https://huggingface.co/datasets/hoangbang/hey-computer-speech-commands.speak-the-digit
Speak the Digit: Spoken Digit Recognition
Dataset Summary
A public, viewer-ready educational challenge dataset. Host-only scoring data and hidden targets are excluded.
Splits
Split
Examples
Description
train
2,400
Labeled training data
test
600
Public inputs with withheld target labels or annotations
Data Fields
Field
Type
audio
Audio
id
string
label
string (test sentinel: unlabeled)… See the full description on the dataset page: https://huggingface.co/datasets/hoangbang/speak-the-digit.m-meldkinh-phap-hoa-ke-trom-huongNormalized using https://github.com/oysterlanguage/emiliapipex
@article{emilia,
title={Emilia: An Extensive, Multilingual, and Diverse Speech Dataset for Large-Scale Speech Generation},
author={He, Haorui and Shang, Zengqiang and Wang, Chaoren and Li, Xuyuan and Gu, Yicheng and Hua, Hua and Liu, Liwei and Yang, Chen and Li, Jiaqi and Shi, Peiyang and Wang, Yuancheng and Chen, Kai and Zhang, Pengyuan and Wu, Zhizheng},
journal={arXiv},
volume={abs/2407.05361}… See the full description on the dataset page: https://huggingface.co/datasets/hr16/kinh-phap-hoa-ke-trom-huong.vi-en-ast-testsetvoice_dataset
Dataset Card for "voice_dataset"
More Information needed
tdtu_voice_dataset
Dataset Card for "tdtu_voice_dataset"
More Information needed
bmeldcntt2-audio-dataset
Dataset Card for "cntt2-audio-dataset"
More Information needed
en-vi-ast-testsetvoice_dataset_yt
Dataset Card for "voice_dataset_yt"
More Information needed
omnivoice-vie
OmniVoice VI — Giọng Việt + SRT lồng tiếng
Dataset chứa 6 giọng tiếng Việt và công cụ speak.py để chạy trên Google Colab với OmniVoice.
Giọng có sẵn
Slug
Tên
ban_mai
Ban Mai
lan_trinh
Lan Trinh
ngan_ha
Ngan Ha
ngoc_huyen
Ngoc Huyen
thao_trinh
Thao Trinh
tuong_vy
Tuong Vy
Mỗi giọng gồm profile.json, voice.pt (prompt cache), audio mẫu và ref_text.txt.
Chạy trên Colab
Mở notebook colab/Omivoice_VI_Colab.ipynb
Đặt HF_REPO =… See the full description on the dataset page: https://huggingface.co/datasets/hoanglinhn0/omnivoice-vie.internal-ast-eval-failuresvivos-processed
VIVOS Processed VoxCPM2 References
This dataset contains 325 quality-first, nested voice-cloning references for the 65 speakers in the VIVOS Vietnamese corpus. Each speaker has nominal 5, 10, 15, 20, and 30-second variants. Whole source utterances are retained, so actual_seconds is the authoritative duration.
Configuration
references is directly usable for VoxCPM2 inference. Use audio as reference_wav_path; for transcript-assisted, highest-fidelity cloning, use… See the full description on the dataset page: https://huggingface.co/datasets/christian-hoang-04/vivos-processed.sodo_hoaimy_newsodo_hoaimy_1AV_FFIA3kmy_dataset_testMMFFIAuser_5476d2c924204b6f9e38713118fdb9b2_datasetuser_03aa5df890b64866be4aef51a01c0a8a_datasetAbstractTTStrinh_hoai_quang_tam77r0dh_datasetuser_35621758bf084337aad673e1cc332d6f_datasetuser_da91d399b47141ccaa812c8b16e8c380_datasetuser_359d53fbd48b405daf1d7a67aed75197_datasetuser_476da26872df492f830a65925d422651_dataset
