datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Custom_common_voice_dataset_using_RVC
Custom Data Augmentation for low resource ASR using Bark and Retrieval-Based Voice Conversion
Custom common_voice_v11 corpus with a custom voice was was created using RVC(Retrieval-Based Voice Conversion)
The model underwent 200 epochs of training, utilizing a total of 1 hour of audio clips. The data was scraped from Youtube.
The audio in the custom generated dataset is of a YouTuber named
Ajay Pandey
Description
license: cc0-1.0
language:
- hi… See the full description on the dataset page: https://huggingface.co/datasets/Aniket-Tathe-08/Custom_common_voice_dataset_using_RVC.VoiceCommandAudioThis is mainly used for fine tune "VoiceCommand" a speech congnition MOD dedicated for SilentHunter game series
Hadou-Voice-Dataset
Hadou Voice Dataset
ハドウ本人が収録した、日本語音声データセットです。
このページで、特徴の異なる2種類のデータセットを公開しています。
配布データ
設定名
内容
音声数
合計時間
v1(おすすめ)
Hadou Calm Voice Dataset v1。落ち着いた中音域、AIキャラクター向けボイスが多めの音声データ
966
約114.02分
v0
Hadou ITA Corpus Dataset v1。ITAコーパスを読み上げた自然な話し声
424
約38.95分
v1 には、AICAコーパス500文、ITAコーパス324文、感情・態度付き90文、同文異演技40文、強度段階12文を収録しています。
v1の詳細: v1/README.txt
v0の詳細: v0/README.txt
読み込み例
from datasets import load_dataset
# 新しい966音声(既定)
dataset =… See the full description on the dataset page: https://huggingface.co/datasets/hadou1225/Hadou-Voice-Dataset.tibetan-voice-benchmarkBenchmark of Tibetan Speech-To-Text dataset created by Monlam AI x Openpecha in 2024
All the transcripts have been reviewed by at least one person in addition to the original transcriber.
Data was taken on 15 July 2024 02∶47∶06 PM.
dept
desc
Count
STT_AB
Audio book
1000
STT_CS
Children Speech
1367
STT_HS
History
1000
STT_MV
Tibetan Movies
1000
STT_NS
Natural Speech
1000
STT_NW
News
1000
STT_PC
Podcast
1000
STT_TT
Tibetan Teachings
1000
grade column is used to… See the full description on the dataset page: https://huggingface.co/datasets/openpecha/tibetan-voice-benchmark.hausa_common_voiceThis dataset is from the common voice corpus 7.0 using the Hausa dataset
ht-voice-dataset
Ansanb done vwa an Kreyòl pou antrene DeepSPeech.
Dataset sa a gen plis pase 7 è tan anrejistreman vwa ak prèske 100 moun an Kreyòl pou bati sistèm ASR ak TTS pou lang Kreyòl la.
Pifò nan done yo soti nan "CMU Haitian Creole Speech Recognition Database" la.
Done sa yo gentan filtre epi òganize pou ka antrene modèl DeepSPeech Mozilaa a.
Si toutfwa ou ta bezwen jwenn plis enfòmasyon sou jan done yo ranje a epi kisa ou ka fè avèk yo, tcheke DeepSpeech Readme.
voiceguard-competition
VoiceGuard — Deepfake Audio Detection Competition
Pelatnas IOAI 2026 | Task 3 of 3
Detect whether a 4-second audio clip is real human speech or AI-generated (TTS/deepfake). Submit probability scores — AUROC is the metric.
Task
Input: .wav audio file (4 seconds, 16 kHz mono)Output: score — probability (0–1) that the audio is fakeMetric: AUROC (Area Under ROC Curve)
Dataset
Split
Real
Fake
Total
Train
2,874
2,874
5,748
Test
627
627
1,254… See the full description on the dataset page: https://huggingface.co/datasets/fassabilf/voiceguard-competition.Prompts_for_Voice_cloning_and_TTScommon-voice-kinyarwanda-english-dataset
Kinyarwanda-English Commonvoice dataset
A compilation of Kinyarwanda-english dataset to be used to train multi-lingual ASR
Note: The audio dataset shall be added in the future
home-assistant-local-llm-voice-benchmark
Home Assistant Local LLM Voice Benchmark
Per-model tool-call accuracy and component latency for running a Home Assistant voice assistant against local LLMs.
Measured per-model tool-call accuracy and component latency for running a Home Assistant voice assistant against local LLMs.
Broken out by pipeline component rather than reported as one opaque round trip, so you can tell whether your latency is wake-word, speech-to-text, the model, or text-to-speech before you go optimizing… See the full description on the dataset page: https://huggingface.co/datasets/iBlessi/home-assistant-local-llm-voice-benchmark.librispeech40emotion-voice-dataset
emotion voice dataset
Developed by Aryan Singh Chandel (Shiro) at Rustamji Institute of Technology (RJIT).
📝 Overview
This repository contains assets for emotion voice dataset. It is a professional research component of the Shiro AI ecosystem.
🚀 Status
The core files are live. Detailed usage instructions and technical benchmarks are currently being compiled for the elite release.
mozilla-common-voice-23-bel-texts-exportzamai-pashto-voice2voice
ZamAI Pashto Voice2Voice
This dataset contains Pashto voice-to-voice preparation metadata for speech and translation experiments. It focuses on Pashto speech records, dialect information, transcript text, and a small viewer-ready sample manifest.
Configs
from datasets import load_dataset
metadata = load_dataset("ZamAI-Pashto/zamai-pashto-voice2voice", "metadata")
sample = load_dataset("ZamAI-Pashto/zamai-pashto-voice2voice", "viewer_sample")
Files… See the full description on the dataset page: https://huggingface.co/datasets/ZamAI-Pashto/zamai-pashto-voice2voice.voices-kokorocol_voiceBo-voice-v1.0.0
Tibetan STT Benchmark Model Card
Bo-voice-v1.0.0 is a high-fidelity benchmark for Tibetan Speech-to-Text (STT) technology. It provides a rigorous, multi-domain evaluation set to measure Automatic Speech Recognition (ASR) performance across diverse acoustic environments and speaking styles.
### Dataset Overview
Snapshot Date: 15 July 2024, 02:47:06 PM
Total Samples: 8,367 audio-transcript pairs.
Verification: Every transcript has been reviewed by at least one expert in… See the full description on the dataset page: https://huggingface.co/datasets/MonlamAI/Bo-voice-v1.0.0.zamai-pashto-voice2voice
ZamAI Pashto Voice2Voice
Languages: psLicense: cc-by-4.0Task categories: automatic-speech-recognition, audio-to-audioSize categories: n<1K
Summary
This dataset is part of the ZamAI Pashto data collection. It is intended for automatic-speech-recognition, audio-to-audio tasks in Pashto.
How to use
from datasets import load_dataset
dataset = load_dataset("tasal9/zamai-pashto-voice2voice")
print(dataset)
Configs
default: load with… See the full description on the dataset page: https://huggingface.co/datasets/tasal9/zamai-pashto-voice2voice.voxclinbench
VoxClinBench (Hugging Face dataset mirror)
Cross-lingual, cross-disease clinical voice biomarker benchmark. This
Hugging Face dataset mirror ships the split manifests, Croissant
metadata, datasheet, and reference baseline prediction CSVs.
Raw audio is NOT distributed here. Each of the five upstream
corpora must be obtained directly from its provider under that
corpus's data use agreement (DUA). See the
GitHub mirror
for the evaluation harness and the voxbench fetch CLI.… See the full description on the dataset page: https://huggingface.co/datasets/voice-bench-submission/voxclinbench.common_voice_13_0_zh_pseudo_labelledcommon-voicecommon_voice_13_0_mn_pseudo_test_smallvoice-acting-instructionscommon_voice_16_0_fa_pseudo_labelledcommon_voice_13_0_thai_small_pseudo_labelledlinkedin-top-voices-market-alpaca
linkedin-top-voices-market-alpaca Dataset
Dataset containing 300 records for fine-tuning language models.
Columns
instruction
input
output
input_tokens
output_tokens
input_cost
output_cost
total_cost
smart-home-voice-commands-v1
Smart Home Voice Commands V1
Description
Dataset of smart home voice assistant commands labeled by device control category.
Data Fields
text: Voice command
intent: Device action label
Intents
lights_on
lights_off
increase_temperature
decrease_temperature
play_music
License
CC-BY-4.0
mdd-voice-biomarker-data
mdd-voice-biomarker-data
Small task-specific mirror of public input files used by Terminal Bench Science task mdd-voice-biomarker.
The files are redistributed here to make benchmark Docker builds faster and more reproducible. See task instructions and original upstream sources for dataset-specific citation and licensing details.
Contents
data/: input files downloaded by the task Dockerfiles.
SHA256SUMS.txt: SHA-256 checksums for all files under data/.
common_voice_16_1_spanish_test_set
Dataset Card for Common Voice Corpus 16 Spanish Dataset
Acknowledgement
The dataset belongs to COMMON VOICE MOZILLA FOUNDATION.
I just uploaded the spanish test set (from HERE : https://huggingface.co/datasets/mozilla-foundation/common_voice_16_1/tree/main)
Dataset Summary
The Common Voice dataset consists of a unique MP3 and corresponding text file.
Languages
Spanish
How to use
The datasets library allows you to load and pre-process… See the full description on the dataset page: https://huggingface.co/datasets/omarsou/common_voice_16_1_spanish_test_set.common_voice_16_1_hi_pseudo_labelled
