datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
VoiceCommandAudioThis is mainly used for fine tune "VoiceCommand" a speech congnition MOD dedicated for SilentHunter game series
Hadou-Voice-Dataset
Hadou Voice Dataset
ハドウ本人が収録した、日本語音声データセットです。
このページで、特徴の異なる2種類のデータセットを公開しています。
配布データ
設定名
内容
音声数
合計時間
v1(おすすめ)
Hadou Calm Voice Dataset v1。落ち着いた中音域、AIキャラクター向けボイスが多めの音声データ
966
約114.02分
v0
Hadou ITA Corpus Dataset v1。ITAコーパスを読み上げた自然な話し声
424
約38.95分
v1 には、AICAコーパス500文、ITAコーパス324文、感情・態度付き90文、同文異演技40文、強度段階12文を収録しています。
v1の詳細: v1/README.txt
v0の詳細: v0/README.txt
読み込み例
from datasets import load_dataset
# 新しい966音声(既定)
dataset =… See the full description on the dataset page: https://huggingface.co/datasets/hadou1225/Hadou-Voice-Dataset.ht-voice-dataset
Ansanb done vwa an Kreyòl pou antrene DeepSPeech.
Dataset sa a gen plis pase 7 è tan anrejistreman vwa ak prèske 100 moun an Kreyòl pou bati sistèm ASR ak TTS pou lang Kreyòl la.
Pifò nan done yo soti nan "CMU Haitian Creole Speech Recognition Database" la.
Done sa yo gentan filtre epi òganize pou ka antrene modèl DeepSPeech Mozilaa a.
Si toutfwa ou ta bezwen jwenn plis enfòmasyon sou jan done yo ranje a epi kisa ou ka fè avèk yo, tcheke DeepSpeech Readme.
voiceguard-competition
VoiceGuard — Deepfake Audio Detection Competition
Pelatnas IOAI 2026 | Task 3 of 3
Detect whether a 4-second audio clip is real human speech or AI-generated (TTS/deepfake). Submit probability scores — AUROC is the metric.
Task
Input: .wav audio file (4 seconds, 16 kHz mono)Output: score — probability (0–1) that the audio is fakeMetric: AUROC (Area Under ROC Curve)
Dataset
Split
Real
Fake
Total
Train
2,874
2,874
5,748
Test
627
627
1,254… See the full description on the dataset page: https://huggingface.co/datasets/fassabilf/voiceguard-competition.emotion-voice-dataset
emotion voice dataset
Developed by Aryan Singh Chandel (Shiro) at Rustamji Institute of Technology (RJIT).
📝 Overview
This repository contains assets for emotion voice dataset. It is a professional research component of the Shiro AI ecosystem.
🚀 Status
The core files are live. Detailed usage instructions and technical benchmarks are currently being compiled for the elite release.
zamai-pashto-voice2voice
ZamAI Pashto Voice2Voice
This dataset contains Pashto voice-to-voice preparation metadata for speech and translation experiments. It focuses on Pashto speech records, dialect information, transcript text, and a small viewer-ready sample manifest.
Configs
from datasets import load_dataset
metadata = load_dataset("ZamAI-Pashto/zamai-pashto-voice2voice", "metadata")
sample = load_dataset("ZamAI-Pashto/zamai-pashto-voice2voice", "viewer_sample")
Files… See the full description on the dataset page: https://huggingface.co/datasets/ZamAI-Pashto/zamai-pashto-voice2voice.col_voicezamai-pashto-voice2voice
ZamAI Pashto Voice2Voice
Languages: psLicense: cc-by-4.0Task categories: automatic-speech-recognition, audio-to-audioSize categories: n<1K
Summary
This dataset is part of the ZamAI Pashto data collection. It is intended for automatic-speech-recognition, audio-to-audio tasks in Pashto.
How to use
from datasets import load_dataset
dataset = load_dataset("tasal9/zamai-pashto-voice2voice")
print(dataset)
Configs
default: load with… See the full description on the dataset page: https://huggingface.co/datasets/tasal9/zamai-pashto-voice2voice.common_voice_19_uk_croppedoil_voices
