datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
genshin-voice-v3.3-mandarin
Dataset Card for Genshin Voice
Dataset Description
Dataset Summary
The Genshin Voice dataset is a text-to-voice dataset of different Genshin Impact characters unpacked from the game.
Languages
The text in the dataset is in Mandarin.
Dataset Creation
Source Data
Initial Data Collection and Normalization
The data was obtained by unpacking the Genshin Impact game.
Who are the source language producers?
The… See the full description on the dataset page: https://huggingface.co/datasets/hanamizuki-ai/genshin-voice-v3.3-mandarin.Cat-Mangahindi_audio_dataset_testrirmega
RIRmega v2 — Dataset card (Hugging Face)
This folder holds the v2 artifacts intended for the Hugging Face dataset mandipgoswami/rirmega. When published, use revision v2.0.0 for the v2 release.
Dataset description
RIRmega v2 extends the existing RIRmega v1 dataset with:
A versioned metadata schema (metadata_v2.parquet) with acoustic metrics (RT60, DRR, C50, C80, D50, EDT), quality-control grades, and provenance.
A QC report (qc_report.parquet) with checks and outlier… See the full description on the dataset page: https://huggingface.co/datasets/mandipgoswami/rirmega.whisper-rirmega-bench
Whisper-RIR-Mega: Paired Clean↔Reverberant Speech Robustness Benchmark
Dataset Summary
Whisper-RIR-Mega is a benchmark dataset of paired clean and reverberant speech for evaluating ASR robustness to room acoustics. Each sample consists of:
audio_clean: Clean speech (LibriSpeech test-clean, 16 kHz)
audio_reverb: Same utterance convolved with one RIR from RIR-Mega (v2)
text_ref: Ground-truth transcript
RIR metadata: rir_id, RT60, DRR, C50, etc. when available
Technical… See the full description on the dataset page: https://huggingface.co/datasets/mandipgoswami/whisper-rirmega-bench.MANGO
MANGO: A Corpus of Human Ratings for Speech
MANGO (MUSHRA Assessment corpus using Native listeners and Guidelines to understand human Opinions at scale) is the first large-scale dataset designed for evaluating Text-to-Speech (TTS) systems in Indian languages.
Key Features:
255,150 human ratings of TTS-generated outputs and ground-truth human speech.
Covers two major Indian languages: Hindi & Tamil, and English.
Based on the MUSHRA (Multiple Stimuli with Hidden Reference… See the full description on the dataset page: https://huggingface.co/datasets/ai4bharat/MANGO.karim-mansouri-40kbps
Karim Mansouri
Part of Maqra, an open, verified archive of verse-by-verse Qur'an recitations mirrored from everyayah.com.
Set
karim-mansouri-40kbps
Style
murattal
Riwayah
hafs
Kind
recitation
Bitrate
40 kbps
Ayah files
6348 (1013 MiB)
Verified against the upstream MD5 list
6348
Ayahs absent upstream
0
Upstream folder
Karim_Mansoori_40kbps
Files
One MP3 per ayah, named SSSAAA.mp3 (surah 3 digits, ayah 3 digits). 001001.mp3 is… See the full description on the dataset page: https://huggingface.co/datasets/maqra-project/karim-mansouri-40kbps.genshin-voice-v3.5-mandarin
Dataset Card for Genshin Voice
Dataset Description
Dataset Summary
The Genshin Voice dataset is a text-to-voice dataset of different Genshin Impact characters unpacked from the game.
Languages
The text in the dataset is in Mandarin.
Dataset Creation
Source Data
Initial Data Collection and Normalization
The data was obtained by unpacking the Genshin Impact game.
Who are the source language producers?
The… See the full description on the dataset page: https://huggingface.co/datasets/hanamizuki-ai/genshin-voice-v3.5-mandarin.nasle-mana-clean-chunked-30s
Nasl-e-Mana Clean Speech Corpus — Sentence-Safe 30s Chunks
Training-oriented WAV chunks derived from the public Nasl-e-Mana magazine audio corpus. Chunks target approximately 30 seconds and are cut at detected acoustic pauses; the labeled configuration additionally assigns only complete source-text sentences to each chunk.
Configuration
Rows
Columns
Meaning
labeled (train/)
9,886
audio, label
Sentence-grouped text/audio pairs from duration-compatible source-text… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/nasle-mana-clean-chunked-30s.nasle-mana-clean-chunked-30s-avasanj
Nasl-e-Mana Clean Persian Speech — corrected 30-second chunks
Corrected, provenance-preserving audio chunks collected from the Nasl-e-Mana magazine website, generated on 2026-08-30. This release supersedes the earlier unreliable proportional-mapping chunk export; that older release was not used here.
Splits
Split
Rows
Audio
Columns
labeled
4,981
41.41 hours
audio, label
to_transcribe
11,127
92.72 hours
audio
The labeled split contains the… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/nasle-mana-clean-chunked-30s-avasanj.coser_v3_manualsna-dataset-annotated
manassehzw/sna-dataset-annotated
An annotated, speaker-relabelled, and loudness-normalised Shona (sna) speech dataset prepared through a reproducible Modal-based data engineering pipeline.
This release addresses speaker label contamination in the original source labels by replacing identity columns with acoustically-derived speaker assignments.
Why this annotated release exists
The original source speaker labels are contaminated (multiple voices assigned to the… See the full description on the dataset page: https://huggingface.co/datasets/manassehzw/sna-dataset-annotated.sna-waxal-annotated-unlabeled
Shona WAXAL annotated-unlabeled checkpoint
This is a self-contained operational checkpoint for pseudo-labeling Shona ASR
data. It contains 90,253 conservatively segmented FLAC clips
(441.585 hours), but intentionally contains no transcripts.
Fields
transcription is intentionally empty.
speaker_id is an approximate source-blind EOM cluster or unknown;
speaker_clip_count is zero for unknown assignments.
gender is always unknown; available classifiers were not… See the full description on the dataset page: https://huggingface.co/datasets/manassehzw/sna-waxal-annotated-unlabeled.genshin-voice-v3.4-mandarin
Dataset Card for Genshin Voice
Dataset Description
Dataset Summary
The Genshin Voice dataset is a text-to-voice dataset of different Genshin Impact characters unpacked from the game.
Languages
The text in the dataset is in Mandarin.
Dataset Creation
Source Data
Initial Data Collection and Normalization
The data was obtained by unpacking the Genshin Impact game.
Who are the source language producers?
The… See the full description on the dataset page: https://huggingface.co/datasets/hanamizuki-ai/genshin-voice-v3.4-mandarin.medical_noise_data
Medical Noise Dataset
Dữ liệu âm thanh tiếng Việt đã được tăng cường nhiễu (noise augmentation + RIR convolution).
Nguồn gốc
Audio gốc: dolly-vn/dolly-audio-1000h-vietnamese
Noise sources: YouTube extracted, WHAM!, Hospital ambient noise
RIR: Real RIR far-field (RVB2014)
Cách load
from datasets import load_dataset
ds = load_dataset("manhcuong2005/medical_noise_data")
shona-bible-bdsc-aligned
Shona Bible Speech Alignment Dataset
Lossless, verse-aligned Shona Bible speech dataset derived from the BDSC
source audio made available by Biblica, Inc. through Open.Bible. This release
contains the complete Bible: 66 books, 1,189 chapters, and 31,284 speech
segments covering approximately 75.55 hours.
Dataset summary
Language: Shona (sna)
Speaker: narrator 1
Speaker sex: male
Books: 66
Clips: 31,284
Audio: approximately 75.55 hours
Audio format: mono 16 kHz… See the full description on the dataset page: https://huggingface.co/datasets/manassehzw/shona-bible-bdsc-aligned.ASVspoof2021_DF
ASVspoof 2021 DF
Benchmark-ready packaging of the DeepFake (DF) evaluation partition from ASVspoof 2021 for speech anti-spoofing and synthetic / deepfake voice detection.
Overview
This dataset contains the DF evaluation subset of the ASVspoof 2021 challenge. The task is binary classification: bonafide (genuine human speech) vs. spoof (synthetic, converted, or otherwise manipulated speech). The original dataset is available at… See the full description on the dataset page: https://huggingface.co/datasets/manojkumarainala152/ASVspoof2021_DF.AnomalyMachine-50K
Dataset Summary
AnomalyMachine-50K is a fully synthetic industrial machine sound anomaly detection dataset designed for research on acoustic monitoring, predictive maintenance, and sound event detection.The dataset contains 50,000 monaural audio clips, each 10 seconds long at 22,050 Hz, covering six industrial machine types, multiple operating conditions, and diverse anomaly types under different signal-to-noise ratios.
The dataset is generated entirely via signal-processing based… See the full description on the dataset page: https://huggingface.co/datasets/mandipgoswami/AnomalyMachine-50K.common_voice_16_AR_pseudo_labelledaishell1-mandarin2024.09.14_AGH_manual_Temperature_tests_20-55_Degrees
Temperature tests (20-55 Degrees)
Dataset author: Oğuzhan Berke Özdil
What was measured?
Methods
Type of needle: Quincke 22G 9cm
Needle Was attached with 3d printed needle attachment directly in front of the sound hole of the microphone
Type of microphone:
How the measurements were taken: manual
Phantoms
thickness of foams : 3 cm
Foam is the same as we used in Magdeburg foam phantom.
Foams were soaked in warm/hot water to create different… See the full description on the dataset page: https://huggingface.co/datasets/VibroNav/2024.09.14_AGH_manual_Temperature_tests_20-55_Degrees.Lahgtna-saudi
Lahgtna Saudi (Mans1611/Lahgtna-saudi)
Saudi-dialect subset prepared for Arabic ASR fine-tuning
(from oddadmix/dialectal-arabic-lahgtna-v2, filtered to language == "sa").
Splits
Split
Rows
train
11,030
test
581
Columns
audio, text, language, duration
Features
{'audio': Audio(sampling_rate=16000, decode=False, num_channels=None, stream_index=None), 'text': Value('string'), 'language': Value('string'), 'duration':… See the full description on the dataset page: https://huggingface.co/datasets/Mans1611/Lahgtna-saudi.shona-waxal-pseudo-labeled
Shona WAXAL pseudo-labelled speech
This release contains 90,253 Shona speech clips, totalling 441.585 hours. Each clip keeps its original FLAC audio and a Sunbird Whisper pseudo-transcript. These are model outputs, not human reference transcriptions.
What this release contains
The source is the unlabeled Shona ASR split from WAXAL NLP, preserved in the operational checkpoint manassehzw/sna-waxal-annotated-unlabeled. The source checkpoint has no transcripts. This… See the full description on the dataset page: https://huggingface.co/datasets/manassehzw/shona-waxal-pseudo-labeled.aishell3-mandarinyoutube-asr-datasetl2-arctic-manual-v5.0-16k
l2-arctic-manual-v5.0-16k
This dataset is a prepared derivative of L2-ARCTIC v5.0 that keeps only
the manually annotated material and converts the audio to 16 kHz mono FLAC.
It is designed to plug into the current peacock-asr training code, which
can consume a Hugging Face dataset with audio plus phonemes.
Included splits
train: 1800 rows, 1.84 hours
validation: 899 rows, 0.94 hours
test: 900 rows, 0.88 hours
suitcase: 22 rows, 0.44 hours
The scripted subset uses the… See the full description on the dataset page: https://huggingface.co/datasets/chikingsley/l2-arctic-manual-v5.0-16k.nasle-mana-clean
Nasl-e-Mana Speech Corpus
Clean, playable Persian speech audio collected from the public Nasl-e-Mana magazine WordPress site. The export contains two intentionally different collections:
Split
Rows
Columns
Meaning
labeled configuration (train/)
809
audio, label
Audio with recovered article text. These clips were identified as the consistent female source-text voice and are kept together.
to_transcribe configuration (to_transcribe/)
626
audio
Playable audio for which… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/nasle-mana-clean.pseudolabel-mandarin-large-v3-timestamp
Pseudolabel Mandarin using Whisper Large V3
Original dataset from https://openslr.org/33/, Heavy postprocessing and post-translation to improve pseudolabeled Whisper Large V3.
Southwestern_Mandarin_Dialectal_Expressions_Speechindic-diarbench
Indic DiarBench
A multilingual joint diarization and ASR benchmark for Indian languages, spanning all 22 scheduled languages of India with approximately 108 hours of natural multi-speaker audio.
Dataset Summary
Indic DiarBench is a conversational speech benchmark designed to evaluate speaker-attributed ASR in realistic multi-speaker settings for Indian languages. All annotations are human-corrected with time-aligned, speaker-attributed transcriptions. The dataset… See the full description on the dataset page: https://huggingface.co/datasets/manojkumarcs/indic-diarbench.
