datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
parler-tts_mls_eng_10k_snac_token_old
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/blanchon/parler-tts_mls_eng_10k_snac_token_old.maleo-short-1.5H
Dataset Card for Maleo Short 1.5H
Dataset Description
Dataset Summary
Maleo Short 1.5H is a manually curated, rigorously annotated speaker diarization dataset designed to benchmark State-of-the-Art (SOTA) models against complex, "in-the-wild" media domains. While modern diarization pipelines excel in controlled acoustic environments (like telephony or reading corpora), they heavily struggle with the overlapping speech, sound effects, and rapid speaker shifts… See the full description on the dataset page: https://huggingface.co/datasets/maleo-ai/maleo-short-1.5H.HALAS
Dataset Card for HALAS
Dataset Summary
HALAS (Hallucination Annotations for Large-scale ASR Systems) is a human-annotated dataset of hallucinations produced by modern automatic speech recognition (ASR) systems on real-world speech recordings. The dataset contains span-level hallucination annotations for ASR outputs generated from recordings in the Earnings22 corpus.
HALAS was introduced to address a key limitation in prior hallucination research: most existing… See the full description on the dataset page: https://huggingface.co/datasets/MatBar99/HALAS.sqp-tts-en
SQP TTS (English)
Synthesized speech for SQPsychConv_qwen-2.5, a synthetic CBT therapist-client
dialogue dataset (English).
Each configuration below corresponds to one TTS model. Load a single model
with:
from datasets import load_dataset
ds = load_dataset("sinselm/sqp-tts-en", "qwen3-tts")
Models included
qwen3-tts: https://huggingface.co/Qwen/Qwen3-TTS-12Hz-0.6B-Base
cosyvoice: https://huggingface.co/FunAudioLLM/Fun-CosyVoice3-0.5B-2512
fishaudio:… See the full description on the dataset page: https://huggingface.co/datasets/marleen-snsl/sqp-tts-en.fama-data
Dataset Description, Collection, and Source
The FAMA training data is the collection of English and Italian datasets for automatic speech recognition (ASR) and speech translation (ST)
used to train the FAMA models family.
The ASR section of FAMA is derived from the MOSEL data collection, including the automatic
transcripts obtained with Whisper and available in the HuggingFace MOSEL Dataset.
The ASR is further augmented with automatically transcribed speech from the… See the full description on the dataset page: https://huggingface.co/datasets/FBK-MT/fama-data.africanvoices-naija-batch1-summary
African Voices Naija Train Metadata Summary
This dataset contains a compact summary of metadata for the Naija training split, provided as CSV tables for inspection and analysis.
Files included:
batch_summary.csv
domain_distribution.csv
The repository contains metadata summaries only and does not include raw audio.
Moroccan-Arabic-Multimodal-Emotion-Recognition
MDER-MA — Moroccan Arabic Multimodal Emotion Recognition (TTS-aligned repackaging)
A repackaging of the MDER-MA dataset that pairs every audio clip with its Arabic (Moroccan dialect / Darija) transcript and ships speaker-disjoint train/validation/test splits.
Original dataset: Ouali, S. & El Garouani, S. (2025). MDER-MA: A multimodal dataset for emotion recognition in low-resource Moroccan Arabic language. Data in Brief. DOI: 10.1016/j.dib.2025.112005. Mendeley:… See the full description on the dataset page: https://huggingface.co/datasets/FatimahEmadEldin/Moroccan-Arabic-Multimodal-Emotion-Recognition.AYDID
AYDID: A Sub-Dialectal Yemeni Arabic Speech Corpus
First speech resource to label Yemeni Arabic at the sub-dialectal level:
17,500 utterances (22.53 h), 350 speakers, 7 classes — six regional varieties
(Adeni, Badawi, Hadrami, Sana'ani, Ta'izzi, Tihami) plus an MSA-proximal
Standard Yemeni control class. Balanced at 50 speakers / 2,500 utterances
per class.
Reproducibility tiers
To respect the copyright of the broadcast source material, raw audio is not
publicly… See the full description on the dataset page: https://huggingface.co/datasets/mansoorSaleh/AYDID.annomi-tts-en
AnnoMI TTS (English)
Synthesized speech for the AnnoMI motivational interviewing dialogues (English).
Each configuration below corresponds to one TTS model. Load a single model
with:
from datasets import load_dataset
ds = load_dataset("sinselm/annomi-tts-en", "qwen3-tts")
Models included
qwen3-tts: https://huggingface.co/Qwen/Qwen3-TTS-12Hz-0.6B-Base
cosyvoice: https://huggingface.co/FunAudioLLM/Fun-CosyVoice3-0.5B-2512
fishaudio:… See the full description on the dataset page: https://huggingface.co/datasets/marleen-snsl/annomi-tts-en.ipa-phonebpe-22lang
22-Language IPA Lexicons and PhoneBPE Vocabulary
This repository contains IPA pronunciation lexicons for 22 languages and a shared multilingual PhoneBPE vocabulary trained on the phoneme transcriptions of all 22 language training sets.
Languages
Code
Language
中文名称
ba
Bashkir
巴什基尔语
be
Belarusian
白俄罗斯语
cs
Czech
捷克语
de
German
德语
el
Greek
希腊语
en
English
英语
es
Spanish
西班牙语
fi
Finnish
芬兰语
fr
French
法语
it
Italian
意大利语
ku
Kurdish
库尔德语
ky… See the full description on the dataset page: https://huggingface.co/datasets/maxwellziweiwei/ipa-phonebpe-22lang.apple-speechanalyzer-vs-whisper-cpp-mac
Apple SpeechAnalyzer vs whisper.cpp on Mac
Four complete speech-recognition benchmark runs over the same deterministic
40-speaker LibriSpeech test-clean snapshot:
Engine
Model path
WER
CER
Repeated median post-speech latency
Repeated p95
Apple SpeechAnalyzer
progressiveTranscription on macOS 26.5
1.98%
1.02%
125–132 ms
194–201 ms
whisper.cpp server
1.8.4 · ggml-small.en
4.28%
1.79%
122–125 ms
152–161 ms
Every run completed 40/40 clips with no failures. Accuracy… See the full description on the dataset page: https://huggingface.co/datasets/researchaudio/apple-speechanalyzer-vs-whisper-cpp-mac.AYDID-audio
AYDID — Audio (gated access)
Segmented 16 kHz mono WAV audio for the AYDID sub-dialectal Yemeni Arabic corpus,
released for non-commercial academic research under a data-use agreement.
Audio is derived from Yemeni broadcast media. To respect source copyright, it is
not publicly redistributed; access is granted to identified researchers who
accept the terms above. Approved users can reproduce the full AYDID pipeline
end-to-end (feature extraction, segmentation, ASR preprocessing… See the full description on the dataset page: https://huggingface.co/datasets/mansoorSaleh/AYDID-audio.lost-in-speech
Lost in Speech
A trilingual benchmark for reference-free classification of synthetically introduced factual and contextual alterations in English, Russian, and Kazakh. It contains 12,013 samples derived from news articles, with text, synthesized speech, and ASR transcript representations used in the study.
Altered samples are LLM-generated rewrites with a controlled alteration type—contradiction, fabrication, or context inconsistency—and severity level—mild, moderate, or severe.… See the full description on the dataset page: https://huggingface.co/datasets/maristombayeva/lost-in-speech.French-Medical-Transcription-Benchmark
🩺 French Medical Transcription Evaluation Dataset
Ce dataset a été créé et ouvert à la communauté dans le cadre du développement R&D de LucioleScribe, la plateforme souveraine de transcription IA 100% locale, spécifiquement conçue pour les milieux médicaux et juridiques (compatibilité RGPD, HDS, et architectures Air-Gapped).
🔗 Découvrir LucioleScribe Édition Santé | ⚙️ Voir le Pipeline Technologique Local
📊 Présentation du Dataset
L'évaluation des modèles de… See the full description on the dataset page: https://huggingface.co/datasets/AWANNABY/French-Medical-Transcription-Benchmark.Kunkado_Maya_V2
Kunkado Maya V2 (Standardized)
Description
Ce dataset est une version transformée du dataset original RobotsMali/Kunkado (configuration human-reviewed).
L'objectif de cette version est de fournir des transcriptions dont les tags d'émotions et de bruits sont 100% compatibles avec les standards d'intégration de Maya One.
Transformations effectuées
Une pipeline de nettoyage automatique a été appliquée sur la colonne corrected-label pour générer la colonne… See the full description on the dataset page: https://huggingface.co/datasets/binaryMao/Kunkado_Maya_V2.swahili_zen_modelmodel:
https://huggingface.co/zenlm/zen3-asr
Code:
from transformers import AutoModelForSpeechSeq2Seq, AutoProcessor
import torch
import librosa
import numpy as np
model_id = "zenlm/zen3-asr"
processor = AutoProcessor.from_pretrained(model_id)
model = AutoModelForSpeechSeq2Seq.from_pretrained(model_id, torch_dtype=torch.float16, device_map="auto")
def transcribe(audio_path):
audio, sr = librosa.load(audio_path, sr=16000)
<!-- Pass raw waveform directly to processor -->… See the full description on the dataset page: https://huggingface.co/datasets/paulinenyaboe/swahili_zen_model.ktt-math-tutor-data
KTT Math Tutor — Data
Data artefacts for the AIMS KTT Hackathon Tier-3 submission
S2.T3.1 AI Math Tutor for Early Learners. Source code:
https://github.com/DrUkachi/ktt-math-tutor.
Contents
T3.1_Math_Tutor/
Core curriculum + seeds.
curriculum.json — 80 items × 5 sub-skills (counting, number
sense, addition, subtraction, word problem) with EN / FR / KIN
stems, difficulty 1–10, age bands 5–6 / 6–7 / 7–8 / 8–9, visual
asset keys, expected integer answer.… See the full description on the dataset page: https://huggingface.co/datasets/DrUkachi/ktt-math-tutor-data.Malay-Speech-Dataset
🎧 Malay Speech Dataset
The Malay Speech Dataset is a comprehensive speech audio dataset designed to deliver high-quality and structured audio data for AI and machine learning applications. It contains 166 hours of audio data across 533 files, provided in MP3 and WAV formats, with a total size of 126 MB. This well-balanced audio dataset offers diverse and representative voice data, with 45% female and 55% male speakers, and an age distribution ranging from 18 to 50+ years. The… See the full description on the dataset page: https://huggingface.co/datasets/Speech-data/Malay-Speech-Dataset.Macedonian-Speech-Dataset
🎧 Macedonian Speech Dataset
The Macedonian Speech Dataset is a high-quality speech audio dataset designed to provide structured and reliable audio data for AI and machine learning systems. It includes 128 hours of audio data distributed across 654 files, delivered in MP3 and WAV formats, with a total size of 279 MB. This carefully curated audio dataset ensures diverse and representative voice data, with 55% female and 45% male speakers, and an age range spanning from 18 to 50+… See the full description on the dataset page: https://huggingface.co/datasets/Speech-data/Macedonian-Speech-Dataset.luel-multilingual-tts-samples
Multilingual TTS Samples (Luel)
License: All Rights Reserved. Proprietary. Access only for authorized parties; no redistribution or use without permission. See LICENSE.
A multilingual text-to-speech / read-speech dataset of short scripted utterances across 7 languages. Each sample is a single-speaker recording of a written prompt, paired with rich speaker and recording metadata. Useful for TTS training and evaluation, ASR adaptation, dialect/accent studies, and read-speech… See the full description on the dataset page: https://huggingface.co/datasets/Luel-ai/luel-multilingual-tts-samples.modicol
MoDiCoL - A Modular Diagnostic Continual Learning Dataset for ASR
MoDiCoL is a speech dataset designed to study the robustness of ASR models to different drift factors in a controlled, continual setting. We construct MoDiCoL using a systematic factorial design that enables a rigorous evaluation of linguistic, speaker, and acoustic variation with clearly defined experimental runs. By combining real-world and synthetic speech with a configuration-dependent augmentation pipeline… See the full description on the dataset page: https://huggingface.co/datasets/TPekarekRosin/modicol.Malayalam-Speech-Dataset
🎧 Malayalam Speech Dataset
The Malayalam Speech Dataset is a high-quality speech audio dataset designed to power AI and machine learning systems with reliable and diverse audio data. It includes 95 hours of recorded speech data across 650 files, available in MP3 and WAV formats, with a total size of 227 MB. This well-structured audio dataset delivers balanced and representative voice data, featuring 54% female and 46% male speakers, with age groups ranging from 18 to 50+ years. The… See the full description on the dataset page: https://huggingface.co/datasets/Speech-data/Malayalam-Speech-Dataset.droidnexus-arabic-editorial-speech-scorecard-mini
DroidNexus Arabic Editorial Speech Scorecard Mini
A public DroidNexus Labs scorecard dataset for Arabic speech workflows: representative editorial scenarios, latency targets, overlap pressure, and the metric stack that decides whether a transcript is usable.
Why this exists
This dataset is the first public speech artifact layer for DroidNexus Labs. It publishes representative editorial workloads and evaluation pressure before claiming a full source-audio benchmark.… See the full description on the dataset page: https://huggingface.co/datasets/driodnexus/droidnexus-arabic-editorial-speech-scorecard-mini.Marathi-Speech-Dataset
Marathi Dataset Metadata
Field
Value
📜 License
CC BY-NC-ND 4.0
🎯 Task Categories
Automatic Speech Recognition
🌍 Language
Marathi (mr)
🏷️ Tags
Audio, Speech, Speech Recognition, Machine, Machine Learning, ML, Marathi
📦 Size Category
n < 1K
AYDID-public
AYDID: Arabic Yemeni Dialect Identification Dataset (public release)
This repository contains a public sample and the held-out test set of AYDID,
the first dedicated speech corpus for Yemeni Arabic at the sub-dialectal level,
supporting both automatic speech recognition (ASR) and dialect identification (DID).
Note on scope. This release contains a representative sample plus the benchmark
test set. It is intended for evaluating models against the published baselines, not
for… See the full description on the dataset page: https://huggingface.co/datasets/mansoorSaleh/AYDID-public.Processed_TTS_Multilingual_Data
Processed TTS Multilingual Data
Validated and quality-checked multilingual speech datasets for TTS training, covering 12+ Indian languages.
Datasets Included
Subset
Samples
Hours
Description
indic_voices_r
239,684
548.8h
Indic Voices_R — IVR recordings
rasa
201,509
361.2h
RASA — read speech (wiki, conv, book, news)
indictts_iitm
155,236
253.6h
Indic TTS (IIT Madras) — studio TTS recordings at 48kHz
Total
596,429
1,163.6h
Languages… See the full description on the dataset page: https://huggingface.co/datasets/PalakEngineerMaster/Processed_TTS_Multilingual_Data.ViSEC-processed
ViSEC Processed
Processed ViSEC speaker audio generated for the Meddies ASR collection.
Contents
processed_audio_by_id/: 147 WAV files named by speaker id.
metadata.csv: per-speaker metadata with duration, clip count, emotion coverage, and source-duration summary.
Schema
metadata.csv contains:
speaker_id: integer speaker identifier.
output_path: relative path to the processed WAV file.
duration_seconds: duration of the processed audio file.… See the full description on the dataset page: https://huggingface.co/datasets/Meddies/ViSEC-processed.my-audio-dataset
Audio Transcription Dataset
This dataset contains audio file paths and their corresponding transcriptions for automatic speech recognition (ASR) tasks.
Dataset Description
This dataset is structured for audio transcription tasks with two main columns:
audio: Audio file paths (type: audio)
transcript: Text transcriptions (type: text)
Files
audio_dataset.csv: Main dataset file containing audio paths and transcriptions
Dataset Structure
audio… See the full description on the dataset page: https://huggingface.co/datasets/Aashish17405/my-audio-dataset.Gleason
TalkBank CHILDES Gleason Raw
This repository mirrors the Gleason corpus from TalkBank CHILDES as raw files
for reproducible local workflows.
Source corpus page: https://talkbank.org/childes/access/Eng-NA/Gleason.html
DOI: doi:10.21415/T5101R
HF repo: MagicLuke/Gleason
Contents
transcripts/Gleason/{Mother,Father,Dinner}/*.cha
media/{Mother,Father,Dinner}/*.mp3
raw/Gleason.zip
metadata.json
metadata_from_cha.json
recordings_from_cha.csv
Data config
The… See the full description on the dataset page: https://huggingface.co/datasets/MagicLuke/Gleason.
