datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Audio2Tool
Audio2Tool: Speak, Call, Act — A Dataset for Benchmarking Speech Tool Use
Authors: Ramit Pahwa1,∗,∗∗, Apoorva Beedu1,∗, Parivesh Priye1, Rutu Gandhi†1, Saloni Takawale†1, Aruna Baijal1, Zengli Yang1
1 Rivian & Volkswagen Technologies · ∗ equal contribution · ∗∗ corresponding author · † equal contribution
📄 Project page / demo: https://audio2tool.github.io/
📦 Dataset: https://huggingface.co/datasets/RVtech/Audio2Tool
✉️ Contact (corresponding… See the full description on the dataset page: https://huggingface.co/datasets/RVtech/Audio2Tool.DeepDialogue-orpheus
DeepDialogue-orpheus
DeepDialogue-orpheus is a large-scale multimodal dataset containing 40,150 high-quality multi-turn dialogues spanning 41 domains and incorporating 20 distinct emotions with coherent emotional progressions. This repository contains the Orpheus variant of the dataset, where speech is generated using Orpheus, a state-of-the-art TTS model that infers emotional expressions implicitly from text.
🚨 Important Notice
This dataset is large (~180GB) due to… See the full description on the dataset page: https://huggingface.co/datasets/SALT-Research/DeepDialogue-orpheus.ytseg
YTSeg: A Benchmark for Audio Chaptering and Video Transcript Segmentation
We present YTSeg, a topically and structurally diverse benchmark for the audio chaptering and transcript segmentation task based on YouTube videos. The dataset comprises 19,299 videos from 393 channels, amounting to 6,533 content hours. The topics are wide-ranging, covering domains such as science, lifestyle, politics, health, economy, and technology. The videos are from various types of content formats… See the full description on the dataset page: https://huggingface.co/datasets/retkowski/ytseg.danish-asr-leaderboard
Open Danish ASR Leaderboard — Results
Benchmark results backing the Open Danish ASR Leaderboard — an open, reproducible comparison of Danish speech-to-text models.
Every model is transcribed and scored identically on the same five independent public Danish test sets, so the numbers compare directly: Word Error Rate (WER) and Character Error Rate (CER) — lower is better — plus speed. Open-weight Danish speech recognition models you can run yourself and hosted transcription APIs… See the full description on the dataset page: https://huggingface.co/datasets/RyeAI/danish-asr-leaderboard.risale-i-nur-sohbet
Risale-i Nur Sohbet
Prof. Dr. Şener Dilek’ten izin alındı.
Türkçe
Risale-i Nur sohbetlerini ses, ham ASR metni ve zaman hizalı segmentler hâlinde
birlikte sunan bağımsız bir veri kümesidir. İlk sürüm izinli ve doğrulanmış
sohbetleri içerir; kitap metni, grounded, çok dilli veya kitap seslendirme veri
kümelerine karıştırılmaz.
Kapsam
2095 sohbet, 954.66 saat 16 kHz mono FLAC ses
Aynı derslerin ölçülmüş 48 kHz kalite katmanı; 786 derste
seçici… See the full description on the dataset page: https://huggingface.co/datasets/risaleinur/risale-i-nur-sohbet.MyMentorLLM-dataset
Dataset Card for MyMentorLLM
This dataset accompanies the MyMentorLLM paper, arXiv:2607.25667, which describes the simulation environment, generation procedure, experimental design and validation analyses. If you use this dataset, please cite the paper (see Citation Information).
Dataset Summary
MyMentorLLM is a multimodal voice- and text-based deliberate-practice environment used to generate 2,100 simulated Cognitive Behavioural Therapy (CBT) training sessions… See the full description on the dataset page: https://huggingface.co/datasets/RodolfoRizzi/MyMentorLLM-dataset.TIE_shorts
Dataset Card for TIE_Shorts
Dataset Summary
TIE_shorts is a derived version of the Technical Indian English (TIE) dataset, a large-scale speech dataset (~ 8K hours) originally consisting of approximately 750 GB of content
sourced from the NPTEL platform. The original TIE dataset contains around 9.8K technical lectures in English delivered by instructors from various regions across India,
with each lecture averaging about 50 minutes. These lectures cover a wide range of… See the full description on the dataset page: https://huggingface.co/datasets/raianand/TIE_shorts.Real-TurnTurk
Real-TurnTurk
English: Real-TurnTurk is a multimodal, two-channel Turkish dyadic conversation dataset built to improve turn-taking prediction in voice-based dialogue systems. Unlike Syn-TurnTurk, the other dataset we built, every conversation here is a real, unscripted exchange between two people, recorded over video calls. Each participant was captured on a separate audio channel, so speaker attribution is exact and requires no diarization model. Alongside the audio, the… See the full description on the dataset page: https://huggingface.co/datasets/tugrulbayrak/Real-TurnTurk.risale-nur-audio
Risale-i Nur Audio–Text Corpus
Gerçek insan okumalarını, aynı satırdaki kaynak metinle birlikte sunan açık bir
ses–metin veri kümesidir. Yeni varsayılan audio-text yapılandırması 15 kitaptan
58.858 oynatılabilir klip ve 127,81 saat ses içerir. Metinler kanonik kaynaktan
değiştirilmeden alınır ve her kayıt byte-exact section_id alıntılarıyla bağlanır.
An open speech corpus pairing human readings with their source text in the same
row. The default audio-text config contains 58,858… See the full description on the dataset page: https://huggingface.co/datasets/risaleinur/risale-nur-audio.APAC-Egocentric-Residential-Voiceover
APAC Egocentric Residential (with Voiceover)
Ten narrated first-person recordings of household chores, each shipping the original capture with spoken voiceover, a burned-in caption render, WebVTT captions, and an ASS annotation track.
This is the only release in the HumynLabs egocentric collection that carries audio narration — the wearer describes each action as they perform it, and the captions align that speech to the video.
Preview: 45 s from the cooking sample, captioned… See the full description on the dataset page: https://huggingface.co/datasets/humyn-labs/APAC-Egocentric-Residential-Voiceover.DeepDialogue-xtts
DeepDialogue-xtts
DeepDialogue-xtts is a large-scale multimodal dataset containing 40,150 high-quality multi-turn dialogues spanning 41 domains and incorporating 20 distinct emotions with coherent emotional progressions.
This repository contains the XTTS-v2 variant of the dataset, where speech is generated using XTTS-v2 with explicit emotional conditioning.
🚨 Important
This dataset is large (~180GB) due to the inclusion of high-quality audio files. When cloning the… See the full description on the dataset page: https://huggingface.co/datasets/SALT-Research/DeepDialogue-xtts.fleurs-r-neucodec-all-languages
FLEURS-R NeuCodec All Languages
FLEURS-R metadata, source audio and precomputed NeuCodec speech tokens for 102
locales, plus a speaker label FLEURS itself does not ship.
Layout
data/{locale}-{split}.parquet — metadata, one row per utterance (this is what the
viewer shows).
audio/{locale}-{split}.zip — source FLEURS-R audio, 24kHz mono PCM16 WAV, members
named audio/{locale}/{split}/{id}.wav (the path column).
neucodec/{locale}-{split}-rank{N}.zip — NeuCodec… See the full description on the dataset page: https://huggingface.co/datasets/malaysia-ai/fleurs-r-neucodec-all-languages.recitation-segmentation-augmented
Automatic Pronunciation Error Detection and Correction of the Holy Quran's Learners Using Deep Learning
Paper | Project Page | Code
Introduction
This dataset is developed as part of the research presented in the paper "Automatic Pronunciation Error Detection and Correction of the Holy Quran's Learners Using Deep Learning". The work introduces a 98% automated pipeline to produce high-quality Quranic datasets, comprising over 850 hours of audio (~300K annotated utterances).… See the full description on the dataset page: https://huggingface.co/datasets/obadx/recitation-segmentation-augmented.Rasa-Annotated-25kHz
Rasa-Annotated
Enhanced version of Rasa with phoneme annotations and Kanade tokenizer features.
Dataset Statistics
Total Samples: 26,102
Total Duration: 46.92 hours
Average Duration: 6.47 seconds
Duration Range: 0.31s - 45.34s
Average Phonemes: 18.4 per sample
Average Kanade Tokens: 530.7 per sample
Global Embedding Dimension: 128
Gender Distribution
Gender
Count
Female
12,583
Male
13,519
Style Distribution
Style
Count… See the full description on the dataset page: https://huggingface.co/datasets/humair025/Rasa-Annotated-25kHz.recitation-segmentation
Recitation Segmentation Dataset for Holy Quran Pronunciation Error Detection
This dataset is used for building models that segment Holy Quran recitations based on pause points (waqf) with high accuracy. The segments are crucial for tasks like Automatic Pronunciation Error Detection and Correction, leveraging the rigorous recitation rules (tajweed) of the Holy Quran.
The dataset was presented in the paper Automatic Pronunciation Error Detection and Correction of the Holy Quran's… See the full description on the dataset page: https://huggingface.co/datasets/nour-world/recitation-segmentation.ark-asr-3b-open-asr-leaderboard-results
ARK-ASR-3B Open ASR Leaderboard Results
Raw JSONL manifests for AutoArk-AI/ARK-ASR-3B on the public English
short-form hf-audio/open-asr-leaderboard splits.
These manifests were generated on a local 8x RTX 4090 machine and scored with
the shared Open ASR Leaderboard scorer:
PYTHONPATH=. python - <<'PY'
from normalizer.eval_utils import score_results
score_results(
'ark_asr/results.AutoArk-AI-ARK-ASR-3B_20260622_official',
'AutoArk-AI/ARK-ASR-3B',
)
PY
Important:… See the full description on the dataset page: https://huggingface.co/datasets/Edge0/ark-asr-3b-open-asr-leaderboard-results.Rasa-Annotated-V1
Rasa-Annotated
Enhanced version of Rasa with phoneme annotations and Kanade tokenizer features.
Dataset Statistics
Total Samples: 26,102
Total Duration: 46.92 hours
Average Duration: 6.47 seconds
Duration Range: 0.31s - 45.34s
Average Phonemes: 18.4 per sample
Average Kanade Tokens: 264.5 per sample
Global Embedding Dimension: 128
Gender Distribution
Gender
Count
Female
12,583
Male
13,519
Style Distribution
Style
Count… See the full description on the dataset page: https://huggingface.co/datasets/humair025/Rasa-Annotated-V1.recitation-segmentation-augmented
Automatic Pronunciation Error Detection and Correction of the Holy Quran's Learners Using Deep Learning
Paper | Project Page | Code
Introduction
This dataset is developed as part of the research presented in the paper "Automatic Pronunciation Error Detection and Correction of the Holy Quran's Learners Using Deep Learning". The work introduces a 98% automated pipeline to produce high-quality Quranic datasets, comprising over 850 hours of audio (~300K annotated… See the full description on the dataset page: https://huggingface.co/datasets/nour-world/recitation-segmentation-augmented.recitation-segmentation
Recitation Segmentation Dataset for Holy Quran Pronunciation Error Detection
This dataset is used for building models that segment Holy Quran recitations based on pause points (waqf) with high accuracy. The segments are crucial for tasks like Automatic Pronunciation Error Detection and Correction, leveraging the rigorous recitation rules (tajweed) of the Holy Quran.
The dataset was presented in the paper Automatic Pronunciation Error Detection and Correction of the Holy Quran's… See the full description on the dataset page: https://huggingface.co/datasets/obadx/recitation-segmentation.egyptian-arabic-tts-diacritized
Egyptian Arabic TTS Corpus (Diacritized)
97,163 utterances / ~334 hours of Egyptian Arabic speech at 24 kHz, with
diacritized transcripts — the short vowels that Arabic script does not
write.
Why diacritics
Arabic is an abjad: short vowels are unwritten, so كتب may be kataba,
kutiba, or kutub. A TTS model with no Arabic pretraining cannot infer which,
and guesses — which native listeners hear as a foreign accent with constant
mispronunciation.
This is invisible to… See the full description on the dataset page: https://huggingface.co/datasets/Rabe3/egyptian-arabic-tts-diacritized.synthetic-speech-diarization-ru
synthetic-speech-diarization-ru
Synthetic speech diarization dataset in Parquet format.
Dataset Details
Number of tracks: 2000
Sampling rate: 16000 Hz
Audio format: Embedded in Parquet files (Audio feature compatible)
Storage: Parquet format for efficient loading
Dataset Structure
The dataset contains audio tracks with speaker diarization annotations, stored directly in Parquet format.
Features
audio: Audio waveform (Audio feature with array and… See the full description on the dataset page: https://huggingface.co/datasets/ivkond/synthetic-speech-diarization-ru.Moroccan-Arabic-Multimodal-Emotion-Recognition
MDER-MA — Moroccan Arabic Multimodal Emotion Recognition (TTS-aligned repackaging)
A repackaging of the MDER-MA dataset that pairs every audio clip with its Arabic (Moroccan dialect / Darija) transcript and ships speaker-disjoint train/validation/test splits.
Original dataset: Ouali, S. & El Garouani, S. (2025). MDER-MA: A multimodal dataset for emotion recognition in low-resource Moroccan Arabic language. Data in Brief. DOI: 10.1016/j.dib.2025.112005. Mendeley:… See the full description on the dataset page: https://huggingface.co/datasets/FatimahEmadEldin/Moroccan-Arabic-Multimodal-Emotion-Recognition.biggest_ru_book_balalaika
Biggest-Ru-Book Annotated by Balalaika
[!IMPORTANT]
Official dataset for our INTERSPEECH 2026 paper
"A Data-Centric Framework for Addressing Phonetic and Prosodic Challenges in Russian Speech Generative Models" (arXiv:2507.13563).
Part of the Balalaika Russian speech data-processing pipeline — code: https://github.com/lab260ru/balalaika.
If you use this resource, please cite it.
A curated Russian speech dataset for advanced speech generative tasks.
Overview… See the full description on the dataset page: https://huggingface.co/datasets/lab260/biggest_ru_book_balalaika.paralingua_ru
Russian Paralinguistic Annotation Dataset
Датасет паралингвистической разметки спикеров из трёх русскоязычных корпусов:
biggest_ru_book,
DeepSpeech и Golos.
Что размечалось
Каждое аудио размечалось вручную по следующим характеристикам:
Поле
Описание
Пример значений
gender
Пол спикера
мужской, женский
age_group
Возрастная группа
молодой, взрослый, пожилой
voice_pitch
Высота голоса
низкий, средний, высокий
loudness
Громкость
тихий, нормальный… See the full description on the dataset page: https://huggingface.co/datasets/turnipseason/paralingua_ru.synthetic-speech-diarization-ru
synthetic-speech-diarization-ru
Synthetic speech diarization dataset in Parquet format.
Dataset Details
Number of tracks: 2000
Sampling rate: 16000 Hz
Audio format: Embedded in Parquet files (Audio feature compatible)
Storage: Parquet format for efficient loading
Dataset Structure
The dataset contains audio tracks with speaker diarization annotations, stored directly in Parquet format.
Features
audio: Audio waveform (Audio feature with array and… See the full description on the dataset page: https://huggingface.co/datasets/niobures/synthetic-speech-diarization-ru.persian-asr-text-2.69M-deduped
🗂️ persian-asr-text-2.69M-deduped
English + فارسی · Part of Shenava 1.0 · Project hub · SLT paper submission
🌟 At a glance | معرفی سریع
English
فارسی
🎯 Purpose
Deduplicated Persian ASR text dataset used by the training stack.
پیکرهٔ متنی فارسیِ حذفتکرارشده برای ساخت واژگان، مدلسازی زبانی و پشتیبانی از آموزش ASR.
🧩 Role
Persian text and linguistic asset
مصنوع متنی و زبانی فارسی
📦 Snapshot
4 files; approximately 109.64 MB
4 فایل؛ حدود 109.64… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/persian-asr-text-2.69M-deduped.BLURB-synth
BLURB-synth: Synthetic audio data based on BLURB corpora
Dataset Summary
Synthetic audio data based on BLURB corpora. More details coming soon...
Supported Tasks and Leaderboards
Biomedical Language Understanding and Reasoning Benchmark (BLURB)
Text-to-Speech
Automatic-Speech-Recognition
Languages
English
Data Structure
Data Instances
Coming soon...
Data Fields
Coming soon...… See the full description on the dataset page: https://huggingface.co/datasets/uy-rrodriguez/BLURB-synth.reazonspeech-v2-quality-index
ReazonSpeech v2 Quality Index
Quality metadata for 21,932,215 ReazonSpeech v2 utterances.
It joins the following two source analyses by exact audio path:
ayousanz/reazon-speech-v2-all-speechMOS-analyze/audio_analysis_results_speechMOS.json
ayousanz/reazon-speech-v2-all-WAND-SNR-analyze/reazonspeech-all-wada-snr.json
The source repositories are not modified and this repository does not contain
the source audio.
Validation
Check
Count
SpeechMOS rows
21… See the full description on the dataset page: https://huggingface.co/datasets/ayousanz/reazonspeech-v2-quality-index.reviewed_sample
Silencio Network: Multilingual Accent Speech Dataset (Sample)
Overview
Silencio data is valuable because it’s collected in the wild from a massive, opt-in community (1.2M users across 180+ countries), giving buyers real-world accents, dialects, devices, and environments that lab or scraped datasets don’t capture. Every recording is tied to explicit, traceable consent and processed with privacy-first pipelines (GDPR/CCPA compliant, anonymized, PII hashed), which… See the full description on the dataset page: https://huggingface.co/datasets/jml2026/reviewed_sample.fleurs-flac
FLEURS-FLAC
A losslessly FLAC-compressed version of Google's FLEURS dataset covering 102 languages.
Overview
This repository contains the Google FLEURS dataset repackaged into Parquet shards with PCM24 FLAC-compressed audio binaries.
Key points:
Audio streams are converted to FLAC (PCM24) with sample-level PCM verification against the source.
Sharded into ~500MB Parquet files per split for efficient I/O and streaming.
Covers all 102 languages from the original… See the full description on the dataset page: https://huggingface.co/datasets/roro128/fleurs-flac.
