CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01RVtech /Audio2Tool Audio2Tool: Speak, Call, Act — A Dataset for Benchmarking Speech Tool Use Authors: Ramit Pahwa1,∗,∗∗, Apoorva Beedu1,∗, Parivesh Priye1, Rutu Gandhi†1, Saloni Takawale†1, Aruna Baijal1, Zengli Yang1 1 Rivian & Volkswagen Technologies &nbsp;·&nbsp; ∗ equal contribution &nbsp;·&nbsp; ∗∗ corresponding author &nbsp;·&nbsp; † equal contribution 📄 Project page / demo: https://audio2tool.github.io/ 📦 Dataset: https://huggingface.co/datasets/RVtech/Audio2Tool ✉️ Contact (corresponding… See the full description on the dataset page: https://huggingface.co/datasets/RVtech/Audio2Tool.audioautomatic-speech-recognition10K<n<100K2 likes6k downloads3mo agoHugging Face02SALT-Research /DeepDialogue-orpheus DeepDialogue-orpheus DeepDialogue-orpheus is a large-scale multimodal dataset containing 40,150 high-quality multi-turn dialogues spanning 41 domains and incorporating 20 distinct emotions with coherent emotional progressions. This repository contains the Orpheus variant of the dataset, where speech is generated using Orpheus, a state-of-the-art TTS model that infers emotional expressions implicitly from text. 🚨 Important Notice This dataset is large (~180GB) due to… See the full description on the dataset page: https://huggingface.co/datasets/SALT-Research/DeepDialogue-orpheus.audioaudio-classification100K<n<1M8 likes4.3k downloads1y agoHugging Face03retkowski /ytseg YTSeg: A Benchmark for Audio Chaptering and Video Transcript Segmentation We present YTSeg, a topically and structurally diverse benchmark for the audio chaptering and transcript segmentation task based on YouTube videos. The dataset comprises 19,299 videos from 393 channels, amounting to 6,533 content hours. The topics are wide-ranging, covering domains such as science, lifestyle, politics, health, economy, and technology. The videos are from various types of content formats… See the full description on the dataset page: https://huggingface.co/datasets/retkowski/ytseg.audiotoken-classification100K<n<1M9 likes2.7k downloads2mo agoHugging Face04RyeAI /danish-asr-leaderboard Open Danish ASR Leaderboard — Results Benchmark results backing the Open Danish ASR Leaderboard — an open, reproducible comparison of Danish speech-to-text models. Every model is transcribed and scored identically on the same five independent public Danish test sets, so the numbers compare directly: Word Error Rate (WER) and Character Error Rate (CER) — lower is better — plus speed. Open-weight Danish speech recognition models you can run yourself and hosted transcription APIs… See the full description on the dataset page: https://huggingface.co/datasets/RyeAI/danish-asr-leaderboard.tabularautomatic-speech-recognition1M<n<10M4 likes1.8k downloads10h agoHugging Face05risaleinur /risale-i-nur-sohbet Risale-i Nur Sohbet Prof. Dr. Şener Dilek’ten izin alındı. Türkçe Risale-i Nur sohbetlerini ses, ham ASR metni ve zaman hizalı segmentler hâlinde birlikte sunan bağımsız bir veri kümesidir. İlk sürüm izinli ve doğrulanmış sohbetleri içerir; kitap metni, grounded, çok dilli veya kitap seslendirme veri kümelerine karıştırılmaz. Kapsam 2095 sohbet, 954.66 saat 16 kHz mono FLAC ses Aynı derslerin ölçülmüş 48 kHz kalite katmanı; 786 derste seçici… See the full description on the dataset page: https://huggingface.co/datasets/risaleinur/risale-i-nur-sohbet.audioautomatic-speech-recognition1M<n<10M1 likes1.2k downloads20d agoHugging Face06RodolfoRizzi /MyMentorLLM-dataset Dataset Card for MyMentorLLM This dataset accompanies the MyMentorLLM paper, arXiv:2607.25667, which describes the simulation environment, generation procedure, experimental design and validation analyses. If you use this dataset, please cite the paper (see Citation Information). Dataset Summary MyMentorLLM is a multimodal voice- and text-based deliberate-practice environment used to generate 2,100 simulated Cognitive Behavioural Therapy (CBT) training sessions… See the full description on the dataset page: https://huggingface.co/datasets/RodolfoRizzi/MyMentorLLM-dataset.audiotext-generation10K<n<100K0 likes955 downloads5d agoHugging Face07raianand /TIE_shorts Dataset Card for TIE_Shorts Dataset Summary TIE_shorts is a derived version of the Technical Indian English (TIE) dataset, a large-scale speech dataset (~ 8K hours) originally consisting of approximately 750 GB of content sourced from the NPTEL platform. The original TIE dataset contains around 9.8K technical lectures in English delivered by instructors from various regions across India, with each lecture averaging about 50 minutes. These lectures cover a wide range of… See the full description on the dataset page: https://huggingface.co/datasets/raianand/TIE_shorts.audioautomatic-speech-recognition1K<n<10K1 likes864 downloads2y agoHugging Face08tugrulbayrak /Real-TurnTurk Real-TurnTurk English: Real-TurnTurk is a multimodal, two-channel Turkish dyadic conversation dataset built to improve turn-taking prediction in voice-based dialogue systems. Unlike Syn-TurnTurk, the other dataset we built, every conversation here is a real, unscripted exchange between two people, recorded over video calls. Each participant was captured on a separate audio channel, so speaker attribution is exact and requires no diarization model. Alongside the audio, the… See the full description on the dataset page: https://huggingface.co/datasets/tugrulbayrak/Real-TurnTurk.tabularaudio-classification100K<n<1M3 likes533 downloads5d agoHugging Face09risaleinur /risale-nur-audio Risale-i Nur Audio–Text Corpus Gerçek insan okumalarını, aynı satırdaki kaynak metinle birlikte sunan açık bir ses–metin veri kümesidir. Yeni varsayılan audio-text yapılandırması 15 kitaptan 58.858 oynatılabilir klip ve 127,81 saat ses içerir. Metinler kanonik kaynaktan değiştirilmeden alınır ve her kayıt byte-exact section_id alıntılarıyla bağlanır. An open speech corpus pairing human readings with their source text in the same row. The default audio-text config contains 58,858… See the full description on the dataset page: https://huggingface.co/datasets/risaleinur/risale-nur-audio.audioautomatic-speech-recognition100K<n<1M1 likes258 downloads4h agoHugging Face10humyn-labs /APAC-Egocentric-Residential-Voiceover APAC Egocentric Residential (with Voiceover) Ten narrated first-person recordings of household chores, each shipping the original capture with spoken voiceover, a burned-in caption render, WebVTT captions, and an ASS annotation track. This is the only release in the HumynLabs egocentric collection that carries audio narration — the wearer describes each action as they perform it, and the captions align that speech to the video. Preview: 45 s from the cooking sample, captioned… See the full description on the dataset page: https://huggingface.co/datasets/humyn-labs/APAC-Egocentric-Residential-Voiceover.tabularroboticsn<1K0 likes249 downloads2mo agoHugging Face11SALT-Research /DeepDialogue-xtts DeepDialogue-xtts DeepDialogue-xtts is a large-scale multimodal dataset containing 40,150 high-quality multi-turn dialogues spanning 41 domains and incorporating 20 distinct emotions with coherent emotional progressions. This repository contains the XTTS-v2 variant of the dataset, where speech is generated using XTTS-v2 with explicit emotional conditioning. 🚨 Important This dataset is large (~180GB) due to the inclusion of high-quality audio files. When cloning the… See the full description on the dataset page: https://huggingface.co/datasets/SALT-Research/DeepDialogue-xtts.audioaudio-classification100K<n<1M8 likes240 downloads1y agoHugging Face12malaysia-ai /fleurs-r-neucodec-all-languages FLEURS-R NeuCodec All Languages FLEURS-R metadata, source audio and precomputed NeuCodec speech tokens for 102 locales, plus a speaker label FLEURS itself does not ship. Layout data/{locale}-{split}.parquet — metadata, one row per utterance (this is what the viewer shows). audio/{locale}-{split}.zip — source FLEURS-R audio, 24kHz mono PCM16 WAV, members named audio/{locale}/{split}/{id}.wav (the path column). neucodec/{locale}-{split}-rank{N}.zip — NeuCodec… See the full description on the dataset page: https://huggingface.co/datasets/malaysia-ai/fleurs-r-neucodec-all-languages.audiotext-to-speech100K<n<1M4 likes239 downloads12d agoHugging Face13obadx /recitation-segmentation-augmented Automatic Pronunciation Error Detection and Correction of the Holy Quran's Learners Using Deep Learning Paper | Project Page | Code Introduction This dataset is developed as part of the research presented in the paper "Automatic Pronunciation Error Detection and Correction of the Holy Quran's Learners Using Deep Learning". The work introduces a 98% automated pipeline to produce high-quality Quranic datasets, comprising over 850 hours of audio (~300K annotated utterances).… See the full description on the dataset page: https://huggingface.co/datasets/obadx/recitation-segmentation-augmented.tabularautomatic-speech-recognition10K<n<100K0 likes234 downloads1y agoHugging Face14humair025 /Rasa-Annotated-25kHz Rasa-Annotated Enhanced version of Rasa with phoneme annotations and Kanade tokenizer features. Dataset Statistics Total Samples: 26,102 Total Duration: 46.92 hours Average Duration: 6.47 seconds Duration Range: 0.31s - 45.34s Average Phonemes: 18.4 per sample Average Kanade Tokens: 530.7 per sample Global Embedding Dimension: 128 Gender Distribution Gender Count Female 12,583 Male 13,519 Style Distribution Style Count… See the full description on the dataset page: https://huggingface.co/datasets/humair025/Rasa-Annotated-25kHz.tabulartext-to-speech10K<n<100K0 likes191 downloads8mo agoHugging Face15nour-world /recitation-segmentation Recitation Segmentation Dataset for Holy Quran Pronunciation Error Detection This dataset is used for building models that segment Holy Quran recitations based on pause points (waqf) with high accuracy. The segments are crucial for tasks like Automatic Pronunciation Error Detection and Correction, leveraging the rigorous recitation rules (tajweed) of the Holy Quran. The dataset was presented in the paper Automatic Pronunciation Error Detection and Correction of the Holy Quran's… See the full description on the dataset page: https://huggingface.co/datasets/nour-world/recitation-segmentation.tabularautomatic-speech-recognition10K<n<100K0 likes190 downloads10d agoHugging Face16Edge0 /ark-asr-3b-open-asr-leaderboard-results ARK-ASR-3B Open ASR Leaderboard Results Raw JSONL manifests for AutoArk-AI/ARK-ASR-3B on the public English short-form hf-audio/open-asr-leaderboard splits. These manifests were generated on a local 8x RTX 4090 machine and scored with the shared Open ASR Leaderboard scorer: PYTHONPATH=. python - <<'PY' from normalizer.eval_utils import score_results score_results( 'ark_asr/results.AutoArk-AI-ARK-ASR-3B_20260622_official', 'AutoArk-AI/ARK-ASR-3B', ) PY Important:… See the full description on the dataset page: https://huggingface.co/datasets/Edge0/ark-asr-3b-open-asr-leaderboard-results.tabularautomatic-speech-recognition10K<n<100K12 likes185 downloads3mo agoHugging Face17humair025 /Rasa-Annotated-V1 Rasa-Annotated Enhanced version of Rasa with phoneme annotations and Kanade tokenizer features. Dataset Statistics Total Samples: 26,102 Total Duration: 46.92 hours Average Duration: 6.47 seconds Duration Range: 0.31s - 45.34s Average Phonemes: 18.4 per sample Average Kanade Tokens: 264.5 per sample Global Embedding Dimension: 128 Gender Distribution Gender Count Female 12,583 Male 13,519 Style Distribution Style Count… See the full description on the dataset page: https://huggingface.co/datasets/humair025/Rasa-Annotated-V1.tabulartext-to-speech10K<n<100K0 likes168 downloads8mo agoHugging Face18nour-world /recitation-segmentation-augmented Automatic Pronunciation Error Detection and Correction of the Holy Quran's Learners Using Deep Learning Paper | Project Page | Code Introduction This dataset is developed as part of the research presented in the paper "Automatic Pronunciation Error Detection and Correction of the Holy Quran's Learners Using Deep Learning". The work introduces a 98% automated pipeline to produce high-quality Quranic datasets, comprising over 850 hours of audio (~300K annotated… See the full description on the dataset page: https://huggingface.co/datasets/nour-world/recitation-segmentation-augmented.tabularautomatic-speech-recognition10K<n<100K0 likes168 downloads10d agoHugging Face19obadx /recitation-segmentation Recitation Segmentation Dataset for Holy Quran Pronunciation Error Detection This dataset is used for building models that segment Holy Quran recitations based on pause points (waqf) with high accuracy. The segments are crucial for tasks like Automatic Pronunciation Error Detection and Correction, leveraging the rigorous recitation rules (tajweed) of the Holy Quran. The dataset was presented in the paper Automatic Pronunciation Error Detection and Correction of the Holy Quran's… See the full description on the dataset page: https://huggingface.co/datasets/obadx/recitation-segmentation.tabularautomatic-speech-recognition10K<n<100K1 likes162 downloads1y agoHugging Face20Rabe3 /egyptian-arabic-tts-diacritized Egyptian Arabic TTS Corpus (Diacritized) 97,163 utterances / ~334 hours of Egyptian Arabic speech at 24 kHz, with diacritized transcripts — the short vowels that Arabic script does not write. Why diacritics Arabic is an abjad: short vowels are unwritten, so كتب may be kataba, kutiba, or kutub. A TTS model with no Arabic pretraining cannot infer which, and guesses — which native listeners hear as a foreign accent with constant mispronunciation. This is invisible to… See the full description on the dataset page: https://huggingface.co/datasets/Rabe3/egyptian-arabic-tts-diacritized.tabulartext-to-speech10K<n<100K0 likes142 downloads1mo agoHugging Face21ivkond /synthetic-speech-diarization-ru synthetic-speech-diarization-ru Synthetic speech diarization dataset in Parquet format. Dataset Details Number of tracks: 2000 Sampling rate: 16000 Hz Audio format: Embedded in Parquet files (Audio feature compatible) Storage: Parquet format for efficient loading Dataset Structure The dataset contains audio tracks with speaker diarization annotations, stored directly in Parquet format. Features audio: Audio waveform (Audio feature with array and… See the full description on the dataset page: https://huggingface.co/datasets/ivkond/synthetic-speech-diarization-ru.tabularautomatic-speech-recognition1K<n<10K0 likes98 downloads10mo agoHugging Face22FatimahEmadEldin /Moroccan-Arabic-Multimodal-Emotion-Recognition MDER-MA — Moroccan Arabic Multimodal Emotion Recognition (TTS-aligned repackaging) A repackaging of the MDER-MA dataset that pairs every audio clip with its Arabic (Moroccan dialect / Darija) transcript and ships speaker-disjoint train/validation/test splits. Original dataset: Ouali, S. & El Garouani, S. (2025). MDER-MA: A multimodal dataset for emotion recognition in low-resource Moroccan Arabic language. Data in Brief. DOI: 10.1016/j.dib.2025.112005. Mendeley:… See the full description on the dataset page: https://huggingface.co/datasets/FatimahEmadEldin/Moroccan-Arabic-Multimodal-Emotion-Recognition.audiotext-to-speech1K<n<10K1 likes84 downloads5mo agoHugging Face23lab260 /biggest_ru_book_balalaika Biggest-Ru-Book Annotated by Balalaika [!IMPORTANT] Official dataset for our INTERSPEECH 2026 paper "A Data-Centric Framework for Addressing Phonetic and Prosodic Challenges in Russian Speech Generative Models" (arXiv:2507.13563). Part of the Balalaika Russian speech data-processing pipeline — code: https://github.com/lab260ru/balalaika. If you use this resource, please cite it. A curated Russian speech dataset for advanced speech generative tasks. Overview… See the full description on the dataset page: https://huggingface.co/datasets/lab260/biggest_ru_book_balalaika.tabulartext-to-speech100K<n<1M3 likes81 downloads3mo agoHugging Face24turnipseason /paralingua_ru Russian Paralinguistic Annotation Dataset Датасет паралингвистической разметки спикеров из трёх русскоязычных корпусов: biggest_ru_book, DeepSpeech и Golos. Что размечалось Каждое аудио размечалось вручную по следующим характеристикам: Поле Описание Пример значений gender Пол спикера мужской, женский age_group Возрастная группа молодой, взрослый, пожилой voice_pitch Высота голоса низкий, средний, высокий loudness Громкость тихий, нормальный… See the full description on the dataset page: https://huggingface.co/datasets/turnipseason/paralingua_ru.tabulartext-to-speech100K<n<1M8 likes76 downloads4mo agoHugging Face25niobures /synthetic-speech-diarization-ru synthetic-speech-diarization-ru Synthetic speech diarization dataset in Parquet format. Dataset Details Number of tracks: 2000 Sampling rate: 16000 Hz Audio format: Embedded in Parquet files (Audio feature compatible) Storage: Parquet format for efficient loading Dataset Structure The dataset contains audio tracks with speaker diarization annotations, stored directly in Parquet format. Features audio: Audio waveform (Audio feature with array and… See the full description on the dataset page: https://huggingface.co/datasets/niobures/synthetic-speech-diarization-ru.tabularautomatic-speech-recognition1K<n<10K0 likes75 downloads5mo agoHugging Face26Reza2kn /persian-asr-text-2.69M-deduped 🗂️ persian-asr-text-2.69M-deduped English + فارسی · Part of Shenava 1.0 · Project hub · SLT paper submission 🌟 At a glance | معرفی سریع English فارسی 🎯 Purpose Deduplicated Persian ASR text dataset used by the training stack. پیکرهٔ متنی فارسیِ حذف‌تکرارشده برای ساخت واژگان، مدل‌سازی زبانی و پشتیبانی از آموزش ASR. 🧩 Role Persian text and linguistic asset مصنوع متنی و زبانی فارسی 📦 Snapshot 4 files; approximately 109.64 MB 4 فایل؛ حدود 109.64… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/persian-asr-text-2.69M-deduped.tabularautomatic-speech-recognition1M<n<10M0 likes67 downloads2mo agoHugging Face27uy-rrodriguez /BLURB-synthgated BLURB-synth: Synthetic audio data based on BLURB corpora Dataset Summary Synthetic audio data based on BLURB corpora. More details coming soon... Supported Tasks and Leaderboards Biomedical Language Understanding and Reasoning Benchmark (BLURB) Text-to-Speech Automatic-Speech-Recognition Languages English Data Structure Data Instances Coming soon... Data Fields Coming soon...… See the full description on the dataset page: https://huggingface.co/datasets/uy-rrodriguez/BLURB-synth.tabulartext-to-speech1M<n<10M0 likes66 downloads2d agoHugging Face28ayousanz /reazonspeech-v2-quality-index ReazonSpeech v2 Quality Index Quality metadata for 21,932,215 ReazonSpeech v2 utterances. It joins the following two source analyses by exact audio path: ayousanz/reazon-speech-v2-all-speechMOS-analyze/audio_analysis_results_speechMOS.json ayousanz/reazon-speech-v2-all-WAND-SNR-analyze/reazonspeech-all-wada-snr.json The source repositories are not modified and this repository does not contain the source audio. Validation Check Count SpeechMOS rows 21… See the full description on the dataset page: https://huggingface.co/datasets/ayousanz/reazonspeech-v2-quality-index.tabularautomatic-speech-recognition10M<n<100M0 likes66 downloads4d agoHugging Face29jml2026 /reviewed_sample Silencio Network: Multilingual Accent Speech Dataset (Sample) Overview Silencio data is valuable because it’s collected in the wild from a massive, opt-in community (1.2M users across 180+ countries), giving buyers real-world accents, dialects, devices, and environments that lab or scraped datasets don’t capture. Every recording is tied to explicit, traceable consent and processed with privacy-first pipelines (GDPR/CCPA compliant, anonymized, PII hashed), which… See the full description on the dataset page: https://huggingface.co/datasets/jml2026/reviewed_sample.audioautomatic-speech-recognition1K<n<10K0 likes57 downloads8mo agoHugging Face30roro128 /fleurs-flac FLEURS-FLAC A losslessly FLAC-compressed version of Google's FLEURS dataset covering 102 languages. Overview This repository contains the Google FLEURS dataset repackaged into Parquet shards with PCM24 FLAC-compressed audio binaries. Key points: Audio streams are converted to FLAC (PCM24) with sample-level PCM verification against the source. Sharded into ~500MB Parquet files per split for efficient I/O and streaming. Covers all 102 languages from the original… See the full description on the dataset page: https://huggingface.co/datasets/roro128/fleurs-flac.tabularautomatic-speech-recognition100K<n<1M0 likes52 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.