CoolFace
15 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01wayu-ai /thai-aligner-bench Thai Aligner Bench 🚧 Development in progress. How accurately can a forced aligner place Thai token and word boundaries in speech? This is a self-contained benchmark: one Python file (aligner_bench.py) plus 1,572 clips of Thai speech with frame-exact timing ground truth. No Thai NLP stack or other code is needed — just numpy soundfile torch torchaudio transformers. The ground truth is what makes the dataset useful: the audio was rendered by a TTS model whose duration predictor… See the full description on the dataset page: https://huggingface.co/datasets/wayu-ai/thai-aligner-bench.audioautomatic-speech-recognition1K<n<10K1 likes805 downloads1mo agoHugging Face02besimple-ai /voice-code-bench VoiceCodeBench VoiceCodeBench is a test-only benchmark for evaluating whether automatic speech recognition (ASR) systems preserve exact structured values in English workplace speech. Paper: VoiceCodeBench: Evaluating Exact Structured-Token Recovery in Automatic Speech Recognition The benchmark targets cases where a transcript is software input: callback numbers, email addresses, command-line flags, file paths, URLs, account identifiers, dates, measurements, and similar values… See the full description on the dataset page: https://huggingface.co/datasets/besimple-ai/voice-code-bench.audioautomatic-speech-recognitionn<1K14 likes744 downloads12d agoHugging Face03akshaya-AI /kimiDatasetaudio1K<n<10K1 likes388 downloads2mo agoHugging Face04besimple-ai /vocal-affect-bench VocalAffectBench VocalAffectBench is a test-only benchmark for evaluating whether AI audio models can identify expressed vocal emotion from raw audio. Paper: VocalAffectBench: Evaluating Vocal Emotion Recognition in AI Audio Models The benchmark targets the expressed emotion — what the speaker conveys through vocal tone, prosody, pace, intensity, and pauses — not inferred internal state. Contents 280 human-recorded English WAV clips, totalling 2.32 hours. 7… See the full description on the dataset page: https://huggingface.co/datasets/besimple-ai/vocal-affect-bench.audioaudio-classificationn<1K7 likes258 downloads15d agoHugging Face05Archit00 /aijockey-public-corpusaudion<1K0 likes200 downloads4mo agoHugging Face06uzinfocom-edu-ai /uzbek-asr-curated-701h Uzbek ASR Curated Dataset (701 hours) A curated multi-source Uzbek speech dataset for automatic speech recognition (ASR) training and evaluation. Dataset Description Language Uzbek (Latin script with okina ʻ) Total utterances 337,920 Total duration ~701 hours Audio format 16 kHz mono WAV (PCM_16) Manifest format NeMo JSONL Splits train (94%) / val (3%) / test (3%) Splits Split Utterances Hours Train 317,655… See the full description on the dataset page: https://huggingface.co/datasets/uzinfocom-edu-ai/uzbek-asr-curated-701h.audioautomatic-speech-recognition100K<n<1M1 likes111 downloads3mo agoHugging Face07ai-ssam /darija-tts-8400 Darija TTS 8400 Synthetic Moroccan Darija speech for TTS fine-tuning: 8,400 single-speaker clips (20.73 hours), 24 kHz mono PCM16 WAV. All audio is generated with Gemini 3.1 Flash TTS (gemini-3.1-flash-tts-preview, voice Kore). Clips are unreviewed; there are no human recordings. Write-up of how this data was used: Training a Voice. At a glance Clips / hours 8,400 / 20.73 Unique texts 4,800 Voice Kore (1 speaker) Sample rate 24 kHz mono PCM16… See the full description on the dataset page: https://huggingface.co/datasets/ai-ssam/darija-tts-8400.audiotext-to-speech1K<n<10K0 likes85 downloads7d agoHugging Face08playwithmino /aishell1mix-ver2-n100-per-mix AISHELL-1 Mix ver2 — 100 clips per mix This is AISHELL-1 Mix ver2, not ver1. 8 kHz mono test subset: 100 mixtures per speaker count (N=1\ldots5) (50 mix_clean + 50 mix_both each) → 500 clips. Derived from the local aishell1mix_ver2 test SCPs (data/scp/scp_aishell1mix_ver2). Includes mixture + oracle speaker stems and transcripts. Split Count 1mix / 2mix / 3mix / 4mix / 5mix 100 each clean / both 250 each Files manifests/test.jsonl —… See the full description on the dataset page: https://huggingface.co/datasets/playwithmino/aishell1mix-ver2-n100-per-mix.audioaudio-to-audion<1K0 likes65 downloads22d agoHugging Face09danielrosehill /multimodal-ai-taxonomy Multimodal AI Taxonomy A comprehensive, structured taxonomy for mapping multimodal AI model capabilities across input and output modalities. Dataset Description This dataset provides a systematic categorization of multimodal AI capabilities, enabling users to: Navigate the complex landscape of multimodal AI models Filter models by specific input/output modality combinations Understand the nuanced differences between similar models (e.g., image-to-video with/without audio… See the full description on the dataset page: https://huggingface.co/datasets/danielrosehill/multimodal-ai-taxonomy.textothern<1K0 likes49 downloads11mo agoHugging Face10krutrim-ai-labs /IndicSTgated IndicST: Indian Multilingual Translation Corpus For Evaluating Speech Large Language Models Introduction IndicST, a new dataset tailored for training and evaluating Speech LLMs for AST tasks (including ASR and TTS), featuring meticulously curated, automatically, and manually verified synthetic data. The dataset offers 10.8k hrs of training data and 1.13k hrs of evaluation data. Use-Cases ASR (Speech-to-Text) Transcribing Indic languages Handling… See the full description on the dataset page: https://huggingface.co/datasets/krutrim-ai-labs/IndicST.audio10M<n<100M8 likes47 downloads2y agoHugging Face11akshaya-AI /kimiaudio1K<n<10K0 likes46 downloads2mo agoHugging Face12sleeping-ai /PlayAI-VoiceExcited to share Play AI Voice Profile. We release 267 unique voice profiles including Israeli, Arabic, Russian, Filipino and many other exclusive voice profiles. Play AI was recently acquired by Meta which sparked our interest in releasing this dataset. tabularn<1K0 likes18 downloads1y agoHugging Face13Ailiance-fr /mascarade-dsp-dataset Mascarade — DSP & Signal Processing Q&A ✅ ATTRIBUTION AUDIT COMPLETED (2026-05-11) Per-sample Stack Exchange Electronics attribution recovered via the SE /search/advanced + /questions/{id} API search : 169 samples (~5.35 %) confirmed as Stack Exchange Electronics (CC-BY-SA-4.0) — fully attributed in metadata.stack_exchange_attribution (URL + author display name + author user_id + post_id + creation_date_unix + match_confidence ≥ 0.60). 535 samples (~16.93 %) marked… See the full description on the dataset page: https://huggingface.co/datasets/Ailiance-fr/mascarade-dsp-dataset.texttext-generation1K<n<10K0 likes12 downloads5mo agoHugging Face14ThisUsernameAlreadyExistsAlreadyExists /aitf-dfk3-synthetic-audio-datasetgatedaudioaudio-classification1K<n<10K0 likes5 downloads6mo agoHugging Face15saidalihon /world-ai-dbaudion<1K0 likes5 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.