CoolFace
17 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01wayu-ai /thai-aligner-bench Thai Aligner Bench 🚧 Development in progress. How accurately can a forced aligner place Thai token and word boundaries in speech? This is a self-contained benchmark: one Python file (aligner_bench.py) plus 1,572 clips of Thai speech with frame-exact timing ground truth. No Thai NLP stack or other code is needed — just numpy soundfile torch torchaudio transformers. The ground truth is what makes the dataset useful: the audio was rendered by a TTS model whose duration predictor… See the full description on the dataset page: https://huggingface.co/datasets/wayu-ai/thai-aligner-bench.audioautomatic-speech-recognition1K<n<10K1 likes802 downloads1mo agoHugging Face02besimple-ai /voice-code-bench VoiceCodeBench VoiceCodeBench is a test-only benchmark for evaluating whether automatic speech recognition (ASR) systems preserve exact structured values in English workplace speech. Paper: VoiceCodeBench: Evaluating Exact Structured-Token Recovery in Automatic Speech Recognition The benchmark targets cases where a transcript is software input: callback numbers, email addresses, command-line flags, file paths, URLs, account identifiers, dates, measurements, and similar values… See the full description on the dataset page: https://huggingface.co/datasets/besimple-ai/voice-code-bench.audioautomatic-speech-recognitionn<1K14 likes728 downloads14d agoHugging Face03akshaya-AI /kimiDatasetaudio1K<n<10K1 likes313 downloads2mo agoHugging Face04besimple-ai /vocal-affect-bench VocalAffectBench VocalAffectBench is a test-only benchmark for evaluating whether AI audio models can identify expressed vocal emotion from raw audio. Paper: VocalAffectBench: Evaluating Vocal Emotion Recognition in AI Audio Models The benchmark targets the expressed emotion — what the speaker conveys through vocal tone, prosody, pace, intensity, and pauses — not inferred internal state. Contents 280 human-recorded English WAV clips, totalling 2.32 hours. 7… See the full description on the dataset page: https://huggingface.co/datasets/besimple-ai/vocal-affect-bench.audioaudio-classificationn<1K7 likes288 downloads2d agoHugging Face05Archit00 /aijockey-public-corpusaudion<1K0 likes177 downloads4mo agoHugging Face06ai-ssam /darija-tts-8400 Darija TTS 8400 Synthetic Moroccan Darija speech for TTS fine-tuning: 8,400 single-speaker clips (20.73 hours), 24 kHz mono PCM16 WAV. All audio is generated with Gemini 3.1 Flash TTS (gemini-3.1-flash-tts-preview, voice Kore). Clips are unreviewed; there are no human recordings. Write-up of how this data was used: Training a Voice. At a glance Clips / hours 8,400 / 20.73 Unique texts 4,800 Voice Kore (1 speaker) Sample rate 24 kHz mono PCM16… See the full description on the dataset page: https://huggingface.co/datasets/ai-ssam/darija-tts-8400.audiotext-to-speech1K<n<10K0 likes122 downloads10d agoHugging Face07uzinfocom-edu-ai /uzbek-asr-curated-701h Uzbek ASR Curated Dataset (701 hours) A curated multi-source Uzbek speech dataset for automatic speech recognition (ASR) training and evaluation. Dataset Description Language Uzbek (Latin script with okina ʻ) Total utterances 337,920 Total duration ~701 hours Audio format 16 kHz mono WAV (PCM_16) Manifest format NeMo JSONL Splits train (94%) / val (3%) / test (3%) Splits Split Utterances Hours Train 317,655… See the full description on the dataset page: https://huggingface.co/datasets/uzinfocom-edu-ai/uzbek-asr-curated-701h.audioautomatic-speech-recognition100K<n<1M1 likes113 downloads3mo agoHugging Face08playwithmino /aishell1mix-ver2-n100-per-mix AISHELL-1 Mix ver2 — 100 clips per mix This is AISHELL-1 Mix ver2, not ver1. 8 kHz mono test subset: 100 mixtures per speaker count (N=1\ldots5) (50 mix_clean + 50 mix_both each) → 500 clips. Derived from the local aishell1mix_ver2 test SCPs (data/scp/scp_aishell1mix_ver2). Includes mixture + oracle speaker stems and transcripts. Split Count 1mix / 2mix / 3mix / 4mix / 5mix 100 each clean / both 250 each Files manifests/test.jsonl —… See the full description on the dataset page: https://huggingface.co/datasets/playwithmino/aishell1mix-ver2-n100-per-mix.audioaudio-to-audion<1K0 likes68 downloads25d agoHugging Face09krutrim-ai-labs /IndicSTgated IndicST: Indian Multilingual Translation Corpus For Evaluating Speech Large Language Models Introduction IndicST, a new dataset tailored for training and evaluating Speech LLMs for AST tasks (including ASR and TTS), featuring meticulously curated, automatically, and manually verified synthetic data. The dataset offers 10.8k hrs of training data and 1.13k hrs of evaluation data. Use-Cases ASR (Speech-to-Text) Transcribing Indic languages Handling… See the full description on the dataset page: https://huggingface.co/datasets/krutrim-ai-labs/IndicST.audio10M<n<100M8 likes48 downloads2y agoHugging Face10danielrosehill /multimodal-ai-taxonomy Multimodal AI Taxonomy A comprehensive, structured taxonomy for mapping multimodal AI model capabilities across input and output modalities. Dataset Description This dataset provides a systematic categorization of multimodal AI capabilities, enabling users to: Navigate the complex landscape of multimodal AI models Filter models by specific input/output modality combinations Understand the nuanced differences between similar models (e.g., image-to-video with/without audio… See the full description on the dataset page: https://huggingface.co/datasets/danielrosehill/multimodal-ai-taxonomy.textothern<1K0 likes47 downloads11mo agoHugging Face11akshaya-AI /kimiaudio1K<n<10K0 likes28 downloads2mo agoHugging Face12sleeping-ai /PlayAI-VoiceExcited to share Play AI Voice Profile. We release 267 unique voice profiles including Israeli, Arabic, Russian, Filipino and many other exclusive voice profiles. Play AI was recently acquired by Meta which sparked our interest in releasing this dataset. tabularn<1K0 likes17 downloads1y agoHugging Face13Ailiance-fr /mascarade-dsp-dataset Mascarade — DSP & Signal Processing Q&A ✅ ATTRIBUTION AUDIT COMPLETED (2026-05-11) Per-sample Stack Exchange Electronics attribution recovered via the SE /search/advanced + /questions/{id} API search : 169 samples (~5.35 %) confirmed as Stack Exchange Electronics (CC-BY-SA-4.0) — fully attributed in metadata.stack_exchange_attribution (URL + author display name + author user_id + post_id + creation_date_unix + match_confidence ≥ 0.60). 535 samples (~16.93 %) marked… See the full description on the dataset page: https://huggingface.co/datasets/Ailiance-fr/mascarade-dsp-dataset.texttext-generation1K<n<10K0 likes10 downloads5mo agoHugging Face14saidalihon /world-ai-dbaudion<1K0 likes5 downloads4mo agoHugging Face15ThisUsernameAlreadyExistsAlreadyExists /aitf-dfk3-synthetic-audio-datasetgatedaudioaudio-classification1K<n<10K0 likes4 downloads6mo agoHugging Face16AIMS-RAIL /RAIL RAIL Audio Benchmark This folder is generated for direct Hugging Face dataset upload. Each row uses relative audio paths rooted at this repository folder. NeurIPS / Croissant Metadata metadata.json is a Croissant-style metadata file with core fields and minimal RAI fields. metadata.json includes the Hugging Face dataset URL, CC-BY-4.0 license URL, checksums, and RAI fields. build_summary.json contains build counts and skipped-source diagnostics.… See the full description on the dataset page: https://huggingface.co/datasets/AIMS-RAIL/RAIL.audio10K<n<100K4 likes1d agoHugging Face17akatz-ai /H3-Character-Swap-v1 H3 Character Swap v1 A reference-conditioned character-replacement dataset for MiniMax H3 Ref2VA LoRA training with Ostris AI Toolkit. It combines synthetic still-image edits with unchanged real-motion regularization videos. 134 examples: 94 character-swap edits and 40 preservation clips. Training has 76 edits + 32 clips; validation has 18 edits + 8 clips. Prepared resolution is 1344×768 at 24 fps. The companion 1,000-step LoRA are available separately. Task and… See the full description on the dataset page: https://huggingface.co/datasets/akatz-ai/H3-Character-Swap-v1.imagen<1K0 likes23h agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.