CoolFace
16 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ground-truth /multichannel-meetings-10h GroundTruth Multi-Channel Meeting Audio Dataset (10h) Summary This dataset contains approximately 10 hours of co-located, multi-speaker meeting recordings, each captured simultaneously via a room (built-in) microphone and individual close-talk lapel microphones worn by each participant. Each meeting includes: One full meeting recording (room microphone) Individual close-talk recordings for each participant (one file per speaker) Structured metadata describing speakers… See the full description on the dataset page: https://huggingface.co/datasets/ground-truth/multichannel-meetings-10h.audioautomatic-speech-recognitionn<1K1 likes172 downloads5mo agoHugging Face02yunqi1766 /voice-code-bench VoiceCodeBench VoiceCodeBench is a test-only benchmark for evaluating whether automatic speech recognition (ASR) systems preserve exact structured values in English workplace speech. Paper: VoiceCodeBench: Evaluating Exact Structured-Token Recovery in Automatic Speech Recognition The benchmark targets cases where a transcript is software input: callback numbers, email addresses, command-line flags, file paths, URLs, account identifiers, dates, measurements, and similar values… See the full description on the dataset page: https://huggingface.co/datasets/yunqi1766/voice-code-bench.audioautomatic-speech-recognitionn<1K1 likes120 downloads2mo agoHugging Face03tsdocode /open-vi-dialog-synthetic-100h OpenDialog Vietnamese Synthetic Dialogue 100h Synthetic Vietnamese two-speaker dialogue for ZipVoice-Dialog experiments. 12,000 chunks 30 seconds per chunk 100.0 hours total Each item contains S1/S2 speaker labels, turn timings, target text, relationship, pronouns, environment, topic, mood, and source reference IDs. Audio renderer: vLLM-Omni VoxCPM2 Audio format: mono WAV, 48 kHz, 30 seconds per chunk This is a research dataset. Review the source/reference licensing and the… See the full description on the dataset page: https://huggingface.co/datasets/tsdocode/open-vi-dialog-synthetic-100h.audiotext-to-speech10K<n<100K0 likes97 downloads1mo agoHugging Face04rustam1221 /uzbek-asr-train-manifests Uzbek ASR Training Manifests The exact training, validation and test splits behind rustam1221/uzbek-asr-gigaam: 974 hours of Uzbek speech drawn from seven public corpora, filtered, text-normalized, and split by speaker. No audio is copied. Each row is a pointer — a parquet file plus a row index in the upstream dataset — and the training dataloader decodes the audio when the batch is built. That keeps the whole corpus definition at 200 MB instead of roughly a terabyte of… See the full description on the dataset page: https://huggingface.co/datasets/rustam1221/uzbek-asr-train-manifests.textautomatic-speech-recognition1K<n<10K0 likes53 downloads21d agoHugging Face05abdo1819 /arabic-english-code-switching-review-annotations Review Annotations for Arabic-English Code-Switching Speech This metadata-only dataset publishes review decisions and transcript-correction deltas for MohamedRashad/arabic-english-code-switching. It contains no human audio, no local file paths, no raw review notes, and no copies of unchanged upstream transcripts. The annotations are pinned to upstream revision 4a3bffc45219c35949470de32b8d4cb328b0ce11 and join by upstream_row_index. Coverage and outcomes The… See the full description on the dataset page: https://huggingface.co/datasets/abdo1819/arabic-english-code-switching-review-annotations.tabularautomatic-speech-recognition10K<n<100K0 likes33 downloads2mo agoHugging Face06ChristophSchuhmann /advanced-soundscapes-stage-1 Advanced Soundscapes Stage 1 — Raw Components This dataset contains Stage 1 output from the LAION Universal Audio Annotation Pipeline (UAAP) data generation plan. Contents 0 shard(s) containing 0 soundscape recipes with raw audio components Each soundscape row includes: recipe.json — full recipe with timeline, events, loudness, speaker IDs, overlap/density settings spkN.flac / spkN.json — raw speech components + full source metadata musicN.flac / musicN.json —… See the full description on the dataset page: https://huggingface.co/datasets/ChristophSchuhmann/advanced-soundscapes-stage-1.tabularaudio-classificationn<1K0 likes32 downloads3mo agoHugging Face07theblackcat102 /audio-alpacatexttext-generation10K<n<100K1 likes27 downloads3y agoHugging Face08theblackcat102 /quantized-common-voice-entextautomatic-speech-recognition1M<n<10M1 likes26 downloads3y agoHugging Face09Abhisingh-18 /hindi-english-codeswitch-dataset Hindi-English Code-Switch ASR Transcripts Text transcripts and metadata for a large bilingual Hindi-English code-switch ASR training corpus, used to train Abhisingh-18/hindi-english-codeswitch-asr. This release contains transcripts and metadata only — no audio files. Audio was sourced from multiple corpora and institutions and is not redistributed here. Credits Speech data collection and curation credit: SPRING Lab, IIT Madras. Contents File… See the full description on the dataset page: https://huggingface.co/datasets/Abhisingh-18/hindi-english-codeswitch-dataset.textautomatic-speech-recognition1M<n<10M0 likes25 downloads1mo agoHugging Face10theblackcat102 /common-voice-en-revoicetextautomatic-speech-recognition10K<n<100K0 likes21 downloads3y agoHugging Face11Shawal777 /yogera_runyankore_ailab_4_0_1imageautomatic-speech-recognition1K<n<10K0 likes16 downloads2y agoHugging Face12Raiff1982 /evalgated Dataset Card for Dataset Name This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Dataset Details Dataset Description Curated by: jonathan harrison Funded by [optional]: [More Information Needed] Shared by [optional]: [More Information Needed] Language(s) (NLP): [More Information Needed] License: [More Information Needed] Dataset Sources [optional] Repository: [More Information… See the full description on the dataset page: https://huggingface.co/datasets/Raiff1982/eval.texttext-classificationn<1K0 likes13 downloads1y agoHugging Face13Tnaot /SPS-Bopha-Voice-Dataset-v1gated VibeVoice Fine-Tuning Dataset: SPS-Bopha-Voice-Dataset-v1 This dataset is formatted for fine-tuning VibeVoice. Structure training_data.jsonl: The main manifest file containing transcriptions and paths. chunks_staging/: Directory containing the audio clips. Usage with VibeVoice Clone this repository: git clone https://huggingface.co/datasets/Tnaot/SPS-Bopha-Voice-Dataset-v1 cd SPS-Bopha-Voice-Dataset-v1 Run the training script pointing to… See the full description on the dataset page: https://huggingface.co/datasets/Tnaot/SPS-Bopha-Voice-Dataset-v1.audiotext-to-speech1K<n<10K0 likes13 downloads10mo agoHugging Face14NbAiLab /nb-asr-qwen3whisperxagreement-v1 nb-asr-qwen3whisperxagreement-v1 Word-level forced alignment training data for Norwegian speech, produced by keeping only examples where two independent aligners — WhisperX and Qwen3 (Lunde forced aligner) — agree within a tight tolerance. Dataset Description This dataset contains 702,067 speech segments drawn from the NB-ASR Norwegian audio corpus. Each record pairs an audio file with a word-level forced alignment in a format suitable for training a… See the full description on the dataset page: https://huggingface.co/datasets/NbAiLab/nb-asr-qwen3whisperxagreement-v1.textautomatic-speech-recognition100K<n<1M0 likes11 downloads4mo agoHugging Face15theblackcat102 /quantized-librispeech-train-360textautomatic-speech-recognition100K<n<1M0 likes7 downloads3y agoHugging Face16gallip0li /medimind-r11-traingated MediMind R11 — ASR training data Unified manifest + packed audio for fine-tuning Whisper-large-v3 on Norwegian clinical and conversational speech. Training manifest: r11_manifest.jsonl — 11,022 packs Held-out eval set: r11_heldout_eval.jsonl — 291 packs (NEVER train on these) ~see manifest audit packs total 11 sources: lege_*, podcasts (motiv/podk/stet), nb_samtale, nb_tale_m3, tts_drugs Schema See r11_manifest.jsonl (one JSON object per line) and… See the full description on the dataset page: https://huggingface.co/datasets/gallip0li/medimind-r11-train.audioautomatic-speech-recognitionn<1K0 likes1 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.