CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01yuekai /speechio_test SpeechIO ASR Test Sets (parquet) Parquet repackaging of the SpeechColab SpeechIO Mandarin ASR benchmark, re-exported from yuekai/speechio (Lhotse cuts) into standard HuggingFace parquet with embedded 16 kHz audio. 27 test sets: SPEECHIO_ASR_ZH00000 ... SPEECHIO_ASR_ZH00026, each a config with a single test split. ~43k utterances, ~66 hours total, evaluation only. Columns column type note segment_id string utterance id speaker string speaker id… See the full description on the dataset page: https://huggingface.co/datasets/yuekai/speechio_test.audioautomatic-speech-recognition10K<n<100K0 likes404 downloads3mo agoHugging Face02rasgaard /fleurs_test FLEURS Test Dataset with Enhanced Metadata This dataset is an enhanced version of the FLEURS (Few-shot Learning Evaluation of Universal Representations of Speech) test set, restructured with complete metadata for easier use in automatic speech recognition (ASR) and multilingual speech processing tasks. Dataset Description FLEURS is a multilingual speech benchmark dataset designed to evaluate universal speech representations. This particular version focuses on 25 European… See the full description on the dataset page: https://huggingface.co/datasets/rasgaard/fleurs_test.audioautomatic-speech-recognition10K<n<100K0 likes171 downloads7mo agoHugging Face03anjalyjayakrishnan /testThe Snow Mountain dataset contains the audio recordings (in .mp3 format) and the corresponding text of The Bible in 11 Indian languages. The recordings were done in a studio setting by native speakers. Each language has a single speaker in the dataset. Most of these languages are geographically concentrated in the Northern part of India around the state of Himachal Pradesh. Being related to Hindi they all use the Devanagari script for transcription.audioautomatic-speech-recognition1K<n<10K0 likes157 downloads4y agoHugging Face04FBK-MT /Speech-MASSIVE-testgated Speech-MASSIVE Test Split This dataset repository is only for test split of Speech-MASSIVE. train and dev splits are available in the separate dataset repository. https://huggingface.co/datasets/FBK-MT/Speech-MASSIVE Dataset Description Speech-MASSIVE is a multilingual Spoken Language Understanding (SLU) dataset comprising the speech counterpart for a portion of the MASSIVE textual corpus. Speech-MASSIVE covers 12 languages (Arabic, German, Spanish, French, Hungarian… See the full description on the dataset page: https://huggingface.co/datasets/FBK-MT/Speech-MASSIVE-test.audioaudio-classification10K<n<100K9 likes139 downloads1y agoHugging Face05Yehor /RS-testaudioautomatic-speech-recognition10K<n<100K0 likes135 downloads2mo agoHugging Face06ggfox00000 /stt-vibravox-fr-test VibraVox FR — test split (mirror of Cnam-LMSSC/vibravox) Mirror public des splits test de VibraVox (CNAM-LMSSC, Paris) pour benchmark ASR français multi-capteur sur audio standard ET non-standard (bone-conduction, in-ear, throat, accéléromètre). Ce repo contient uniquement les configs speech_clean + speech_noisy (les seules avec transcription). Les configs speechless_* upstream sont exclues car sans texte → pas de WER possible. Configs Config Test rows Test… See the full description on the dataset page: https://huggingface.co/datasets/ggfox00000/stt-vibravox-fr-test.audioautomatic-speech-recognition1K<n<10K0 likes127 downloads5mo agoHugging Face07Steveeeeeeen /edacc_testaudioautomatic-speech-recognition1K<n<10K0 likes117 downloads2y agoHugging Face08Steveeeeeeen /edacc_test_cleanaudioautomatic-speech-recognition1K<n<10K0 likes114 downloads2y agoHugging Face09shrikanth-19 /dhravani-mit-testCheck out the configuration reference at https://huggingface.co/docs/hub/spaces-config-reference Dataset Preparation Interface for Fine-tuning Whisper A web-based interface for preparing audio datasets to fine-tune OpenAI's Whisper model. This tool helps in recording, managing, and organizing voice recordings with their corresponding transcriptions, with support for cloud storage and authentication. Features 🔐 User authentication via Pocketbase ☁️ Cloud storage… See the full description on the dataset page: https://huggingface.co/datasets/shrikanth-19/dhravani-mit-test.audioautomatic-speech-recognitionn<1K0 likes111 downloads5d agoHugging Face10ciempiess /ciempiess_testThe CIEMPIESS TEST Corpus is a gender balanced corpus destined to test acoustic models for the speech recognition task. The corpus was manually transcribed and it contains audio recordings from 10 male and 10 female speakers. The CIEMPIESS TEST is one of the three corpora included at the LDC's \"CIEMPIESS Experimentation\" (LDC2019S07).audioautomatic-speech-recognition1K<n<10K3 likes105 downloads3y agoHugging Face11shrikanth-19 /dhravani-iitpatna-testCheck out the configuration reference at https://huggingface.co/docs/hub/spaces-config-reference Dataset Preparation Interface for Fine-tuning Whisper A web-based interface for preparing audio datasets to fine-tune OpenAI's Whisper model. This tool helps in recording, managing, and organizing voice recordings with their corresponding transcriptions, with support for cloud storage and authentication. Features 🔐 User authentication via Pocketbase ☁️ Cloud storage… See the full description on the dataset page: https://huggingface.co/datasets/shrikanth-19/dhravani-iitpatna-test.audioautomatic-speech-recognitionn<1K0 likes103 downloads5d agoHugging Face12shrikanth-19 /dhravani-IIT_Guwahati-testCheck out the configuration reference at https://huggingface.co/docs/hub/spaces-config-reference Dataset Preparation Interface for Fine-tuning Whisper A web-based interface for preparing audio datasets to fine-tune OpenAI's Whisper model. This tool helps in recording, managing, and organizing voice recordings with their corresponding transcriptions, with support for cloud storage and authentication. Features 🔐 User authentication via Pocketbase ☁️ Cloud storage… See the full description on the dataset page: https://huggingface.co/datasets/shrikanth-19/dhravani-IIT_Guwahati-test.audioautomatic-speech-recognitionn<1K0 likes95 downloads5d agoHugging Face13sfsmcnulty /stt_test_audio Mobile Voice Platform STT Test Audio Versioned benchmark audio for the MobileVoicePlatform-Android SampleApp. The repository intentionally contains two views of the same clips: data/ is an AudioFolder-compatible view with metadata.csv for Hugging Face tooling. packs/ contains checksum-pinned ZIPs optimized for bounded download and validation on Android. catalog.json is the machine-readable index used to discover datasets, categories, languages, checksums, clip references, and… See the full description on the dataset page: https://huggingface.co/datasets/sfsmcnulty/stt_test_audio.audioautomatic-speech-recognitionn<1K0 likes93 downloads14d agoHugging Face14shrikanth-19 /dhravani-IGDTUW_Delhi-testCheck out the configuration reference at https://huggingface.co/docs/hub/spaces-config-reference Dataset Preparation Interface for Fine-tuning Whisper A web-based interface for preparing audio datasets to fine-tune OpenAI's Whisper model. This tool helps in recording, managing, and organizing voice recordings with their corresponding transcriptions, with support for cloud storage and authentication. Features 🔐 User authentication via Pocketbase ☁️ Cloud storage… See the full description on the dataset page: https://huggingface.co/datasets/shrikanth-19/dhravani-IGDTUW_Delhi-test.audioautomatic-speech-recognitionn<1K0 likes90 downloads5d agoHugging Face15shrikanth-19 /dhravani-iitdelhi-testCheck out the configuration reference at https://huggingface.co/docs/hub/spaces-config-reference Dataset Preparation Interface for Fine-tuning Whisper A web-based interface for preparing audio datasets to fine-tune OpenAI's Whisper model. This tool helps in recording, managing, and organizing voice recordings with their corresponding transcriptions, with support for cloud storage and authentication. Features 🔐 User authentication via Pocketbase ☁️ Cloud storage… See the full description on the dataset page: https://huggingface.co/datasets/shrikanth-19/dhravani-iitdelhi-test.audioautomatic-speech-recognitionn<1K0 likes87 downloads5d agoHugging Face16shrikanth-19 /dhravani-iiitdelhi-testCheck out the configuration reference at https://huggingface.co/docs/hub/spaces-config-reference Dataset Preparation Interface for Fine-tuning Whisper A web-based interface for preparing audio datasets to fine-tune OpenAI's Whisper model. This tool helps in recording, managing, and organizing voice recordings with their corresponding transcriptions, with support for cloud storage and authentication. Features 🔐 User authentication via Pocketbase ☁️ Cloud storage… See the full description on the dataset page: https://huggingface.co/datasets/shrikanth-19/dhravani-iiitdelhi-test.audioautomatic-speech-recognitionn<1K0 likes86 downloads5d agoHugging Face17JacobLinCool /audio-testing audio-testing Overview This is a small, open dataset designed for quick validation of audio-related pipelines and applications, especially for Text-to-Speech (TTS) and Speech-to-Text (STT) systems. It provides a few short, diverse audio clips and corresponding text transcripts, allowing developers to verify input/output handling, audio processing, and transcription logic without downloading large datasets. Contents 3 short audio samples (.mp3, .wav)… See the full description on the dataset page: https://huggingface.co/datasets/JacobLinCool/audio-testing.audioautomatic-speech-recognitionn<1K0 likes84 downloads11mo agoHugging Face18ggfox00000 /stt-covost25-test-fr Common Voice FR — read + spontaneous + CoVoST 2 test mirror Mirror combiné de trois ressources Mozilla / Facebook AI Research dans le même repo pour benchmark WER/STT français multi-paradigme. Config read (défaut) Source upstream : Mozilla Common Voice 25.0 FR (paradigme lecture). 16 149 utterances, ~21 h. Locuteurs très variés (crowdsourced). Paradigme : contributeurs lisent à voix haute des phrases écrites d'un pool partagé. Débit régulier, peu d'hésitations. Référence… See the full description on the dataset page: https://huggingface.co/datasets/ggfox00000/stt-covost25-test-fr.audioautomatic-speech-recognition10K<n<100K0 likes77 downloads5mo agoHugging Face19humair025 /h-test@misc{humair025/h-test, title = {h-test}, author = {Humair Munir}, year = {2025}, howpublished = {\url{https://huggingface.co/datasets/humair025/h-test}}, note = {Synthetic dataset of speech (with emotions). Licensed under CC-BY 4.0.} } audiotext-to-speechn<1K0 likes75 downloads9mo agoHugging Face20Yehor /RS-test-fix2audioautomatic-speech-recognition10K<n<100K0 likes74 downloads2mo agoHugging Face21arda-argmax /fastmss-v0.5.0-test FastMSS synthetic multi-speaker meetings - parquet edition Streaming-friendly parquet shards of the FastMSS synthetic multi-speaker conversational corpus. Each row is one mixture with the audio bytes embedded inline (16 kHz mono WAV) plus per-segment diarization timestamps, per-word transcript and the full lhotse cut as a JSON blob. See fastmss/hf_dataset.py for the schema docstring. Subsets and splits v0.5.0_test — splits: train — 1000 mixtures, 1081.0 min total, 3609… See the full description on the dataset page: https://huggingface.co/datasets/arda-argmax/fastmss-v0.5.0-test.audioautomatic-speech-recognition1K<n<10K0 likes70 downloads4mo agoHugging Face22g-group-ai-lab /vi-asr-tech-test Vietnamese ASR Test Set - Technology A Vietnamese speech-recognition benchmark for the technology domain (Công nghệ), released by G-Group AI Lab. Audio is real-world Vietnamese speech covering consumer electronics reviews, software tutorials, programming and IT walkthroughs — dense in English loanwords and product names. Listen & explore Every utterance is playable inline in the viewer above — hit play on any row to stream the clip. For full-text search across… See the full description on the dataset page: https://huggingface.co/datasets/g-group-ai-lab/vi-asr-tech-test.audioautomatic-speech-recognitionn<1K0 likes67 downloads2mo agoHugging Face23XRXRX /X-Voice-TestsetX-Voice Multilingual Test Set High-Fidelity Test Set for Multilingual Text-to-Speech across 30 Languages This test set is built as part of the research: X-Voice: One Speaker, 30+ Languages with Zero-Shot Voice Cloning, serving as the evaluation benchmark for our model. Dataset Summary 30 languages European: bg (Bulgarian), cs (Czech), da (Danish), de (German), el (Greek), en (English), es (Spanish), et (Estonian), fi (Finnish), fr (French), hr (Croatian), hu (Hungarian), it… See the full description on the dataset page: https://huggingface.co/datasets/XRXRX/X-Voice-Testset.audiotext-to-speech10K<n<100K4 likes65 downloads5mo agoHugging Face24Codyfederer /test4 test4 This is a merged speech dataset containing 345 audio segments from 2 source datasets. Dataset Information Total Segments: 345 Speakers: 7 Languages: en Emotions: neutral, sad, angry, happy Original Datasets: 2 Dataset Structure Each example contains: audio: Audio file (WAV format, 16kHz sampling rate) text: Transcription of the audio speaker_id: Unique speaker identifier (made unique across all merged datasets) emotion: Detected emotion (neutral, happy… See the full description on the dataset page: https://huggingface.co/datasets/Codyfederer/test4.audioautomatic-speech-recognitionn<1K0 likes63 downloads1y agoHugging Face25g-group-ai-lab /vi-asr-finance-test Vietnamese ASR Test Set - Finance A Vietnamese speech-recognition benchmark for the finance domain (Tài chính), released by G-Group AI Lab. Audio is real-world Vietnamese speech covering stock market commentary, trading platforms, banking and crypto — dense in tickers, numbers and financial jargon. Listen & explore Every utterance is playable inline in the viewer above — hit play on any row to stream the clip. For full-text search across transcripts, duration… See the full description on the dataset page: https://huggingface.co/datasets/g-group-ai-lab/vi-asr-finance-test.audioautomatic-speech-recognitionn<1K0 likes60 downloads2mo agoHugging Face26avemio /ASR-GERMAN-MIXED-TEST Dataset Beschreibung Dieser Datensatz und die Beschreibung wurde von flozi00/asr-german-mixed übernommen und nur der Test-Split hier hochgeladen, da Hugging Face native erst einmal alle Splits herunterlädt. Für eine Evaluation von Speech-to-Text Modellen ist ein Download von 136 GB allerdings etwas zeit- & speicherraubend, weshalb wir hier nur den Test-Split für Evaluationen anbieten möchten. Die Arbeit und die Anerkennung sollten deshalb weiter bei primeline & flozi00 für die… See the full description on the dataset page: https://huggingface.co/datasets/avemio/ASR-GERMAN-MIXED-TEST.audioautomatic-speech-recognition1K<n<10K3 likes57 downloads2y agoHugging Face27Trelis /ami-2speaker-test AMI 2-Speaker Test Set Need a voice model for your domain? Trelis builds custom ASR, TTS, and voice agent pipelines for specialist verticals (legal, medical, finance, construction) and low-resource languages. Enquire or book a consultation → A 50-clip benchmark for 2-speaker overlapping speech recognition, derived from the AMI Meeting Corpus test split. Each clip is 8–28 seconds of real conversational meeting audio reconstructed as a 2-speaker virtual meeting, with separate… See the full description on the dataset page: https://huggingface.co/datasets/Trelis/ami-2speaker-test.audioautomatic-speech-recognitionn<1K0 likes56 downloads5mo agoHugging Face28g-group-ai-lab /vi-asr-edu-test Vietnamese ASR Test Set - Education A Vietnamese speech-recognition benchmark for the education domain (Giáo dục), released by G-Group AI Lab. Audio is real-world Vietnamese speech covering study-abroad consulting, exam and certification guidance, university and training-course introductions. Listen & explore Every utterance is playable inline in the viewer above — hit play on any row to stream the clip. For full-text search across transcripts, duration filters… See the full description on the dataset page: https://huggingface.co/datasets/g-group-ai-lab/vi-asr-edu-test.audioautomatic-speech-recognitionn<1K0 likes55 downloads2mo agoHugging Face29adalat-ai /vividh-test-hindi 🎙️ Vividh-ASR Benchmark — Hindi (Test Split) How well does your ASR model actually work in the wild? Vividh-ASR is a complexity-stratified benchmark that tells you exactly where your model succeeds — and where it falls apart. Most Indic ASR benchmarks evaluate models on clean, studio-recorded speech. Real-world audio is not that. Vividh-ASR organises evaluation by acoustic complexity rather than domain, exposing the studio-bias that plagues models fine-tuned predominantly on read… See the full description on the dataset page: https://huggingface.co/datasets/adalat-ai/vividh-test-hindi.audioautomatic-speech-recognition10K<n<100K1 likes54 downloads4mo agoHugging Face30g-group-ai-lab /vi-asr-pubadmin-test Vietnamese ASR Test Set - Public Administration A Vietnamese speech-recognition benchmark for the public administration domain (Hành chính công), released by G-Group AI Lab. Audio is real-world Vietnamese speech covering administrative procedures, paperwork and licensing guidance, civil records — dense in place names and legal terminology. Listen & explore Every utterance is playable inline in the viewer above — hit play on any row to stream the clip. For… See the full description on the dataset page: https://huggingface.co/datasets/g-group-ai-lab/vi-asr-pubadmin-test.audioautomatic-speech-recognitionn<1K0 likes53 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.