CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ggfox00000 /stt-vibravox-fr-test VibraVox FR — test split (mirror of Cnam-LMSSC/vibravox) Mirror public des splits test de VibraVox (CNAM-LMSSC, Paris) pour benchmark ASR français multi-capteur sur audio standard ET non-standard (bone-conduction, in-ear, throat, accéléromètre). Ce repo contient uniquement les configs speech_clean + speech_noisy (les seules avec transcription). Les configs speechless_* upstream sont exclues car sans texte → pas de WER possible. Configs Config Test rows Test… See the full description on the dataset page: https://huggingface.co/datasets/ggfox00000/stt-vibravox-fr-test.audioautomatic-speech-recognition1K<n<10K0 likes380 downloads5mo agoHugging Face02rasgaard /fleurs_test FLEURS Test Dataset with Enhanced Metadata This dataset is an enhanced version of the FLEURS (Few-shot Learning Evaluation of Universal Representations of Speech) test set, restructured with complete metadata for easier use in automatic speech recognition (ASR) and multilingual speech processing tasks. Dataset Description FLEURS is a multilingual speech benchmark dataset designed to evaluate universal speech representations. This particular version focuses on 25 European… See the full description on the dataset page: https://huggingface.co/datasets/rasgaard/fleurs_test.audioautomatic-speech-recognition10K<n<100K0 likes181 downloads7mo agoHugging Face03FBK-MT /Speech-MASSIVE-testgated Speech-MASSIVE Test Split This dataset repository is only for test split of Speech-MASSIVE. train and dev splits are available in the separate dataset repository. https://huggingface.co/datasets/FBK-MT/Speech-MASSIVE Dataset Description Speech-MASSIVE is a multilingual Spoken Language Understanding (SLU) dataset comprising the speech counterpart for a portion of the MASSIVE textual corpus. Speech-MASSIVE covers 12 languages (Arabic, German, Spanish, French, Hungarian… See the full description on the dataset page: https://huggingface.co/datasets/FBK-MT/Speech-MASSIVE-test.audioaudio-classification10K<n<100K9 likes139 downloads1y agoHugging Face04shrikanth-19 /dhravani-mit-testCheck out the configuration reference at https://huggingface.co/docs/hub/spaces-config-reference Dataset Preparation Interface for Fine-tuning Whisper A web-based interface for preparing audio datasets to fine-tune OpenAI's Whisper model. This tool helps in recording, managing, and organizing voice recordings with their corresponding transcriptions, with support for cloud storage and authentication. Features 🔐 User authentication via Pocketbase ☁️ Cloud storage… See the full description on the dataset page: https://huggingface.co/datasets/shrikanth-19/dhravani-mit-test.audioautomatic-speech-recognitionn<1K0 likes118 downloads7d agoHugging Face05Yehor /RS-testaudioautomatic-speech-recognition10K<n<100K0 likes116 downloads3mo agoHugging Face06shrikanth-19 /dhravani-iitpatna-testCheck out the configuration reference at https://huggingface.co/docs/hub/spaces-config-reference Dataset Preparation Interface for Fine-tuning Whisper A web-based interface for preparing audio datasets to fine-tune OpenAI's Whisper model. This tool helps in recording, managing, and organizing voice recordings with their corresponding transcriptions, with support for cloud storage and authentication. Features 🔐 User authentication via Pocketbase ☁️ Cloud storage… See the full description on the dataset page: https://huggingface.co/datasets/shrikanth-19/dhravani-iitpatna-test.audioautomatic-speech-recognitionn<1K0 likes105 downloads7d agoHugging Face07shrikanth-19 /dhravani-IIT_Guwahati-testCheck out the configuration reference at https://huggingface.co/docs/hub/spaces-config-reference Dataset Preparation Interface for Fine-tuning Whisper A web-based interface for preparing audio datasets to fine-tune OpenAI's Whisper model. This tool helps in recording, managing, and organizing voice recordings with their corresponding transcriptions, with support for cloud storage and authentication. Features 🔐 User authentication via Pocketbase ☁️ Cloud storage… See the full description on the dataset page: https://huggingface.co/datasets/shrikanth-19/dhravani-IIT_Guwahati-test.audioautomatic-speech-recognitionn<1K0 likes99 downloads7d agoHugging Face08ggfox00000 /stt-cefc-fr-test CEFC-Orfeo FR — long-form oral test mirror Mirror non-officiel du Corpus d'Études du Français Contemporain (CEFC) agrégé par le projet Orfeo (ANR), tel que distribué sur le portail projet-orfeo.fr (release 13). Long-form : 1 row = 1 fichier audio entier (30-60 min en moyenne). 12 sous-corpus oraux du français contemporain, 303 heures au total, 901 fichiers. Idéal pour bench Whisper / Canary en conditions réelles (chunked decoding, dérive temporelle, multi-locuteurs).… See the full description on the dataset page: https://huggingface.co/datasets/ggfox00000/stt-cefc-fr-test.audioautomatic-speech-recognitionn<1K0 likes97 downloads5mo agoHugging Face09shrikanth-19 /dhravani-IGDTUW_Delhi-testCheck out the configuration reference at https://huggingface.co/docs/hub/spaces-config-reference Dataset Preparation Interface for Fine-tuning Whisper A web-based interface for preparing audio datasets to fine-tune OpenAI's Whisper model. This tool helps in recording, managing, and organizing voice recordings with their corresponding transcriptions, with support for cloud storage and authentication. Features 🔐 User authentication via Pocketbase ☁️ Cloud storage… See the full description on the dataset page: https://huggingface.co/datasets/shrikanth-19/dhravani-IGDTUW_Delhi-test.audioautomatic-speech-recognitionn<1K0 likes93 downloads7d agoHugging Face10shrikanth-19 /dhravani-iitdelhi-testCheck out the configuration reference at https://huggingface.co/docs/hub/spaces-config-reference Dataset Preparation Interface for Fine-tuning Whisper A web-based interface for preparing audio datasets to fine-tune OpenAI's Whisper model. This tool helps in recording, managing, and organizing voice recordings with their corresponding transcriptions, with support for cloud storage and authentication. Features 🔐 User authentication via Pocketbase ☁️ Cloud storage… See the full description on the dataset page: https://huggingface.co/datasets/shrikanth-19/dhravani-iitdelhi-test.audioautomatic-speech-recognitionn<1K0 likes91 downloads7d agoHugging Face11shrikanth-19 /dhravani-iiitdelhi-testCheck out the configuration reference at https://huggingface.co/docs/hub/spaces-config-reference Dataset Preparation Interface for Fine-tuning Whisper A web-based interface for preparing audio datasets to fine-tune OpenAI's Whisper model. This tool helps in recording, managing, and organizing voice recordings with their corresponding transcriptions, with support for cloud storage and authentication. Features 🔐 User authentication via Pocketbase ☁️ Cloud storage… See the full description on the dataset page: https://huggingface.co/datasets/shrikanth-19/dhravani-iiitdelhi-test.audioautomatic-speech-recognitionn<1K0 likes88 downloads7d agoHugging Face12Steveeeeeeen /edacc_testaudioautomatic-speech-recognition1K<n<10K0 likes87 downloads2y agoHugging Face13ggfox00000 /stt-covost25-test-fr Common Voice FR — read + spontaneous + CoVoST 2 test mirror Mirror combiné de trois ressources Mozilla / Facebook AI Research dans le même repo pour benchmark WER/STT français multi-paradigme. Config read (défaut) Source upstream : Mozilla Common Voice 25.0 FR (paradigme lecture). 16 149 utterances, ~21 h. Locuteurs très variés (crowdsourced). Paradigme : contributeurs lisent à voix haute des phrases écrites d'un pool partagé. Débit régulier, peu d'hésitations. Référence… See the full description on the dataset page: https://huggingface.co/datasets/ggfox00000/stt-covost25-test-fr.audioautomatic-speech-recognition10K<n<100K0 likes86 downloads5mo agoHugging Face14ggfox00000 /stt-summre-fr-test SUMM-RE — French test split (mirror of linagora/SUMM-RE) Mirror public du split test de SUMM-RE (LINAGORA / Aix-Marseille LPL), pour benchmark ASR français conversationnel (parole de réunion, 3-4 locuteurs, ~20 min par session). ⚠ Ce repo ne contient que le split test (124 tracks individuelles = 37 réunions × 3-4 micros). Pour les splits train / dev, voir le repo upstream linagora/SUMM-RE. Contenu 124 pistes audio individuelles (1 piste = 1 microphone d'un locuteur… See the full description on the dataset page: https://huggingface.co/datasets/ggfox00000/stt-summre-fr-test.audioautomatic-speech-recognitionn<1K0 likes85 downloads5mo agoHugging Face15Yehor /RS-test-fix2audioautomatic-speech-recognition10K<n<100K0 likes72 downloads3mo agoHugging Face16g-group-ai-lab /vi-asr-tech-test Vietnamese ASR Test Set - Technology A Vietnamese speech-recognition benchmark for the technology domain (Công nghệ), released by G-Group AI Lab. Audio is real-world Vietnamese speech covering consumer electronics reviews, software tutorials, programming and IT walkthroughs — dense in English loanwords and product names. Listen & explore Every utterance is playable inline in the viewer above — hit play on any row to stream the clip. For full-text search across… See the full description on the dataset page: https://huggingface.co/datasets/g-group-ai-lab/vi-asr-tech-test.audioautomatic-speech-recognitionn<1K0 likes66 downloads2mo agoHugging Face17arda-argmax /fastmss-v0.5.0-test FastMSS synthetic multi-speaker meetings - parquet edition Streaming-friendly parquet shards of the FastMSS synthetic multi-speaker conversational corpus. Each row is one mixture with the audio bytes embedded inline (16 kHz mono WAV) plus per-segment diarization timestamps, per-word transcript and the full lhotse cut as a JSON blob. See fastmss/hf_dataset.py for the schema docstring. Subsets and splits v0.5.0_test — splits: train — 1000 mixtures, 1081.0 min total, 3609… See the full description on the dataset page: https://huggingface.co/datasets/arda-argmax/fastmss-v0.5.0-test.audioautomatic-speech-recognition1K<n<10K0 likes61 downloads4mo agoHugging Face18Steveeeeeeen /edacc_test_cleanaudioautomatic-speech-recognition1K<n<10K0 likes59 downloads2y agoHugging Face19adalat-ai /vividh-test-hindi 🎙️ Vividh-ASR Benchmark — Hindi (Test Split) How well does your ASR model actually work in the wild? Vividh-ASR is a complexity-stratified benchmark that tells you exactly where your model succeeds — and where it falls apart. Most Indic ASR benchmarks evaluate models on clean, studio-recorded speech. Real-world audio is not that. Vividh-ASR organises evaluation by acoustic complexity rather than domain, exposing the studio-bias that plagues models fine-tuned predominantly on read… See the full description on the dataset page: https://huggingface.co/datasets/adalat-ai/vividh-test-hindi.audioautomatic-speech-recognition10K<n<100K1 likes59 downloads4mo agoHugging Face20avemio /ASR-GERMAN-MIXED-TEST Dataset Beschreibung Dieser Datensatz und die Beschreibung wurde von flozi00/asr-german-mixed übernommen und nur der Test-Split hier hochgeladen, da Hugging Face native erst einmal alle Splits herunterlädt. Für eine Evaluation von Speech-to-Text Modellen ist ein Download von 136 GB allerdings etwas zeit- & speicherraubend, weshalb wir hier nur den Test-Split für Evaluationen anbieten möchten. Die Arbeit und die Anerkennung sollten deshalb weiter bei primeline & flozi00 für die… See the full description on the dataset page: https://huggingface.co/datasets/avemio/ASR-GERMAN-MIXED-TEST.audioautomatic-speech-recognition1K<n<10K3 likes56 downloads2y agoHugging Face21Trelis /ami-2speaker-test AMI 2-Speaker Test Set Need a voice model for your domain? Trelis builds custom ASR, TTS, and voice agent pipelines for specialist verticals (legal, medical, finance, construction) and low-resource languages. Enquire or book a consultation → A 50-clip benchmark for 2-speaker overlapping speech recognition, derived from the AMI Meeting Corpus test split. Each clip is 8–28 seconds of real conversational meeting audio reconstructed as a 2-speaker virtual meeting, with separate… See the full description on the dataset page: https://huggingface.co/datasets/Trelis/ami-2speaker-test.audioautomatic-speech-recognitionn<1K0 likes54 downloads5mo agoHugging Face22g-group-ai-lab /vi-asr-finance-test Vietnamese ASR Test Set - Finance A Vietnamese speech-recognition benchmark for the finance domain (Tài chính), released by G-Group AI Lab. Audio is real-world Vietnamese speech covering stock market commentary, trading platforms, banking and crypto — dense in tickers, numbers and financial jargon. Listen & explore Every utterance is playable inline in the viewer above — hit play on any row to stream the clip. For full-text search across transcripts, duration… See the full description on the dataset page: https://huggingface.co/datasets/g-group-ai-lab/vi-asr-finance-test.audioautomatic-speech-recognitionn<1K0 likes54 downloads2mo agoHugging Face23g-group-ai-lab /vi-asr-edu-test Vietnamese ASR Test Set - Education A Vietnamese speech-recognition benchmark for the education domain (Giáo dục), released by G-Group AI Lab. Audio is real-world Vietnamese speech covering study-abroad consulting, exam and certification guidance, university and training-course introductions. Listen & explore Every utterance is playable inline in the viewer above — hit play on any row to stream the clip. For full-text search across transcripts, duration filters… See the full description on the dataset page: https://huggingface.co/datasets/g-group-ai-lab/vi-asr-edu-test.audioautomatic-speech-recognitionn<1K0 likes52 downloads2mo agoHugging Face24olympusmons /librispeech_asr_test_clean_word_timestamp Word-level timestamp annotated Librispeech ASR test set This dataset contains word-level timestamp information for the Librispeech ASR test (clean) dataset. It contains 2620 short files that have been force-aligned with its text to get reasonably accurate word-level timestamp information. Suitable for use in timestamp benchmarking of ASR models or audio dataset preprocessing. To request access to more datasets like this, please fill out this form:… See the full description on the dataset page: https://huggingface.co/datasets/olympusmons/librispeech_asr_test_clean_word_timestamp.audioautomatic-speech-recognition1K<n<10K1 likes50 downloads2y agoHugging Face25g-group-ai-lab /vi-asr-pubadmin-test Vietnamese ASR Test Set - Public Administration A Vietnamese speech-recognition benchmark for the public administration domain (Hành chính công), released by G-Group AI Lab. Audio is real-world Vietnamese speech covering administrative procedures, paperwork and licensing guidance, civil records — dense in place names and legal terminology. Listen & explore Every utterance is playable inline in the viewer above — hit play on any row to stream the clip. For… See the full description on the dataset page: https://huggingface.co/datasets/g-group-ai-lab/vi-asr-pubadmin-test.audioautomatic-speech-recognitionn<1K0 likes48 downloads2mo agoHugging Face26Yehor /cv10-uk-testset-clean The cleaned Common Voice 10 (test set) that has been checked by a human for Ukrainian 🇺🇦 Overview This repository contains the archive of Common Voice 10 (test set) with checked Ukrainian transcriptions and audios. All audios have been checked by a human to be sure that they are correct. This archive is used to test all ASR models listed here: https://github.com/egorsmkv/speech-recognition-uk Community Discord: https://bit.ly/discord-uds Speech… See the full description on the dataset page: https://huggingface.co/datasets/Yehor/cv10-uk-testset-clean.audioautomatic-speech-recognition1K<n<10K3 likes45 downloads2y agoHugging Face27ymoslem /IWSLT2025-Test Dataset details This is the blind test set of IWSLT 2025's model compression track. It consists of audio extracted from ACL presentations. For training data in the same domain, the ACL 60/60 dataset can be used. Citation @INPROCEEDINGS{Abdulmumin2025-IWSLT, title = "{Findings of the IWSLT 2025 Evaluation Campaign}", author = "Abdulmumin, Idris and Agostinelli, Victor and Alumäe, Tanel and Anastasopoulos, Antonios and {Ashwin} and Bentivogli… See the full description on the dataset page: https://huggingface.co/datasets/ymoslem/IWSLT2025-Test.audioautomatic-speech-recognitionn<1K0 likes44 downloads1y agoHugging Face28adalat-ai /vividh-test-malayalam 🎙️ Vividh-ASR Benchmark — Malayalam (Test Split) How well does your ASR model actually work in the wild?Vividh-ASR is a complexity-stratified benchmark that tells you exactly where your model succeeds — and where it falls apart. Most Indic ASR benchmarks evaluate models on clean, studio-recorded speech. Real-world audio is not that. Vividh-ASR organises evaluation by acoustic complexity rather than domain, exposing the studio-bias that plagues models fine-tuned predominantly on… See the full description on the dataset page: https://huggingface.co/datasets/adalat-ai/vividh-test-malayalam.audioautomatic-speech-recognition10K<n<100K1 likes44 downloads4mo agoHugging Face29freyavoice /covost2-tr-test covost2-tr-test Turkish test split of CoVoST 2 (tr_en, Turkish source) (Common Voice–based ST/ASR corpus), re-hosted for Turkish STT benchmarking. Rows: 1629 Columns: client_id, file, audio, sentence, translation, id (sentence = Turkish transcript, translation = English) Source: https://github.com/facebookresearch/covost (audio mirror: fixie-ai/covost2) License: cc0-1.0 Only the Turkish source test split is included, extracted as-is. audioautomatic-speech-recognition1K<n<10K0 likes38 downloads3mo agoHugging Face30Lauler /riksdagen_testtabularautomatic-speech-recognition1K<n<10K0 likes36 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.