CoolFace
22 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01hf-audio /open-asr-leaderboard-resultstabularn<1K0 likes4.6k downloads11h agoHugging Face02ysdede /asr_benchmark_storetabularn<1K1 likes371 downloads2mo agoHugging Face03sophia8888 /clipquill-asr-benchmark Measuring whisper-tiny vs whisper-base in a browser tab Word error rate, wall-clock timing, transfer size and peak memory for two quantised Whisper tiers running entirely client-side in a real Chrome window, with the scripts that produced every number. If you are building an in-browser transcription page, the two results worth knowing before you pick a model tier: On clean synthetic audio the two tiers tie. If that is all you test, you will conclude the tier does not matter… See the full description on the dataset page: https://huggingface.co/datasets/sophia8888/clipquill-asr-benchmark.tabularautomatic-speech-recognitionn<1K0 likes135 downloads4d agoHugging Face04s512757 /polish-tedx-asr-eval Polish-TEDx-ASR-Eval A dataset for evaluating automatic speech recognition (ASR) systems for Polish in the domain of TEDx public talks. Contains audio segments from Polish TEDx talks available on YouTube (CC BY-NC-ND 4.0) and synthetic speech generated with KugelAudio (MIT), with manually created and cross-verified transcriptions. Created as part of the course "Workshops on Evaluation of Speech Recognition Systems" (ZWESUI, AMU 2026) by Group 1. Statistics… See the full description on the dataset page: https://huggingface.co/datasets/s512757/polish-tedx-asr-eval.audioautomatic-speech-recognitionn<1K0 likes88 downloads3mo agoHugging Face05kvest /Swedia-ASR-Dataset Swedia ASR Dataset This repository contains a small Swedish ASR evaluation dataset based on speech transcriptions from Swedia 2000. It was assembled to compare automatic speech-recognition output against manually corrected reference transcriptions for Swedish dialectal speech. The dataset is useful for quick experiments with Swedish ASR systems, especially when you want to inspect recognition quality on spontaneous speech from different regions, speakers, ages, and genders.… See the full description on the dataset page: https://huggingface.co/datasets/kvest/Swedia-ASR-Dataset.tabularautomatic-speech-recognitionn<1K1 likes50 downloads5mo agoHugging Face06almaz-nlp /almaz-asr-roster The ALMAZ ASR Roster Archived at Zenodo: 10.5281/zenodo.22761871 (concept DOI, always resolves to the latest version). A curated catalog of Azerbaijani speech-to-text artifacts: corpora, models, services, benchmarks and tools. Companion to the ALMAZ Resource Roster, which does the same for text. Schema matches the text roster so the two join, plus three columns speech needs and text does not: hours, condition, and verified. The verified column The standard way a… See the full description on the dataset page: https://huggingface.co/datasets/almaz-nlp/almaz-asr-roster.tabularn<1K0 likes50 downloads8d agoHugging Face07sidleal /CORAA-MUPE-ASRtabular100K<n<1M0 likes44 downloads2y agoHugging Face08witcheer /sovereign-asr-bench Sovereign ASR Bench — RTX 5090 Local, self-hosted automatic speech recognition benchmarks on one RTX 5090 32GB. Part of the WITCHEER local-AI rig. Methodology that matters: load-once measurement (so RTFx times transcription, not model load), one shared text normalizer applied to every model output and reference, and micro-averaged WER (total errors / total reference words — the LibriSpeech standard). The board lives as data in board.csv (shown in the viewer). Board —… See the full description on the dataset page: https://huggingface.co/datasets/witcheer/sovereign-asr-bench.tabularn<1K0 likes38 downloads3mo agoHugging Face09bengaliAI /ben10-asr-results Ben-10 Regional ASR — public results Score rows for the maintainer-run Ben-10 regional dialect ASR leaderboard. Field Meaning model_id Hub id or slug model_url Link to weights / paper wer Corpus Word Error Rate on private ben-10-test (lower better) wer_by_region JSON map region → WER backend Decode stack used by maintainers scorer_commit / decode_commit Git SHAs in BengaliAI/reg-speech-aacl evaluated_at ISO date requested_by Who asked, or maintainer if… See the full description on the dataset page: https://huggingface.co/datasets/bengaliAI/ben10-asr-results.tabularautomatic-speech-recognitionn<1K0 likes37 downloads2mo agoHugging Face10uam-wmi-asr-eval-labs /2026-dwesui-g01-neurologia DWESUI 2026 - Grupa 1 - neurologia Robocza/archiwalna kopia zbioru ewaluacyjnego ASR zbudowanego przez studentow kursu Warsztaty z ewaluacji systemow rozpoznawania mowy (UAM WMI), edycja 2026, tryb dzienny. Zespol (atrybucja): Grupa 1 (DWESUI 2026) Zrodlo oryginalne: https://huggingface.co/datasets/JankesTNJ/dwesui-grupa-1-neurologia Domena: neurologia Licencja zrodla: nagrania YouTube CC-BY + synteza TTS Status: kopia publiczna w organizacji kursowej (zespół opublikował zbiór… See the full description on the dataset page: https://huggingface.co/datasets/uam-wmi-asr-eval-labs/2026-dwesui-g01-neurologia.audioautomatic-speech-recognitionn<1K0 likes34 downloads1mo agoHugging Face11uam-wmi-asr-eval-labs /2026-dwesui-g02-kulinarna DWESUI 2026 - Grupa 2 - kulinarna (PIEROGA) Robocza/archiwalna kopia zbioru ewaluacyjnego ASR zbudowanego przez studentow kursu Warsztaty z ewaluacji systemow rozpoznawania mowy (UAM WMI), edycja 2026, tryb dzienny. Zespol (atrybucja): Grupa 2 (DWESUI 2026) Zrodlo oryginalne: https://huggingface.co/datasets/s479246/dwesui-grupa-2-kulinarna Domena: kulinarna Licencja zrodla: nagrania YouTube CC-BY/CC-BY-SA + TTS Status: kopia publiczna w organizacji kursowej (zespół opublikował… See the full description on the dataset page: https://huggingface.co/datasets/uam-wmi-asr-eval-labs/2026-dwesui-g02-kulinarna.audioautomatic-speech-recognitionn<1K0 likes21 downloads1mo agoHugging Face12reach-vb /open-asr-leaderboard-evals-alltabularn<1K0 likes17 downloads3y agoHugging Face13SabrinaSadiekh /responses-and-asr-labels-small-models LLM Responses and ASR Labels — Small Models Model responses to harmful prompts, labelled by 4 LLM-as-judge guards.Companion dataset for the master's thesis ASR Signal Geometry: Dense Representations vs. SAE Features (HSE, 2025). Dataset composition N = 4 326 prompts per model, (no adversarial suffix). Two sources: Source N Description JailbreakBench () 100 Curated harmful behaviours Anthropic HH-RLHF red-team-attempts () 4 226 Red-team conversations… See the full description on the dataset page: https://huggingface.co/datasets/SabrinaSadiekh/responses-and-asr-labels-small-models.tabulartext-classification10K<n<100K0 likes14 downloads3mo agoHugging Face14Subu19 /Devnagari-ASRaudio1K<n<10K0 likes11 downloads2y agoHugging Face15reach-vb /open-asr-leaderboard-evals-ex-cvtabularn<1K0 likes9 downloads3y agoHugging Face16cyingliu /partial-asrtabular100K<n<1M0 likes8 downloads3y agoHugging Face17prakhargupta94 /data_llama_vision_20_asrtabular1K<n<10K0 likes5 downloads3y agoHugging Face18sulaimank /voco-four-asrgatedtabular1K<n<10K0 likes5 downloads1y agoHugging Face19lisdfdf /open-asr-leaderboard-evals-alltabularn<1K0 likes4 downloads4mo agoHugging Face20quinnlue /asr-evalstabular100K<n<1M0 likes2 downloads6mo agoHugging Face21Asrithaa29 /model-capability-classificationtabular1K<n<10K0 likes2 downloads6mo agoHugging Face22enesssssw /child-asr-phoneme-pseudolabelstabular10K<n<100K0 likes1 downloads7mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.