CoolFace
14 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01RVtech /Audio2Tool Audio2Tool: Speak, Call, Act — A Dataset for Benchmarking Speech Tool Use Authors: Ramit Pahwa1,∗,∗∗, Apoorva Beedu1,∗, Parivesh Priye1, Rutu Gandhi†1, Saloni Takawale†1, Aruna Baijal1, Zengli Yang1 1 Rivian & Volkswagen Technologies &nbsp;·&nbsp; ∗ equal contribution &nbsp;·&nbsp; ∗∗ corresponding author &nbsp;·&nbsp; † equal contribution 📄 Project page / demo: https://audio2tool.github.io/ 📦 Dataset: https://huggingface.co/datasets/RVtech/Audio2Tool ✉️ Contact (corresponding… See the full description on the dataset page: https://huggingface.co/datasets/RVtech/Audio2Tool.audioautomatic-speech-recognition10K<n<100K2 likes5.9k downloads3mo agoHugging Face02Edge0 /ark-asr-3b-open-asr-leaderboard-results ARK-ASR-3B Open ASR Leaderboard Results Raw JSONL manifests for AutoArk-AI/ARK-ASR-3B on the public English short-form hf-audio/open-asr-leaderboard splits. These manifests were generated on a local 8x RTX 4090 machine and scored with the shared Open ASR Leaderboard scorer: PYTHONPATH=. python - <<'PY' from normalizer.eval_utils import score_results score_results( 'ark_asr/results.AutoArk-AI-ARK-ASR-3B_20260622_official', 'AutoArk-AI/ARK-ASR-3B', ) PY Important:… See the full description on the dataset page: https://huggingface.co/datasets/Edge0/ark-asr-3b-open-asr-leaderboard-results.tabularautomatic-speech-recognition10K<n<100K12 likes184 downloads3mo agoHugging Face03google /red_ace_asr_error_detection_and_correction RED-ACE Dataset Summary This dataset can be used to train and evaluate ASR Error Detection or Correction models. It was introduced in the RED-ACE paper (Gekhman et al, 2022). The dataset contains ASR outputs on the LibriSpeech corpus (Panayotov et al., 2015) with annotated transcription errors. Dataset Details The LibriSpeech corpus was decoded using Google Cloud Speech-to-Text API, with the default and video models. The word-level confidence was enabled… See the full description on the dataset page: https://huggingface.co/datasets/google/red_ace_asr_error_detection_and_correction.textautomatic-speech-recognition100K<n<1M6 likes126 downloads3y agoHugging Face04ansh-rohilla /verbalyze-stt-bench Verbalyze: Indic Speech & ITN Benchmark (12 Languages) Verbalyze is a scenario-weighted benchmark dataset designed to evaluate and train Speech-to-Text (ASR) and Inverse Text Normalization (ITN) models on real-world Indic speech phenomena. It covers 12 major Indian languages with 172,800 balanced utterances categorized across 11 edge-case scenarios where standard speech models typically fail. Languages Covered Language Code Samples Script Assamese as 14… See the full description on the dataset page: https://huggingface.co/datasets/ansh-rohilla/verbalyze-stt-bench.textautomatic-speech-recognition100K<n<1M1 likes108 downloads12d agoHugging Face05NbAiLab /lunde_nor_nob_reading_optimisedTest only - not for training. First version - 0.1 of lunde_nor_nob_reading_optimised This dataset does not contain any audio data. Export Details Train samples: 10040932 Validation samples: 0 Test samples: 0 Dataset created using search datasets:lunde_nor_nob_reading_optimised. textautomatic-speech-recognition10M<n<100M0 likes97 downloads8mo agoHugging Face06rustam1221 /uzbek-asr-train-manifests Uzbek ASR Training Manifests The exact training, validation and test splits behind rustam1221/uzbek-asr-gigaam: 974 hours of Uzbek speech drawn from seven public corpora, filtered, text-normalized, and split by speaker. No audio is copied. Each row is a pointer — a parquet file plus a row index in the upstream dataset — and the training dataloader decodes the audio when the batch is built. That keeps the whole corpus definition at 200 MB instead of roughly a terabyte of… See the full description on the dataset page: https://huggingface.co/datasets/rustam1221/uzbek-asr-train-manifests.textautomatic-speech-recognition1K<n<10K0 likes58 downloads22d agoHugging Face07danieldzikunuofmarvel /bibletts-asante-twi-repaired BibleTTS Asante Twi — Repaired Transcripts The Asante Twi transcripts released with BibleTTS have had the characters ɛ (U+025B) and ɔ (U+0254) stripped out. This dataset restores them. Audio is not included. This is a drop-in replacement for the .txt files that ship with the BibleTTS Asante Twi package, matched by clip ID. The problem Both are Twi vowels, and both are required by the orthography. Measured across the released Asante Twi transcripts: Character… See the full description on the dataset page: https://huggingface.co/datasets/danieldzikunuofmarvel/bibletts-asante-twi-repaired.tabularautomatic-speech-recognition10K<n<100K0 likes35 downloads2mo agoHugging Face08abdo1819 /arabic-english-code-switching-review-annotations Review Annotations for Arabic-English Code-Switching Speech This metadata-only dataset publishes review decisions and transcript-correction deltas for MohamedRashad/arabic-english-code-switching. It contains no human audio, no local file paths, no raw review notes, and no copies of unchanged upstream transcripts. The annotations are pinned to upstream revision 4a3bffc45219c35949470de32b8d4cb328b0ce11 and join by upstream_row_index. Coverage and outcomes The… See the full description on the dataset page: https://huggingface.co/datasets/abdo1819/arabic-english-code-switching-review-annotations.tabularautomatic-speech-recognition10K<n<100K0 likes32 downloads2mo agoHugging Face09Quran-Lab /quranic-asr-cloud-rawdatagated Quranic ASR Provider Benchmark Results Professional benchmark artifacts for comparing commercial and official ASR providers on the Quranic ASR benchmark hosted at Quran-Lab/quranic-asr-benchmark. This repository contains metadata, normalized result tables, raw provider responses, unchanged run scripts, scoring outputs, Tarteel streaming probes, and reports. It does not duplicate the source audio. What Is Included Area Path Purpose Benchmark split… See the full description on the dataset page: https://huggingface.co/datasets/Quran-Lab/quranic-asr-cloud-rawdata.tabularautomatic-speech-recognition1K<n<10K1 likes31 downloads29d agoHugging Face10theblackcat102 /common-voice-en-revoicetextautomatic-speech-recognition10K<n<100K0 likes21 downloads3y agoHugging Face11Shawal777 /yogera_runyankore_ailab_4_0_1imageautomatic-speech-recognition1K<n<10K0 likes16 downloads2y agoHugging Face12Raiff1982 /evalgated Dataset Card for Dataset Name This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Dataset Details Dataset Description Curated by: jonathan harrison Funded by [optional]: [More Information Needed] Shared by [optional]: [More Information Needed] Language(s) (NLP): [More Information Needed] License: [More Information Needed] Dataset Sources [optional] Repository: [More Information… See the full description on the dataset page: https://huggingface.co/datasets/Raiff1982/eval.texttext-classificationn<1K0 likes14 downloads1y agoHugging Face13Shawal777 /yogera_runyankore_ailabimageautomatic-speech-recognition1K<n<10K0 likes10 downloads2y agoHugging Face14gallip0li /medimind-r11-traingated MediMind R11 — ASR training data Unified manifest + packed audio for fine-tuning Whisper-large-v3 on Norwegian clinical and conversational speech. Training manifest: r11_manifest.jsonl — 11,022 packs Held-out eval set: r11_heldout_eval.jsonl — 291 packs (NEVER train on these) ~see manifest audit packs total 11 sources: lege_*, podcasts (motiv/podk/stet), nb_samtale, nb_tale_m3, tts_drugs Schema See r11_manifest.jsonl (one JSON object per line) and… See the full description on the dataset page: https://huggingface.co/datasets/gallip0li/medimind-r11-train.audioautomatic-speech-recognitionn<1K0 likes1 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.