CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01huseyin-karaca /hit-asrtabular1M<n<10M0 likes8.5k downloads5d agoHugging Face02Reza2kn /persian-asr-audio-text-2.69M-chizzled 🗂️ persian-asr-audio-text-2.69M-chizzled English + فارسی · Part of Shenava 1.0 · Project hub · SLT paper submission 🌟 At a glance | معرفی سریع English فارسی 🎯 Purpose Phase A-scale audio/text dataset. پیکرهٔ بزرگ جفت‌های صوت و متنِ پالایش‌شده برای آموزش در مقیاس فاز A. 🧩 Role Persian text and linguistic asset مصنوع متنی و زبانی فارسی 📦 Snapshot 417 files; approximately 236.86 GB 417 فایل؛ حدود 236.86 GB 🧱 Packaging 414 Parquet files and 0… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/persian-asr-audio-text-2.69M-chizzled.tabular1M<n<10M2 likes4.7k downloads2mo agoHugging Face03syvai /danish-asr-unified-hviske-v5-tinygated danish-asr-unified — two-model labels and a quality manifest Transcriptions, per-token confidences, and a per-row quality verdict for every row of syvai/danish-asr-unified (3,414,589 rows, 8 sources). Two independently trained models labelled the whole corpus: model architecture vocabulary syvai/hviske-v5-tiny encoder-decoder 16,384 BPE 3dio-ai/svale-110M RNN-T (Parakeet) 44 characters Both decode greedily. Shards mirror the source data/train-*.parquet by name… See the full description on the dataset page: https://huggingface.co/datasets/syvai/danish-asr-unified-hviske-v5-tiny.tabularautomatic-speech-recognition1M<n<10M0 likes1.9k downloads9d agoHugging Face04RyeAI /danish-asr-leaderboard Open Danish ASR Leaderboard — Results Benchmark results backing the Open Danish ASR Leaderboard — an open, reproducible comparison of Danish speech-to-text models. Every model is transcribed and scored identically on the same five independent public Danish test sets, so the numbers compare directly: Word Error Rate (WER) and Character Error Rate (CER) — lower is better — plus speed. Open-weight Danish speech recognition models you can run yourself and hosted transcription APIs… See the full description on the dataset page: https://huggingface.co/datasets/RyeAI/danish-asr-leaderboard.tabularautomatic-speech-recognition1M<n<10M4 likes1.8k downloads21m agoHugging Face05anuj-inavlabs /kupe-asr-en-data kupe-asr-en-mini-150m — data Two loadable configs, packed into ~20-25 bunch_*.parquet files each (Hub-quota friendly): raw — 24 kHz mono English audio (flac bytes) + text. The encode stage reads this. mimi — Mimi c0..c7 codes (12.5 Hz) + text. Training reads this. Ledgers under ledger/ (data.json, mimi.json) track collected/encoded hours and resume state. from datasets import load_dataset ds = load_dataset("anuj-inavlabs/kupe-asr-en-data", "mimi", split="train") tabularautomatic-speech-recognition1M<n<10M0 likes1k downloads16d agoHugging Face06addy88 /sanskrit-asr-84tabular10K<n<100K0 likes545 downloads5y agoHugging Face07syvai /danish-asr-verified danish-asr-verified ALL rows of syvai/danish-asr-unified transcribed by the syv-transcribe ensemble (hviske-v5.3 + hviske-v5, confidence-weighted ROVER), each annotated with: verified — True when the ensemble independently reproduced the reference exactly (compared after lowercasing, punctuation-strip, whitespace-collapse). Two independent witnesses agree => near-certain label. wer_teacher_vs_ref / cer_teacher_vs_ref — word/character error rate between normalized teacher output… See the full description on the dataset page: https://huggingface.co/datasets/syvai/danish-asr-verified.tabularautomatic-speech-recognition1M<n<10M0 likes455 downloads2mo agoHugging Face08treble-technologies /librispeech_asr_sliced Librispeech Slices Description Librispeech is a large corpus of read English utterances derived from the LibriVox public domain audiobook project. It was assembled to assist in Automatic Speech Recognition tasks, and contains approximately 1000 Hours of utterances recorded at 16kHz. A subset of the original Librispeech dataset was created to support the development of automatic audio scene creation in the Treble SDK environment. To better imitate the natural flow of… See the full description on the dataset page: https://huggingface.co/datasets/treble-technologies/librispeech_asr_sliced.tabular100K<n<1M0 likes335 downloads7mo agoHugging Face09Atika88 /Indonesian-ASR-11-Class-Dataset Indonesian ASR 11-Class Dataset Public Hugging Face repository for an Indonesian ASR corpus and its paper-supporting benchmark artifacts. Dataset summary Audio files: 104,500 WAV files Real/human recordings: 104,368 Synthetic repair files: 132 Sentence classes: 11 Indonesian sentence categories Canonical balanced sentence slots: 209 (11 categories × 19 retained slots) Public speaker labels: M1..M12, F1..F8, plus synthetic labels Ms*/Fs* Audio format: 16 kHz… See the full description on the dataset page: https://huggingface.co/datasets/Atika88/Indonesian-ASR-11-Class-Dataset.tabularautomatic-speech-recognition100K<n<1M0 likes306 downloads16d agoHugging Face10wuff-mann /ASR-CYGNSS-HGMM-Reproducibility ASR CYGNSS-SMAP H-GMM reproducibility repository This public dataset repository stores derived CYGNSS-SMAP collocations, model checkpoints, evaluation statistics, and manuscript figures for the manuscript on year-adaptive CYGNSS sea-surface wind retrieval. Upstream public data CYGNSS Level-2 Science Data Record v3.2 (CYGNSS_L2_V3.2), NASA PO.DAAC, DOI: 10.5067/CYGNS-L2X32. JPL SMAP Level-2B CAP Sea Surface Salinity and extreme-wind product v5.0… See the full description on the dataset page: https://huggingface.co/datasets/wuff-mann/ASR-CYGNSS-HGMM-Reproducibility.tabulartabular-regression10M<n<100M0 likes276 downloads2mo agoHugging Face11AS-Robotics /robot-meet-gemma-record-medicineThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "so101_follower", "total_episodes": 13, "total_frames": 6119, "total_tasks": 1, "total_videos": 26, "total_chunks": 1, "chunks_size": 1000, "fps": 25, "splits": { "train": "0:13" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/AS-Robotics/robot-meet-gemma-record-medicine.tabularrobotics10K<n<100K0 likes262 downloads1y agoHugging Face12addy88 /sanskrit-asr-84-evaltabular1K<n<10K1 likes225 downloads5y agoHugging Face13jq /salt-asr-data-transcriptionstabular10K<n<100K0 likes209 downloads2y agoHugging Face14AS-Robotics /Pick-and-Place-dataset_so100_pickplaceThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "so100_follower", "total_episodes": 10, "total_frames": 3764, "total_tasks": 1, "total_videos": 20, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:10" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/AS-Robotics/Pick-and-Place-dataset_so100_pickplace.tabularrobotics10K<n<100K0 likes208 downloads1y agoHugging Face15sanchit-gandhi /librispeech_asr_dummy Dataset Card for librispeech_asr_dummy Dataset Summary This is a truncated version of the LibriSpeech dataset. It contains 20 samples from each of the splits. To view the full dataset, visit: https://huggingface.co/datasets/librispeech_asr LibriSpeech is a corpus of approximately 1000 hours of 16kHz read English speech, prepared by Vassil Panayotov with the assistance of Daniel Povey. The data is derived from read audiobooks from the LibriVox project, and has been… See the full description on the dataset page: https://huggingface.co/datasets/sanchit-gandhi/librispeech_asr_dummy.audioautomatic-speech-recognitionn<1K0 likes193 downloads3y agoHugging Face16frankie137 /sd_asr_synthesis_datatabularn<1K0 likes160 downloads3mo agoHugging Face17sonalsannigrahi /voxpopuli_asr_norm_curatortabular100K<n<1M0 likes151 downloads3mo agoHugging Face18cmu-mlsp /librispeech960-encodec1024_asr Dataset Card for "librispeech960-encodec1024_asr" More Information needed tabular100K<n<1M0 likes120 downloads3y agoHugging Face19asr-malayalam /indicvoices-v1atabular10K<n<100K0 likes115 downloads2y agoHugging Face20frankie137 /sd_asr_synthesis_data_v0_less_silencetabularn<1K0 likes112 downloads3mo agoHugging Face21pere /nb-asr-numerics-harvested Norwegian Bokmål Numeric Expression Harvesting Dataset This dataset contains cleaned, high-quality Norwegian Bokmål sentences containing numeric expressions harvested from both the Norwegian Colossal Corpus (NbAiLab/NCC) and the FineWeb-2 Norwegian Bokmål subset (HuggingFaceFW/fineweb-2 config nob_Latn). This is a combined high-volume intermediate dataset built for the first stage of a template-based Norwegian synthetic speech (TTS) generation pipeline to improve number/digit… See the full description on the dataset page: https://huggingface.co/datasets/pere/nb-asr-numerics-harvested.tabulartext-generation1M<n<10M0 likes101 downloads3mo agoHugging Face22cmu-mlsp /librispeech960-wavlm-large-km1000_asr Dataset Card for "librispeech960-wavlm-large-km1000_asr" More Information needed tabular100K<n<1M0 likes95 downloads3y agoHugging Face23nyalpatel /condensed_librispeech_asr Condensed LibriSpeech ASR This dataset is a condensed version of the LibriSpeech ASR dataset, created by subsampling approximately 10% of the original data from each split. It is intended for quick experimentation, prototyping, and debugging when working with Automatic Speech Recognition (ASR) tasks. Dataset Details Original Dataset: LibriSpeech ASR Condensation Ratio: Approximately 10% of the full dataset Splits Included: train.clean.100 train.clean.360 train.other.500… See the full description on the dataset page: https://huggingface.co/datasets/nyalpatel/condensed_librispeech_asr.tabular1K<n<10K0 likes95 downloads2y agoHugging Face24humanify /real_data_sd_asrtabularn<1K0 likes84 downloads3mo agoHugging Face25nhatminh /korean-asr korean-asr — Korean ASR pseudo-labels for YODAS2 This repository contains transcripts and segment metadata only. It does not contain audio. Every row points into espnet/yodas2 by (shard, video_id, start, end), so you fetch the audio from the upstream dataset and cut it yourself. See Reconstructing the audio. split utterances hours train 1,034,181 6,974.3 heldout 46,542 314.6 dev (subset of heldout) 3,000 20.4 Total labelled: 1,080,723 utterances / 7,288.9… See the full description on the dataset page: https://huggingface.co/datasets/nhatminh/korean-asr.tabularautomatic-speech-recognition1M<n<10M0 likes71 downloads23d agoHugging Face26Reza2kn /persian-asr-text-2.69M-deduped 🗂️ persian-asr-text-2.69M-deduped English + فارسی · Part of Shenava 1.0 · Project hub · SLT paper submission 🌟 At a glance | معرفی سریع English فارسی 🎯 Purpose Deduplicated Persian ASR text dataset used by the training stack. پیکرهٔ متنی فارسیِ حذف‌تکرارشده برای ساخت واژگان، مدل‌سازی زبانی و پشتیبانی از آموزش ASR. 🧩 Role Persian text and linguistic asset مصنوع متنی و زبانی فارسی 📦 Snapshot 4 files; approximately 109.64 MB 4 فایل؛ حدود 109.64… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/persian-asr-text-2.69M-deduped.tabularautomatic-speech-recognition1M<n<10M0 likes67 downloads2mo agoHugging Face27Vikhrmodels /russian-asr-leaderboardtabularn<1K0 likes61 downloads1y agoHugging Face28Peacockery /tajik-asr-youtube tajik-asr-youtube Every Tajik-language clip from this project's YouTube scrape — 41 channels of news, talk shows, podcasts, audiobooks, and learning content — with machine transcripts and the verification scores left in as columns instead of applied as a filter. Pick your own quality threshold; the training corpus this project actually ships (tajik-asr-corpus-v3) is the gated subset. Layout Parquet shards under data/ with an audio struct column (16 kHz mono FLAC… See the full description on the dataset page: https://huggingface.co/datasets/Peacockery/tajik-asr-youtube.tabularautomatic-speech-recognition100K<n<1M0 likes59 downloads3mo agoHugging Face29Navana-AI /indic-asr-benchmark Indic ASR Benchmark — Nine Languages Speech, human references, and side-by-side transcripts from multiple speech-to-text systems across nine Indian languages — the evaluation data behind Navana's public Bodhi ASR benchmark. Every clip comes from openly available research datasets, and every system is scored the same way, on the same audio. Released by Navana Tech, the team behind Bodhi, our Indian-language speech-to-text engine. A companion Hindi-only benchmark, with the full… See the full description on the dataset page: https://huggingface.co/datasets/Navana-AI/indic-asr-benchmark.audio10K<n<100K1 likes52 downloads2mo agoHugging Face30Aryan95614 /aerograph-asrs AeroGraph ASRS Dataset 2,000 real NASA Aviation Safety Reporting System (ASRS) incident reports with LLM-extracted entities and relations for knowledge graph construction. Dataset Description This dataset contains processed ASRS incident narratives along with structured entity and relation extractions conforming to an aviation safety ontology (10 entity types, 8 edge types). Reports Split 2000 reports from the NASA ASRS database Fields: id, text, aircraft_type… See the full description on the dataset page: https://huggingface.co/datasets/Aryan95614/aerograph-asrs.tabularquestion-answering1K<n<10K1 likes35 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.