CoolFace
28 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01facebook /omnilingual-asr-corpus Meta Omnilingual ASR Corpus The Omnilingual ASR Corpus is a collection of spontaneous speech recordings and their transcriptions for 348 under-served languages. The corpus was collected as part of Meta FAIR’s Omnilingual ASR project (blog, model, paper) for the purposes of training automatic speech recognition (ASR) and spoken language identification models. Data schema { `language`: "lij_Latn", `iso_639_3`: "lij", `iso_15924`: "Latn", `glottocode`:… See the full description on the dataset page: https://huggingface.co/datasets/facebook/omnilingual-asr-corpus.audioautomatic-speech-recognition100K<n<1M211 likes4.9k downloads11mo agoHugging Face02vnahata /OmnilingualASR-retrieval Omnilingual ASR speech-text retrieval (MTEB) Read speech paired with its human transcription, for languages that no existing MTEB audio task covers. Source: facebook/omnilingual-asr-corpus at revision 8648ba8, cc-by-4.0, official test split. Recordings are re-encoded from FLAC to Opus at 16 kHz. Repeated transcripts are dropped, since one would otherwise be relevant to several recordings while only one is marked correct. Built by… See the full description on the dataset page: https://huggingface.co/datasets/vnahata/OmnilingualASR-retrieval.audioautomatic-speech-recognition1K<n<10K0 likes1.2k downloads24d agoHugging Face03KathleenKunLiu /omnilingual-asr-corpus Meta Omnilingual ASR Corpus The Omnilingual ASR Corpus is a collection of spontaneous speech recordings and their transcriptions for 348 under-served languages. The corpus was collected as part of Meta FAIR’s Omnilingual ASR project (blog, model, paper) for the purposes of training automatic speech recognition (ASR) and spoken language identification models. Data schema { `language`: "lij_Latn", `iso_639_3`: "lij", `iso_15924`: "Latn"… See the full description on the dataset page: https://huggingface.co/datasets/KathleenKunLiu/omnilingual-asr-corpus.audioautomatic-speech-recognition100K<n<1M0 likes787 downloads3mo agoHugging Face04Sanghyang00 /omniasr-molge OmniASR Molge Aligned Training-friendly re-segmentation of Meta’s facebook/omnilingual-asr-corpus: long utterances are segmented / aligned into ≤30s clips with transcripts, then packed as Parquet shards with embedded FLAC. Source facebook/omnilingual-asr-corpus Configs omniasr_aligned_v1, omniasr_aligned_v2 Splits train / validation (dev-*.parquet) / test Scale ~2.56M utts · ~839 shards · ~439GB If this dataset is useful for your work, we’d appreciate a… See the full description on the dataset page: https://huggingface.co/datasets/Sanghyang00/omniasr-molge.tabularautomatic-speech-recognition1M<n<10M0 likes553 downloads2mo agoHugging Face05SynDataLab-EN /omnivoice-tr Sample rate 24 kHz Voice-designed 10,000 Voice-cloned 10,000 Total 20,000 audiotext-to-speech10K<n<100K0 likes340 downloads6mo agoHugging Face06SynDataLab-EN /omnivoice-fr Sample rate 24 kHz Voice-designed 10,000 Voice-cloned 10,000 Total 20,000 audiotext-to-speech10K<n<100K0 likes338 downloads6mo agoHugging Face07SynDataLab-EN /omnivoice-th Sample rate 24 kHz Voice-designed 10,000 Voice-cloned 9,833 Total 19,833 audiotext-to-speech10K<n<100K1 likes333 downloads6mo agoHugging Face08OmniEvalKit /omnievalkit-dataset OmniEvalKit Evaluation Datasets Evaluation datasets for OmniEvalKit, a comprehensive evaluation framework for omni-modal (audio + video + image + text) models. Overview Total subsets: 65 Total samples: 315,264 Total size: 620.3 GB (Parquet with embedded audio/image/video) Subsets with embedded video: 15 Subsets requiring external video download: 2 Usage from datasets import load_dataset ds = load_dataset("OmniEvalKit/omnievalkit-dataset", "aishell1_test")… See the full description on the dataset page: https://huggingface.co/datasets/OmniEvalKit/omnievalkit-dataset.audioaudio-classification100K<n<1M0 likes329 downloads6mo agoHugging Face09SynDataLab-EN /omnivoice-zh Sample rate 24 kHz Voice-designed 10,000 Voice-cloned 9,946 Total 19,946 audiotext-to-speech10K<n<100K0 likes282 downloads6mo agoHugging Face10SynDataLab-EN /omnivoice-it Sample rate 24 kHz Voice-designed 10,000 Voice-cloned 10,000 Total 20,000 audiotext-to-speech10K<n<100K0 likes281 downloads6mo agoHugging Face11SynDataLab-EN /omnivoice-es Sample rate 24 kHz Voice-designed 10,000 Voice-cloned 10,000 Total 20,000 audiotext-to-speech10K<n<100K0 likes230 downloads6mo agoHugging Face12SynDataLab-EN /omnivoice-ja Sample rate 24 kHz Voice-designed 10,000 Voice-cloned 9,958 Total 19,958 audiotext-to-speech10K<n<100K0 likes177 downloads6mo agoHugging Face13SynDataLab-EN /omnivoice-ru Sample rate 24 kHz Voice-designed 10,000 Voice-cloned 9,999 Total 19,999 audiotext-to-speech10K<n<100K0 likes162 downloads6mo agoHugging Face14Goekdeniz-Guelmez /mlx-omni-lora-stt-tts-demoWill be used in the development of the trainer backend of mlx-omni by Neywa Labs. audioautomatic-speech-recognitionn<1K0 likes110 downloads1mo agoHugging Face15lindonghello /omnilingual-asr-corpus Meta Omnilingual ASR Corpus The Omnilingual ASR Corpus is a collection of spontaneous speech recordings and their transcriptions for 348 under-served languages. The corpus was collected as part of Meta FAIR’s Omnilingual ASR project (blog, model, paper) for the purposes of training automatic speech recognition (ASR) and spoken language identification models. Data schema { `language`: "lij_Latn", `iso_639_3`: "lij", `iso_15924`: "Latn", `glottocode`:… See the full description on the dataset page: https://huggingface.co/datasets/lindonghello/omnilingual-asr-corpus.audioautomatic-speech-recognition100K<n<1M0 likes96 downloads5mo agoHugging Face16Sellopale /omnilingualpaleoi Meta Omnilingual ASR Corpus The Omnilingual ASR Corpus is a collection of spontaneous speech recordings and their transcriptions for 348 under-served languages. The corpus was collected as part of Meta FAIR’s Omnilingual ASR project (blog, model, paper) for the purposes of training automatic speech recognition (ASR) and spoken language identification models. Data schema { `language`: "lij_Latn", `iso_639_3`: "lij", `iso_15924`: "Latn"… See the full description on the dataset page: https://huggingface.co/datasets/Sellopale/omnilingualpaleoi.audioautomatic-speech-recognition100K<n<1M2 likes86 downloads9mo agoHugging Face17Chiz /omniASR-igbo-blindspots omniASR Igbo Blind Spot Dataset Research Questions This dataset investigates three interrelated questions about multilingual ASR performance on tonal languages: Operational Definition: What does "language support" mean when a model lists 1,600+ languages? Does coverage imply functional accuracy on linguistically meaningful distinctions? Diagnostic Validity: Can tonal diacritic preservation serve as a diagnostic for acoustic competence vs. orthographic pattern matching… See the full description on the dataset page: https://huggingface.co/datasets/Chiz/omniASR-igbo-blindspots.audioautomatic-speech-recognitionn<1K1 likes67 downloads7mo agoHugging Face18xiaofff /omnievalkit-data-test OmniEvalKit Evaluation Datasets Evaluation datasets for OmniEvalKit, a comprehensive evaluation framework for omni-modal (audio + video + image + text) models. Overview Total subsets: 89 Total samples: 353,610 Total size: 352.3 GB (Parquet with embedded audio/image, no video) Subsets requiring video download: 42 Note: Video files are NOT embedded in the Parquet files due to size constraints. Usage from datasets import load_dataset ds =… See the full description on the dataset page: https://huggingface.co/datasets/xiaofff/omnievalkit-data-test.audioaudio-classification10K<n<100K0 likes33 downloads6mo agoHugging Face19Omniloy /fonendo-benchgated fonendo-bench: clinical subsets Spanish clinical dictation for evaluating speech-to-text systems. This repository holds the two clinical subsets of fonendo-bench, a reproducible Spanish clinical speech-to-text benchmark published by Omniloy. The benchmark code, the tools that rebuild the public (real speech) subsets and the leaderboard live in the GitHub repository; this dataset contains only the clinical audio, its reference transcripts and its metadata. config clips… See the full description on the dataset page: https://huggingface.co/datasets/Omniloy/fonendo-bench.audioautomatic-speech-recognitionn<1K0 likes26 downloads10d agoHugging Face20SynDataLab-EN /omnivoice-de Sample rate 24 kHz Voice-designed 10,000 Voice-cloned 10,000 Total 20,000 audiotext-to-speech10K<n<100K0 likes8 downloads6mo agoHugging Face21ShiniChien /OmniEdugated Overview This dataset is being developed for training and evaluating omni-modal speech models, with a primary focus on audio-to-audio tasks. The repository is part of an ongoing research effort to build high-quality speech interaction datasets for next-generation conversational AI systems. The dataset is still under active development. Data organization, validation, annotation, and quality control are continuously being improved before a stable public release.… See the full description on the dataset page: https://huggingface.co/datasets/ShiniChien/OmniEdu.audioautomatic-speech-recognition10K<n<100K1 likes8 downloads3mo agoHugging Face22Harshkmr /omniscribe_corpus OmniScribe Corpus A multilingual speech transcription corpus designed for fine-tuning ASR models on Indian medical and general-domain speech. It covers Hindi, Marathi, and Indian English, with a focus on clinical and healthcare contexts. Overview Split Rows (after oversampling) Approx. Duration train ~30750 ~230 hrs benchmark ~4,089 ~25 hrs Audio samples average 20–30 seconds each. All samples are at least 5 seconds… See the full description on the dataset page: https://huggingface.co/datasets/Harshkmr/omniscribe_corpus.audioautomatic-speech-recognition10K<n<100K0 likes7 downloads5mo agoHugging Face23instinct-org /omni_chunked_speech_restorisedgated omni_chunked_speech_restorised This is a gated Uzbek speech-restorised chunked speech dataset from instinct-org. This repository contains speech audio and transcripts for speech-to-text training, evaluation, or data preparation workflows. Language Primary language: uz (Uzbek) Intended Use speech-to-text training and evaluation Internal dataset curation, quality checks, and model evaluation Research or commercial use only after access approval and license… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/omni_chunked_speech_restorised.audioautomatic-speech-recognition1K<n<10K0 likes7 downloads4mo agoHugging Face24Batuka0901 /omni_source Omni source dataset Mongolian audio/text pairs used as source material for the omni training set. Dataset Statistics Total samples: 69 Total duration: 0h 15m 47s (0.26 h) Per-split breakdown Split Samples Total Duration Avg Duration train 69 0h 15m 47s (0.26 h) 13.73 s Single split (no held-out test set) -- this is source/reference data, not a benchmark split. audioautomatic-speech-recognitionn<1K0 likes6 downloads4mo agoHugging Face25toiar /Khasi-OmniVoice-TTS-Datagated Khasi Omni Voice TTS Dataset The Khasi Omni Voice dataset is a comprehensive, high-quality audio collection designed specifically for Text-to-Speech (TTS) research and model training in the Khasi language. It features nearly 50 hours of speech data targeting realistic, modern Khasi speech patterns, including natural code-switching. Key Statistics Total Duration: 49 hours, 53 minutes, 19.98 seconds Total Samples: 18,874 distinct audio utterances Language: Khasi… See the full description on the dataset page: https://huggingface.co/datasets/toiar/Khasi-OmniVoice-TTS-Data.audiotext-to-speech10K<n<100K0 likes6 downloads4mo agoHugging Face26ShiniChien /OmniDistilgated Overview This dataset is being developed for training and evaluating omni-modal speech models, with a primary focus on audio-to-audio tasks. The repository is part of an ongoing research effort to build high-quality speech interaction datasets for next-generation conversational AI systems. The dataset is still under active development. Data organization, validation, annotation, and quality control are continuously being improved before a stable public release.… See the full description on the dataset page: https://huggingface.co/datasets/ShiniChien/OmniDistil.audioautomatic-speech-recognition100K<n<1M1 likes5 downloads3mo agoHugging Face27instinct-org /omni_chunkedgated omni_chunked This is a gated Uzbek chunked speech dataset from instinct-org. This repository contains speech audio and transcripts for speech-to-text training, evaluation, or data preparation workflows. Language Primary language: uz (Uzbek) Intended Use speech-to-text training and evaluation Internal dataset curation, quality checks, and model evaluation Research or commercial use only after access approval and license review Data Notes Contains… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/omni_chunked.audioautomatic-speech-recognition1K<n<10K0 likes4 downloads4mo agoHugging Face28instinct-org /omni_chunked_nfa_alignedgated Forced-Aligned STT Dataset Source dataset: instinct-org/omni_chunked Aligned dataset: instinct-org/omni_chunked_nfa_aligned Rows: 2699 successfully aligned rows This dataset adds NeMo Forced Aligner metadata for STT training and timestamp quality control. Source rows that did not produce usable CTM alignments are excluded from the published data. Added columns: nfa_token_alignments: NeMo token/subword CTM spans nfa_word_alignments: word-level CTM spans nfa_segment_alignments:… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/omni_chunked_nfa_aligned.textautomatic-speech-recognition1K<n<10K0 likes3 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.