datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
omnilingual-asr-corpus
Meta Omnilingual ASR Corpus
The Omnilingual ASR Corpus is a collection of spontaneous speech recordings and their transcriptions for 348 under-served languages. The corpus was collected as part of Meta FAIR’s Omnilingual ASR project (blog, model, paper) for the purposes of training automatic speech recognition (ASR) and spoken language identification models.
Data schema
{
`language`: "lij_Latn",
`iso_639_3`: "lij",
`iso_15924`: "Latn",
`glottocode`:… See the full description on the dataset page: https://huggingface.co/datasets/facebook/omnilingual-asr-corpus.OmnilingualASR-retrieval
Omnilingual ASR speech-text retrieval (MTEB)
Read speech paired with its human transcription, for languages that no existing
MTEB audio task covers.
Source: facebook/omnilingual-asr-corpus at revision 8648ba8, cc-by-4.0, official
test split. Recordings are re-encoded from FLAC to Opus at 16 kHz. Repeated
transcripts are dropped, since one would otherwise be relevant to several
recordings while only one is marked correct.
Built by… See the full description on the dataset page: https://huggingface.co/datasets/vnahata/OmnilingualASR-retrieval.omnilingual-asr-corpus
Meta Omnilingual ASR Corpus
The Omnilingual ASR Corpus is a collection of spontaneous speech recordings and their transcriptions for 348 under-served languages. The corpus was collected as part of Meta FAIR’s Omnilingual ASR project (blog, model, paper) for the purposes of training automatic speech recognition (ASR) and spoken language identification models.
Data schema
{
`language`: "lij_Latn",
`iso_639_3`: "lij",
`iso_15924`: "Latn"… See the full description on the dataset page: https://huggingface.co/datasets/KathleenKunLiu/omnilingual-asr-corpus.omniasr-molge
OmniASR Molge Aligned
Training-friendly re-segmentation of Meta’s facebook/omnilingual-asr-corpus: long utterances are segmented / aligned into ≤30s clips with transcripts, then packed as Parquet shards with embedded FLAC.
Source
facebook/omnilingual-asr-corpus
Configs
omniasr_aligned_v1, omniasr_aligned_v2
Splits
train / validation (dev-*.parquet) / test
Scale
~2.56M utts · ~839 shards · ~439GB
If this dataset is useful for your work, we’d appreciate a… See the full description on the dataset page: https://huggingface.co/datasets/Sanghyang00/omniasr-molge.omnivoice-tr
Sample rate
24 kHz
Voice-designed
10,000
Voice-cloned
10,000
Total
20,000
omnivoice-fr
Sample rate
24 kHz
Voice-designed
10,000
Voice-cloned
10,000
Total
20,000
omnivoice-th
Sample rate
24 kHz
Voice-designed
10,000
Voice-cloned
9,833
Total
19,833
omnievalkit-dataset
OmniEvalKit Evaluation Datasets
Evaluation datasets for OmniEvalKit,
a comprehensive evaluation framework for omni-modal (audio + video + image + text) models.
Overview
Total subsets: 65
Total samples: 315,264
Total size: 620.3 GB (Parquet with embedded audio/image/video)
Subsets with embedded video: 15
Subsets requiring external video download: 2
Usage
from datasets import load_dataset
ds = load_dataset("OmniEvalKit/omnievalkit-dataset", "aishell1_test")… See the full description on the dataset page: https://huggingface.co/datasets/OmniEvalKit/omnievalkit-dataset.omnivoice-zh
Sample rate
24 kHz
Voice-designed
10,000
Voice-cloned
9,946
Total
19,946
omnivoice-it
Sample rate
24 kHz
Voice-designed
10,000
Voice-cloned
10,000
Total
20,000
omnivoice-es
Sample rate
24 kHz
Voice-designed
10,000
Voice-cloned
10,000
Total
20,000
omnivoice-ja
Sample rate
24 kHz
Voice-designed
10,000
Voice-cloned
9,958
Total
19,958
omnivoice-ru
Sample rate
24 kHz
Voice-designed
10,000
Voice-cloned
9,999
Total
19,999
mlx-omni-lora-stt-tts-demoWill be used in the development of the trainer backend of mlx-omni by Neywa Labs.
omnilingual-asr-corpus
Meta Omnilingual ASR Corpus
The Omnilingual ASR Corpus is a collection of spontaneous speech recordings and their transcriptions for 348 under-served languages. The corpus was collected as part of Meta FAIR’s Omnilingual ASR project (blog, model, paper) for the purposes of training automatic speech recognition (ASR) and spoken language identification models.
Data schema
{
`language`: "lij_Latn",
`iso_639_3`: "lij",
`iso_15924`: "Latn",
`glottocode`:… See the full description on the dataset page: https://huggingface.co/datasets/lindonghello/omnilingual-asr-corpus.omnilingualpaleoi
Meta Omnilingual ASR Corpus
The Omnilingual ASR Corpus is a collection of spontaneous speech recordings and their transcriptions for 348 under-served languages. The corpus was collected as part of Meta FAIR’s Omnilingual ASR project (blog, model, paper) for the purposes of training automatic speech recognition (ASR) and spoken language identification models.
Data schema
{
`language`: "lij_Latn",
`iso_639_3`: "lij",
`iso_15924`: "Latn"… See the full description on the dataset page: https://huggingface.co/datasets/Sellopale/omnilingualpaleoi.omniASR-igbo-blindspots
omniASR Igbo Blind Spot Dataset
Research Questions
This dataset investigates three interrelated questions about multilingual ASR performance on tonal languages:
Operational Definition: What does "language support" mean when a model lists 1,600+ languages? Does coverage imply functional accuracy on linguistically meaningful distinctions?
Diagnostic Validity: Can tonal diacritic preservation serve as a diagnostic for acoustic competence vs. orthographic pattern matching… See the full description on the dataset page: https://huggingface.co/datasets/Chiz/omniASR-igbo-blindspots.omnievalkit-data-test
OmniEvalKit Evaluation Datasets
Evaluation datasets for OmniEvalKit,
a comprehensive evaluation framework for omni-modal (audio + video + image + text) models.
Overview
Total subsets: 89
Total samples: 353,610
Total size: 352.3 GB (Parquet with embedded audio/image, no video)
Subsets requiring video download: 42
Note: Video files are NOT embedded in the Parquet files due to size constraints.
Usage
from datasets import load_dataset
ds =… See the full description on the dataset page: https://huggingface.co/datasets/xiaofff/omnievalkit-data-test.fonendo-bench
fonendo-bench: clinical subsets
Spanish clinical dictation for evaluating speech-to-text systems. This repository holds the
two clinical subsets of fonendo-bench, a
reproducible Spanish clinical speech-to-text benchmark published by
Omniloy. The benchmark code, the tools that rebuild the public (real
speech) subsets and the leaderboard live in the GitHub repository; this dataset contains
only the clinical audio, its reference transcripts and its metadata.
config
clips… See the full description on the dataset page: https://huggingface.co/datasets/Omniloy/fonendo-bench.omnivoice-de
Sample rate
24 kHz
Voice-designed
10,000
Voice-cloned
10,000
Total
20,000
OmniEdu
Overview
This dataset is being developed for training and evaluating omni-modal speech models, with a primary focus on audio-to-audio tasks. The repository is part of an ongoing research effort to build high-quality speech interaction datasets for next-generation conversational AI systems.
The dataset is still under active development. Data organization, validation, annotation, and quality control are continuously being improved before a stable public release.… See the full description on the dataset page: https://huggingface.co/datasets/ShiniChien/OmniEdu.omniscribe_corpus
OmniScribe Corpus
A multilingual speech transcription corpus designed for fine-tuning ASR models on Indian medical and general-domain speech. It covers Hindi, Marathi, and Indian English, with a focus on clinical and healthcare contexts.
Overview
Split
Rows (after oversampling)
Approx. Duration
train
~30750
~230 hrs
benchmark
~4,089
~25 hrs
Audio samples average 20–30 seconds each. All samples are at least 5 seconds… See the full description on the dataset page: https://huggingface.co/datasets/Harshkmr/omniscribe_corpus.omni_chunked_speech_restorised
omni_chunked_speech_restorised
This is a gated Uzbek speech-restorised chunked speech dataset from instinct-org.
This repository contains speech audio and transcripts for speech-to-text training, evaluation, or data preparation workflows.
Language
Primary language: uz (Uzbek)
Intended Use
speech-to-text training and evaluation
Internal dataset curation, quality checks, and model evaluation
Research or commercial use only after access approval and license… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/omni_chunked_speech_restorised.omni_source
Omni source dataset
Mongolian audio/text pairs used as source material for the omni training set.
Dataset Statistics
Total samples: 69
Total duration: 0h 15m 47s (0.26 h)
Per-split breakdown
Split
Samples
Total Duration
Avg Duration
train
69
0h 15m 47s (0.26 h)
13.73 s
Single split (no held-out test set) -- this is source/reference data, not a
benchmark split.
Khasi-OmniVoice-TTS-Data
Khasi Omni Voice TTS Dataset
The Khasi Omni Voice dataset is a comprehensive, high-quality audio collection designed specifically for Text-to-Speech (TTS) research and model training in the Khasi language. It features nearly 50 hours of speech data targeting realistic, modern Khasi speech patterns, including natural code-switching.
Key Statistics
Total Duration: 49 hours, 53 minutes, 19.98 seconds
Total Samples: 18,874 distinct audio utterances
Language: Khasi… See the full description on the dataset page: https://huggingface.co/datasets/toiar/Khasi-OmniVoice-TTS-Data.OmniDistil
Overview
This dataset is being developed for training and evaluating omni-modal speech models, with a primary focus on audio-to-audio tasks. The repository is part of an ongoing research effort to build high-quality speech interaction datasets for next-generation conversational AI systems.
The dataset is still under active development. Data organization, validation, annotation, and quality control are continuously being improved before a stable public release.… See the full description on the dataset page: https://huggingface.co/datasets/ShiniChien/OmniDistil.omni_chunked
omni_chunked
This is a gated Uzbek chunked speech dataset from instinct-org.
This repository contains speech audio and transcripts for speech-to-text training, evaluation, or data preparation workflows.
Language
Primary language: uz (Uzbek)
Intended Use
speech-to-text training and evaluation
Internal dataset curation, quality checks, and model evaluation
Research or commercial use only after access approval and license review
Data Notes
Contains… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/omni_chunked.omni_chunked_nfa_aligned
Forced-Aligned STT Dataset
Source dataset: instinct-org/omni_chunked
Aligned dataset: instinct-org/omni_chunked_nfa_aligned
Rows: 2699 successfully aligned rows
This dataset adds NeMo Forced Aligner metadata for STT training and timestamp quality
control. Source rows that did not produce usable CTM alignments are excluded from the
published data.
Added columns:
nfa_token_alignments: NeMo token/subword CTM spans
nfa_word_alignments: word-level CTM spans
nfa_segment_alignments:… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/omni_chunked_nfa_aligned.
