datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
VaaniVAANI is an India-representative multi-modal multi-lingual dataset.
The current version (phase 1- 80 districts, phase 2- 85 districts) contains ~31278 hours of spontaenous,image-prompted speech by 156K speakers across 165 districts, talking about 288K images covering 105 languages.
From this audio data, 2,122 hours of transcribed data(text) is available, spanning almost evenly across the 165 districts.
Project Vaani, by IISc, Bangalore and ARTPARK, is capturing the true diversity of India’s… See the full description on the dataset page: https://huggingface.co/datasets/ARTPARK-IISc/Vaani.short_video_ocr_dataset
Short Video OCR / ASR Dataset
An actively curated research dataset for building OCR, ASR, subtitle-alignment,
and video-transcript pipelines for short social videos. It combines source
videos and extracted frames with human review artifacts and model-generated
text candidates. The primary languages are Ukrainian and Russian; English or
mixed-language content may also occur.
Status: work in progress. Model outputs and pseudo-label candidates are
not ground truth. Only… See the full description on the dataset page: https://huggingface.co/datasets/ElectronicHug/short_video_ocr_dataset.emova-alignment-7m
EMOVA-Alignment-7M
🤗 EMOVA-Models | 🤗 EMOVA-Datasets | 🤗 EMOVA-Demo
📄 Paper | 🌐 Project-Page | 💻 Github | 💻 EMOVA-Speech-Tokenizer-Github
Overview
EMOVA-Alignment-7M is a comprehensive dataset curated for omni-modal pre-training, including vision-language and speech-language alignment.
This dataset is created using open-sourced image-text pre-training datasets, OCR datasets, and 2,000 hours of ASR and TTS data.
This dataset is part of the EMOVA-Datasets… See the full description on the dataset page: https://huggingface.co/datasets/Emova-ollm/emova-alignment-7m.emova-sft-4m
EMOVA-SFT-4M
🤗 EMOVA-Models | 🤗 EMOVA-Datasets | 🤗 EMOVA-Demo
📄 Paper | 🌐 Project-Page | 💻 Github | 💻 EMOVA-Speech-Tokenizer-Github
Overview
EMOVA-SFT-4M is a comprehensive dataset curated for omni-modal instruction tuning, including textual, visual, and audio interactions. This dataset is created by gathering open-sourced multi-modal instruction datasets and synthesizing high-quality omni-modal conversation data to enhance user experience. This dataset is… See the full description on the dataset page: https://huggingface.co/datasets/Emova-ollm/emova-sft-4m.NEXUS-temporal_hierarchical_multi-modal
NEXUS: Neural Evolution for eXtensible Universal Semantics Dataset
(Temporal Multimodal Slices)
This dataset is a multi-modal, hierarchical, temporal representation derived from HuggingFaceFV/finevideo. It is designed for streaming training where the primary unit is a 10 ms "slice" that aggregates upward into moments (100 ms), seconds (1 s), experiences (10 s), and minutes (60 s).
It is meant to represent an extensible stream of "experience" as there are… See the full description on the dataset page: https://huggingface.co/datasets/Ardea/NEXUS-temporal_hierarchical_multi-modal.omnievalkit-dataset
OmniEvalKit Evaluation Datasets
Evaluation datasets for OmniEvalKit,
a comprehensive evaluation framework for omni-modal (audio + video + image + text) models.
Overview
Total subsets: 65
Total samples: 315,264
Total size: 620.3 GB (Parquet with embedded audio/image/video)
Subsets with embedded video: 15
Subsets requiring external video download: 2
Usage
from datasets import load_dataset
ds = load_dataset("OmniEvalKit/omnievalkit-dataset", "aishell1_test")… See the full description on the dataset page: https://huggingface.co/datasets/OmniEvalKit/omnievalkit-dataset.emova-sft-speech-231k
EMOVA-SFT-Speech-231K
🤗 EMOVA-Models | 🤗 EMOVA-Datasets | 🤗 EMOVA-Demo
📄 Paper | 🌐 Project-Page | 💻 Github | 💻 EMOVA-Speech-Tokenizer-Github
Overview
EMOVA-SFT-Speech-231K is a comprehensive dataset curated for omni-modal instruction tuning and emotional spoken dialogue. This dataset is created by converting existing text and visual instruction datasets via Text-to-Speech (TTS) tools. EMOVA-SFT-Speech-231K is part of EMOVA-Datasets collection and is used in… See the full description on the dataset page: https://huggingface.co/datasets/Emova-ollm/emova-sft-speech-231k.Vaani-Benchmark-V1.0
Vaani-Benchmark-V1.0
A curated ASR evaluation set drawn from the Vaani project. This benchmark contains 5,050 audio segments from 1,103 speakers across 104 Indian districts, each with three independent human transcriptions.
Evaluation Toolkit
A standalone toolkit implementing this benchmark's scoring methodology, plus
Latin-script normalization for code-switched predictions and one-command
publishing of results to a model's HF card, is available at… See the full description on the dataset page: https://huggingface.co/datasets/ARTPARK-IISc/Vaani-Benchmark-V1.0.ScreenASR-Bench
ScreenASR-Bench
Data
Item
Value
Split
test
Cases
2,002
Audio clips
2,002
Keyframes
2,469
Languages
Chinese
Structure
Field
Type
Description
caseid
string
Unique case identifier
ref
string
Reference transcription
target
string
Target text in TN form
level
string
Difficulty level: L1, L2, or L3
audio
audio
Audio clip
keyframes
list[image]
Keyframes associated with the case
frame_captions
list[string]… See the full description on the dataset page: https://huggingface.co/datasets/MingweiFu/ScreenASR-Bench.una-fraza-al-diya
Una fraza al diya
Ladino language learning sentences prepared by Karen Sarhon of Sephardic Center of Istanbul. Each sentence has translations in Turkish, English, Spanish. Includes audio and image. 307 sentences in total.
Source: https://sefarad.com.tr/judeo-espanyolladino/frazadeldia/
Citation
If you use this dataset, please cite:
Preparing an Endangered Language for the Digital Age: The Case of Judeo-Spanish
Preparing an endangered language for the digital age: The… See the full description on the dataset page: https://huggingface.co/datasets/collectivat/una-fraza-al-diya.emova-sft-speech-eval
EMOVA-SFT-Speech-Eval
🤗 EMOVA-Models | 🤗 EMOVA-Datasets | 🤗 EMOVA-Demo
📄 Paper | 🌐 Project-Page | 💻 Github | 💻 EMOVA-Speech-Tokenizer-Github
Overview
EMOVA-SFT-Speech-Eval is an evaluation dataset curated for omni-modal instruction tuning and emotional spoken dialogue. This dataset is created by converting existing text and visual instruction datasets via Text-to-Speech (TTS) tools. EMOVA-SFT-Speech-Eval is part of EMOVA-Datasets collection, and the training… See the full description on the dataset page: https://huggingface.co/datasets/Emova-ollm/emova-sft-speech-eval.Vaani-Atypical-Speech-CorpusProject Euphonia is a public initiative led by Google that aims to improve Automatic Speech Recognition (ASR) for individuals with atypical speech. To date, most of Project Euphonia’s work has focused on English, resulting in outcomes such as the Android application Project Relate, which generates personalized speech recognition models in English.
In recent years, the project has expanded its data collection efforts to additional languages, including French, Spanish, Japanese, and Hindi.
The… See the full description on the dataset page: https://huggingface.co/datasets/ARTPARK-IISc/Vaani-Atypical-Speech-Corpus.Sagarmatha-ASR-Nepali-Diamond-V3
Dataset Card for Sagarmatha ASR Nepali Diamond V3
Dataset Summary
Sagarmatha ASR Nepali Diamond V3 is a large-scale, production-grade Automatic Speech Recognition (ASR) dataset designed for the Nepali language. The corpus contains 265.7 hours of verified, 16 kHz audio paired with strictly normalized Devanagari transcriptions. It was compiled and curated primarily for the fine-tuning of state-of-the-art multilingual acoustic models, including OpenAI's Whisper… See the full description on the dataset page: https://huggingface.co/datasets/tonibirat/Sagarmatha-ASR-Nepali-Diamond-V3.
