datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
short_video_ocr_dataset
Short Video OCR / ASR Dataset
An actively curated research dataset for building OCR, ASR, subtitle-alignment,
and video-transcript pipelines for short social videos. It combines source
videos and extracted frames with human review artifacts and model-generated
text candidates. The primary languages are Ukrainian and Russian; English or
mixed-language content may also occur.
Status: work in progress. Model outputs and pseudo-label candidates are
not ground truth. Only… See the full description on the dataset page: https://huggingface.co/datasets/ElectronicHug/short_video_ocr_dataset.propagator-multimodal-pretraining-data
Propagator Multimodal Pretraining Data
This public dataset contains tokenized multimodal pretraining data prepared for the Propagator model family. It combines language, image-grounded, and speech/audio-token examples into a single training format.
This is not a raw text or image browsing dataset. The examples have already been converted into compact binary token frames for model training, with a manifest that records the source groups and file layout.
Source Code… See the full description on the dataset page: https://huggingface.co/datasets/ken-sungmin/propagator-multimodal-pretraining-data.omnievalkit-dataset
OmniEvalKit Evaluation Datasets
Evaluation datasets for OmniEvalKit,
a comprehensive evaluation framework for omni-modal (audio + video + image + text) models.
Overview
Total subsets: 65
Total samples: 315,264
Total size: 620.3 GB (Parquet with embedded audio/image/video)
Subsets with embedded video: 15
Subsets requiring external video download: 2
Usage
from datasets import load_dataset
ds = load_dataset("OmniEvalKit/omnievalkit-dataset", "aishell1_test")… See the full description on the dataset page: https://huggingface.co/datasets/OmniEvalKit/omnievalkit-dataset.omnievalkit-data-test
OmniEvalKit Evaluation Datasets
Evaluation datasets for OmniEvalKit,
a comprehensive evaluation framework for omni-modal (audio + video + image + text) models.
Overview
Total subsets: 89
Total samples: 353,610
Total size: 352.3 GB (Parquet with embedded audio/image, no video)
Subsets requiring video download: 42
Note: Video files are NOT embedded in the Parquet files due to size constraints.
Usage
from datasets import load_dataset
ds =… See the full description on the dataset page: https://huggingface.co/datasets/xiaofff/omnievalkit-data-test.video-dataset
📚 Russian Storytelling Video Dataset (700 participants)
This dataset contains full-body videos of 700 native Russian speakers engaged in unscripted storytelling. Participants freely tell personal stories, express a wide range of emotions, and naturally use facial expressions and hand gestures. Each video captures authentic human behavior in high resolution with high-quality audio.
📊 Sample
📺 Preview Video (10 Participants)
To get a quick impression of… See the full description on the dataset page: https://huggingface.co/datasets/MaratDV/video-dataset.
