datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Omni-MATH
Dataset Card for Omni-MATH
Recent advancements in AI, particularly in large language models (LLMs), have led to significant breakthroughs in mathematical reasoning capabilities. However, existing benchmarks like GSM8K or MATH are now being solved with high accuracy (e.g., OpenAI o1 achieves 94.8% on MATH dataset), indicating their inadequacy for truly challenging these models. To mitigate this limitation, we propose a comprehensive and challenging benchmark specifically designed… See the full description on the dataset page: https://huggingface.co/datasets/KbsdJames/Omni-MATH.pubmed-abstracts-36M
OmniBioAI PubMed Abstracts — 37.8M+
The most comprehensive open collection of
PubMed biomedical abstracts.
Stats
37,846,388 abstracts (full PubMed coverage)
150 biomedical domains
56 general corpus chunks
207 total files
JSONL.gz format (human readable)
FREE and open access
Coverage
Complete PubMed database as of 2026.
Format
Each line = one abstract in JSON:
{"pmid": "...", "title": "...",
"abstract": "...", "authors": [...]… See the full description on the dataset page: https://huggingface.co/datasets/omnibioai/pubmed-abstracts-36M.Daily-OmniThis is the official dataset for Daily-Omni. Check code repository for instructions.
Omni-EuropatWe flattened the multilingual pairs in Europat into document pairs. And we generated sythetic images with captions. This is intended to teach a model multilingual, multimodal abilities related to science, technology and patents.
Note that the images and flattening are not always accurate. This dataset is meant for pretraining, and not high accuracy post-training.
We are still in the process of uploading this dataset, but due to upload limitations, it will take us a bit of time. Please be… See the full description on the dataset page: https://huggingface.co/datasets/EmpathicRobotics/Omni-Europat.OmniDocBenchForked from opendatalab/OmniDocBench.
Sampler
We have added a simple Python tool for filtering and performing stratified sampling on OmniDocBench data.
Features
Filter JSON entries based on custom criteria
Perform stratified sampling based on multiple categories
Handle nested JSON fields
Installation
Local Development Install (Recommended)
git clone https://huggingface.co/Quivr/OmniDocBench.git
cd OmniDocBench
pip install -r requirements.txt #… See the full description on the dataset page: https://huggingface.co/datasets/Quivr/OmniDocBench.OmniVideo-Test
OmniVideo-Test
Official repository for OmniVideo-Test, the human-verified test set introduced in our paper: "OmniVideo-100K: A Dataset for Audio-Visual Reasoning through Structured Scripts and Evidence Chains".
This repository includes:
videos/: Raw video files.
test_505.jsonl: The test set containing 505 multiple-choice QA pairs, complete with task taxonomies, ground-truth answers, and options.
OmniVideo-Test serves as the evaluation companion to the OmniVideo-100K… See the full description on the dataset page: https://huggingface.co/datasets/MiG-NJU/OmniVideo-Test.OmniStar-RNGOmniPro
OmniPro
A comprehensive benchmark for evaluating proactive video understanding capabilities of omni multimodal large language models (MLLMs). Unlike traditional reactive QA benchmarks where models respond to explicit questions after watching a video, OmniPro evaluates whether models can proactively monitor video streams and respond at the right moment when specific conditions are met.
OmniPro is designed around three core capabilities that define a good omni-proactive model:… See the full description on the dataset page: https://huggingface.co/datasets/RuixiangZhao/OmniPro.S1-Omni-Corpus-10K
S1-Omni-Corpus-10K
An open-source scientific multimodal reasoning dataset subset for S1-Omni
🧬 Model Introduction
S1-Omni is a unified scientific multimodal reasoning model for scientific understanding, prediction, and generation. It is developed by the ScienceOne AI team of the Chinese Academy of Sciences.
S1-Omni addresses fragmented scientific AI capabilities with a shared backbone for cross-disciplinary, cross-modal, and cross-task understanding and reasoning… See the full description on the dataset page: https://huggingface.co/datasets/ScienceOne-AI/S1-Omni-Corpus-10K.OmniVChat
OmniVChat: Synthesizing, Benchmarking, and Training for Native Audio-Visual Dialogue
1 The Chinese University of Hong Kong 2 Alibaba Token Hub, Alibaba Group
3 Shanghai Jiao Tong University 4 Shanghai Innovation Institute 5 Zhejiang University
OmniVChat (Omni Video Chat) is the task of native audio-visual dialogue: an omni model
directly and simultaneously receives audio and video from a user and returns text. The user's
query is inside the audio and… See the full description on the dataset page: https://huggingface.co/datasets/Harland/OmniVChat.Omni-Edu
Omni-Edu — Core V6 SFT mixture
69,999 supervised instruction examples (~158M characters) covering K-12 subject
competence, curriculum grounding, diagnostic reasoning, pedagogical action and
general-purpose instruction. 12,146 rows (17.4%) are multimodal; every image
referenced by the JSONL ships in this repository under images/.
This is the system-prompted assembly of the v6 core mixture: every row carries
an explicit system message, and the non-system turns are byte-identical… See the full description on the dataset page: https://huggingface.co/datasets/lhpku20010120/Omni-Edu.OmniVideo-100K
OmniVideo-100K
Official repository for OmniVideo-100K, an instruction-tuning dataset introduced in our paper: "OmniVideo-100K: A Dataset for Audio-Visual Reasoning through Structured Scripts and Evidence Chains".
This repository includes:
videos.tar.part_xx: Raw video files.
train_oe_70k.jsonl: Original Open-Ended (OE) training samples.
train_mcq_30k.jsonl: Original Multiple-Choice (MCQ) training samples.
train_oe_70k_formatted.jsonl: Instruction-formatted OE samples (ready for… See the full description on the dataset page: https://huggingface.co/datasets/MiG-NJU/OmniVideo-100K.OmniEdit-Bench
OmniEdit-Bench
This repository contains a comprehensive benchmark for instruction-based video editing. It evaluates whether a model can edit an input video according to a natural-language instruction while preserving unrelated content and maintaining visual, temporal, and audio consistency.
The benchmark covers five tracks: spatial track, temporal track, reasoning track, audio track, and reference-based track. Each task provides a source video and an editing instruction.… See the full description on the dataset page: https://huggingface.co/datasets/OmniEdit-Bench/OmniEdit-Bench.DREAM-1K
DREAM-1K
DREAM-1K (Description
with Rich Events, Actions, and Motions) is a challenging video description benchmark. It contains a collection of 1,000 short (around 10 seconds) video clips with diverse complexities from five different origins: live-action movies, animated movies, stock videos, long YouTube videos, and TikTok-style short videos. We provide a fine-grained manual annotation for each video.
Bellow is the dataset statistics:
omnidocbench-render-compare
OmniDocBench Render-and-Compare
This dataset contains the rendered HTML reconstructions and comparison images produced
by a render-and-compare pipeline — a reference-free visual similarity evaluation
framework for OCR systems.
Overview
The pipeline processes each page of OmniDocBench through
a Qwen3.5-122B-A10B OCR model, renders the structured output back to a PNG via HTML
(reconstructed.png), and compares it against the original page scan (masked_original.png)
using… See the full description on the dataset page: https://huggingface.co/datasets/gt-free-ocr-metrics/omnidocbench-render-compare.omniact-gui-trajectories
OmniACT
OmniACT is a GUI trajectory dataset with single-step action traces grounded in screenshots.
Dataset Structure
.
├── README.md
├── .gitattributes
├── data/
│ └── train.jsonl
├── observations/
│ └── OmniACT_pilot_*/000/screenshot.jpg
└── env_meta/
└── OmniACT_pilot_*/000/metadata.json
Each row in data/train.jsonl is one trajectory. The main image path is stored in the top-level image field, and the same relative path is also used inside… See the full description on the dataset page: https://huggingface.co/datasets/Dhscl/omniact-gui-trajectories.OmniCharacterThis is the official data collection for paper "OmniCharacter: Towards Immersive Role-Playing Agents with Seamless Speech-Language Personality Interaction".
Please see paper & code for more information:
paper: https://www.arxiv.org/abs/2505.20277
code: https://github.com/AlibabaResearch/DAMO-ConvAI/tree/main/OmniCharacter
OmniRef-trainingOmni-MATH-2
Dataset Card for Omni-MATH-2
Benchmarks are important tools for tracking progress in the development of large language models (LLMs). However, inaccuracies in datasets and evaluation methods often undermine their effectiveness. Here, we present Omni-MATH-2: a manually revised version of the original Omni-MATH dataset that preserves its size (n = 4,428) while significantly improving LaTeX compilability, solvability, and verifiability. A total of 647 problems were edited (14.6%) and… See the full description on the dataset page: https://huggingface.co/datasets/martheballon/Omni-MATH-2.OmniCap-IF-54K
OmniCap-IF-54K
OmniCap-IF-54K is a large-scale instruction-tuning dataset for improving instruction-following abilities in omni-modal video captioning. It contains 54K curated video-instruction-response triplets covering format constraints, temporal grounding, visual and audio content constraints, and audio-visual synergy.
The dataset is constructed through a three-stage pipeline: video curation, constraint-aware instruction synthesis… See the full description on the dataset page: https://huggingface.co/datasets/NJU-LINK/OmniCap-IF-54K.OmniAlign-V-DPO
Introduction
Paper: Paper,
Github: Github,
Page: Page,
SFT Dataset: OmniAlign-V,
MM-AlignBench: VLMEvalkit, Huggingface
Checkpoints: LLaVANext-OA-7B, LLaVANext-OA-32B, LLaVANext-OA-32B-DPO
This is the official repo of OmniAlign-V-DPO datasets in OmniAlign-V: Towards Enhanced Alignment of MLLMs with Human Preference
OmniAlign-V-DPO datasets contains 150k high-qulity positive-negative pairs for Direct Preference Optimization(DPO) based on OmniAlign-V datasets. It utilizes the answer… See the full description on the dataset page: https://huggingface.co/datasets/PhoenixZ/OmniAlign-V-DPO.OmniEAR
OmniEAR Expert Trajectory Dataset
Dataset Summary
The OmniEAR Expert Trajectory SFT Dataset is a comprehensive collection of high-quality expert demonstration trajectories specifically designed for supervised fine-tuning (SFT) of embodied reasoning models. This dataset contains 1,982 instruction-following examples across single-agent and multi-agent scenarios, focusing on physical interactions, tool usage, and collaborative reasoning in embodied environments.… See the full description on the dataset page: https://huggingface.co/datasets/wangzx1210/OmniEAR.OmniEarth-Bench_MCQ_VLMOmniParsingBench
🤗 Model | 📑 Technical Report | 💻 GitHub
OmniParsingBench is a comprehensive, large-scale, and high-quality evaluation corpus designed to rigorously evaluate the unified parsing capabilities of Multimodal Large Language Models (MLLMs) across diverse modalities.
Unlike traditional single-task benchmarks, OmniParsingBench assesses the full spectrum of parsing performance—from fundamental signal detection to complex semantic reasoning—across six primary domains: Document… See the full description on the dataset page: https://huggingface.co/datasets/Logics-MLLM/OmniParsingBench.tinygiant-omni-test-featuresLongText-Bench
📊Dataset Card for LongText-Bench
LongText-Bench, proposed in X-Omni, focuses on evaluating the performance on rendering longer texts in both English and Chinese.
Leaderboard
Method
Open-source
Avg.
English
Chinese
Seedream 3.0
0.887
0.896
0.878
X-Omni
✓
0.857
0.900
0.814
GPT-4o
0.788
0.956
0.619
BAGEL
✓
0.342
0.373
0.310
OmniGen2
✓
0.310
0.561
0.059
FLUX.1-dev
✓
0.306
0.607
0.005
Kolors 2.0
0.294
0.258
0.329
HiDream-I1-Full
✓
0.284
0.543
0.024… See the full description on the dataset page: https://huggingface.co/datasets/X-Omni/LongText-Bench.OmniCoding
OmniCoding
A multimodal terminal-tool-use SFT/RL dataset. Each record is a
question + verifiable answer + media (video/audio/image) — the target
agent is expected to operate on the media via a Linux terminal (ffmpeg,
ffprobe, whisper, python, etc.) rather than a GUI.
Aggregated and filtered from four upstream sources, with a single unified
schema, global dedup, and category-balanced sampling.
Records
Source
n
Omnimodal-Agent-SFT-2K (RUC-NLPIR) — agentic… See the full description on the dataset page: https://huggingface.co/datasets/shuaishuaicdp/OmniCoding.OmniPro
OmniPro
A comprehensive benchmark for evaluating proactive video understanding capabilities of omni multimodal large language models (MLLMs). Unlike traditional reactive QA benchmarks where models respond to explicit questions after watching a video, OmniPro evaluates whether models can proactively monitor video streams and respond at the right moment when specific conditions are met.
OmniPro is designed around three core capabilities that define a good omni-proactive model:… See the full description on the dataset page: https://huggingface.co/datasets/omniproact-bench/OmniPro.lm-eval-results-paulml-OmniBeagleSquaredMBX-v3-7B-v2-private
Dataset Card for Evaluation run of paulml/OmniBeagleSquaredMBX-v3-7B-v2
Dataset automatically created during the evaluation run of model paulml/OmniBeagleSquaredMBX-v3-7B-v2
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-paulml-OmniBeagleSquaredMBX-v3-7B-v2-private.Omni-DeepSearch
✨ Focus on Multimodal Audio Deep‑Search: Automated generation, filtering, and evaluation of multi‑hop reasoning benchmarks ✨
| 🧩 Fully Automated Pipeline | 🎵 Rich Audio Domains | 🧠 Multi‑Hop Reasoning QA | 🤖 Agentic Evaluation |
Omni-DeepSearch Benchmark
🎧 Multimodal Audio Focus – Designed for audio‑centric deep‑search tasks covering speech, music, bio‑acoustics, and environmental sounds.
🔄 Fully Automated Pipeline – End‑to‑end generation, multi‑stage… See the full description on the dataset page: https://huggingface.co/datasets/Kirito-Lab/Omni-DeepSearch.
