datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
OmniDocBenchForked from opendatalab/OmniDocBench.
Sampler
We have added a simple Python tool for filtering and performing stratified sampling on OmniDocBench data.
Features
Filter JSON entries based on custom criteria
Perform stratified sampling based on multiple categories
Handle nested JSON fields
Installation
Local Development Install (Recommended)
git clone https://huggingface.co/Quivr/OmniDocBench.git
cd OmniDocBench
pip install -r requirements.txt #… See the full description on the dataset page: https://huggingface.co/datasets/Quivr/OmniDocBench.OmniVideo-Test
OmniVideo-Test
Official repository for OmniVideo-Test, the human-verified test set introduced in our paper: "OmniVideo-100K: A Dataset for Audio-Visual Reasoning through Structured Scripts and Evidence Chains".
This repository includes:
videos/: Raw video files.
test_505.jsonl: The test set containing 505 multiple-choice QA pairs, complete with task taxonomies, ground-truth answers, and options.
OmniVideo-Test serves as the evaluation companion to the OmniVideo-100K… See the full description on the dataset page: https://huggingface.co/datasets/MiG-NJU/OmniVideo-Test.Omni-Edu
Omni-Edu — Core V6 SFT mixture
69,999 supervised instruction examples (~158M characters) covering K-12 subject
competence, curriculum grounding, diagnostic reasoning, pedagogical action and
general-purpose instruction. 12,146 rows (17.4%) are multimodal; every image
referenced by the JSONL ships in this repository under images/.
This is the system-prompted assembly of the v6 core mixture: every row carries
an explicit system message, and the non-system turns are byte-identical… See the full description on the dataset page: https://huggingface.co/datasets/lhpku20010120/Omni-Edu.S1-Omni-Corpus-10K
S1-Omni-Corpus-10K
An open-source scientific multimodal reasoning dataset subset for S1-Omni
🧬 Model Introduction
S1-Omni is a unified scientific multimodal reasoning model for scientific understanding, prediction, and generation. It is developed by the ScienceOne AI team of the Chinese Academy of Sciences.
S1-Omni addresses fragmented scientific AI capabilities with a shared backbone for cross-disciplinary, cross-modal, and cross-task understanding and reasoning… See the full description on the dataset page: https://huggingface.co/datasets/ScienceOne-AI/S1-Omni-Corpus-10K.OmniVideo-100K
OmniVideo-100K
Official repository for OmniVideo-100K, an instruction-tuning dataset introduced in our paper: "OmniVideo-100K: A Dataset for Audio-Visual Reasoning through Structured Scripts and Evidence Chains".
This repository includes:
videos.tar.part_xx: Raw video files.
train_oe_70k.jsonl: Original Open-Ended (OE) training samples.
train_mcq_30k.jsonl: Original Multiple-Choice (MCQ) training samples.
train_oe_70k_formatted.jsonl: Instruction-formatted OE samples (ready for… See the full description on the dataset page: https://huggingface.co/datasets/MiG-NJU/OmniVideo-100K.omnidocbench-render-compare
OmniDocBench Render-and-Compare
This dataset contains the rendered HTML reconstructions and comparison images produced
by a render-and-compare pipeline — a reference-free visual similarity evaluation
framework for OCR systems.
Overview
The pipeline processes each page of OmniDocBench through
a Qwen3.5-122B-A10B OCR model, renders the structured output back to a PNG via HTML
(reconstructed.png), and compares it against the original page scan (masked_original.png)
using… See the full description on the dataset page: https://huggingface.co/datasets/gt-free-ocr-metrics/omnidocbench-render-compare.omniact-gui-trajectories
OmniACT
OmniACT is a GUI trajectory dataset with single-step action traces grounded in screenshots.
Dataset Structure
.
├── README.md
├── .gitattributes
├── data/
│ └── train.jsonl
├── observations/
│ └── OmniACT_pilot_*/000/screenshot.jpg
└── env_meta/
└── OmniACT_pilot_*/000/metadata.json
Each row in data/train.jsonl is one trajectory. The main image path is stored in the top-level image field, and the same relative path is also used inside… See the full description on the dataset page: https://huggingface.co/datasets/Dhscl/omniact-gui-trajectories.OmniRef-trainingOmniParsingBench
🤗 Model | 📑 Technical Report | 💻 GitHub
OmniParsingBench is a comprehensive, large-scale, and high-quality evaluation corpus designed to rigorously evaluate the unified parsing capabilities of Multimodal Large Language Models (MLLMs) across diverse modalities.
Unlike traditional single-task benchmarks, OmniParsingBench assesses the full spectrum of parsing performance—from fundamental signal detection to complex semantic reasoning—across six primary domains: Document… See the full description on the dataset page: https://huggingface.co/datasets/Logics-MLLM/OmniParsingBench.Omni-Edu
Omni-Edu — instruction-tuning mixture
69,999 supervised instruction examples (~158M characters) covering K-12 subject
competence, curriculum grounding, diagnostic reasoning, pedagogical action and
general-purpose instruction. 12,146 rows (17.4%) are multimodal; every image
referenced by the JSONL ships in this repository under images/.
Composition
Capability family
Examples
Subject competence
31,855
Pedagogical action and scaffolding
14,226… See the full description on the dataset page: https://huggingface.co/datasets/OmniEdu/Omni-Edu.OmniAVSOmniVideo-100K
OmniVideo-100K
Official repository for OmniVideo-100K, an instruction-tuning dataset introduced in our paper: "OmniVideo-100K: A Dataset for Audio-Visual Reasoning through Structured Scripts and Evidence Chains".
This repository includes:
videos.tar.part_xx: Raw video files.
train_oe_70k.jsonl: Original Open-Ended (OE) training samples.
train_mcq_30k.jsonl: Original Multiple-Choice (MCQ) training samples.
train_oe_70k_formatted.jsonl: Instruction-formatted OE samples… See the full description on the dataset page: https://huggingface.co/datasets/brolydfgh/OmniVideo-100K.Omni-Edu
Omni-Edu — instruction-tuning mixture
69,999 supervised instruction examples (~158M characters) covering K-12 subject
competence, curriculum grounding, diagnostic reasoning, pedagogical action and
general-purpose instruction. 12,146 rows (17.4%) are multimodal; every image
referenced by the JSONL ships in this repository under images/.
Composition
Capability family
Examples
Subject competence
31,855
Pedagogical action and scaffolding
14,226… See the full description on the dataset page: https://huggingface.co/datasets/OpenDCAI/Omni-Edu.omnidocbench-render-compare-sample
OmniDocBench Render-and-Compare — Sample
This is a 60-page stratified sample of
gt-free-ocr-metrics/omnidocbench-render-compare
(the full dataset is ~10 GB).
It is provided to help reviewers explore the data without downloading the full dataset,
as recommended by the NeurIPS 2025 Datasets & Benchmarks Track guidelines.
Sampling Methodology
Pages were selected by stratified random sampling from the full dataset:
Each page in ocr_all (1 355 pages) was assigned to one… See the full description on the dataset page: https://huggingface.co/datasets/gt-free-ocr-metrics/omnidocbench-render-compare-sample.omniedit-dataset
OmniEdit Dataset
数据集包含用于图像编辑任务的训练和测试数据。
📁 文件结构
omniedit_train.jsonl: 训练集数据 (5.2MB)
omniedit_test.jsonl: 测试集数据 (5.4KB)
omniedit/: 相关资源文件夹
🚀 使用方法
下载整个数据集
from huggingface_hub import snapshot_download
# 下载到本地
local_path = snapshot_download(
repo_id="your-username/omniedit-dataset", # 将会自动替换为实际用户名
repo_type="dataset"
)
print(f"数据集下载到: {local_path}")
下载单个文件
from huggingface_hub import hf_hub_download
# 下载训练集
train_file =… See the full description on the dataset page: https://huggingface.co/datasets/xuelu/omniedit-dataset.omnieditOmniText-Bench
OmniText-Bench
OmniText-Bench is a benchmark for controllable text-image manipulation (TIM). It evaluates five
applications on the same 150 input images: text removal, rescaling, repositioning,
editing (content and style-reference variants) and insertion (content and style-reference variants).
Every sample ships with the input image, task masks, ground-truth results and, where applicable,
a reference image for style transfer.
It was introduced in the ICLR 2026 paper
OmniText: A… See the full description on the dataset page: https://huggingface.co/datasets/agusgun/OmniText-Bench.WebVoyager-Trajectories-Qwen3.5-Omni
WebVoyager + GAIA Agent Trajectories (Qwen3.5-Omni)
Full browser-agent trajectories for 733 tasks (643 WebVoyager + 90 GAIA-web), produced by a browser-use agent driven by qwen3.5-omni-plus-2026-03-15 (multimodal, via Alibaba DashScope), run in a real headed browser. Every step records the exact LLM context (including the screenshot the model saw) and the action taken, plus a reference-grounded success verdict.
Results
Judged by the WebVoyager reference-grounded… See the full description on the dataset page: https://huggingface.co/datasets/shiqihe/WebVoyager-Trajectories-Qwen3.5-Omni.VisuLogic-Train
VisuLogic-Train (OmniTool)
Short description of the dataset…
Usage
from datasets import load_dataset
ds = load_dataset("OmniTool/VisuLogic-Train", "solution", split="train")"internvl"
print(ds[0])
