datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
OmniVideoBench
OmniVideoBench: Towards Audio-Visual Understanding Evaluation for Omni MLLMs
✨ Overview
Recent advances in multimodal large language models (MLLMs) have brought remarkable progress in video understanding.However, most existing benchmarks fail to jointly evaluate both audio and visual reasoning — often focusing on one modality or overlooking their interaction.
🎬 OmniVideoBench fills this gap.It’s a large-scale, rigorously curated… See the full description on the dataset page: https://huggingface.co/datasets/NJU-LINK/OmniVideoBench.Light-Omni-Training
Light-Omni Training Dataset
This repository contains the training data used by Light-Omni, a multimodal
agent framework for reflexive video understanding with long-term memory.
Light-Omni uses memory-augmented multimodal streams to train adapters for
memory construction, response generation, and reaction/action control.
Links
Project page: https://clare-nie.github.io/Light-Omni/
Code: https://github.com/Clare-Nie/Light-Omni
Dataset:… See the full description on the dataset page: https://huggingface.co/datasets/ClareNie/Light-Omni-Training.anonymous-storybench
Omni-StoryBench
Omni-StoryBench is a context-aware omnimodal story generation benchmark.Each sample provides a current story page and requires generating the next page's image, narration text, and speech utterance.
Dataset Structure
The dataset contains:
data/testset.jsonl: Main benchmark file.
images/: Page images.
texts/: Page text files.
speech/: Generated speech audio files.
instruction/: Source-level instruction metadata.
Data Fields
Each JSONL sample… See the full description on the dataset page: https://huggingface.co/datasets/omnibench/anonymous-storybench.OmniRewardBench
Omni-Reward: Towards Generalist Omni-Modal Reward Modeling with Free-Form Preferences
📄 Paper |
💻 Code |
🤗 Benchmark (This Dataset) |
🤗 Training Data |
🤗 Model |
🏠 Homepage
Reward models (RMs) play a critical role in aligning AI behaviors with human preferences, yet they face two fundamental challenges: (1) Modality Imbalance, where most RMs are mainly focused on text and image modalities, offering limited support for video, audio, and other modalities; and… See the full description on the dataset page: https://huggingface.co/datasets/HongbangYuan/OmniRewardBench.OpenDialog_English
OpenDialog English
This dataset contains English dialog and conversation data.
Dataset Structure
The dataset is provided in Parquet format with 153 splits for efficient loading.
Data Files
Format: Parquet
Splits: 153 files (train-00001-of-00153.parquet through train-00153-of-00153.parquet)
Total Size: ~72.8 GB
Loading the Dataset
from datasets import load_dataset
# Load the full dataset
dataset =… See the full description on the dataset page: https://huggingface.co/datasets/zuhri025/OpenDialog_English.OmniFair
OmniFair: A Unified Fairness Benchmark Across Tasks and Modalities
OmniFair is the first fairness benchmark with instance-level cross-task and cross-modal alignment. It introduces the Bias Semantic Unit (BSU), a task- and modality-independent representation of bias, and uses it to construct:
BiasAtlas: a semantic space of 416K unique BSUs distilled from 143 fairness benchmarks and 11M multilingual news articles.
OmniFair Benchmark: 46,125 evaluation instances spanning 5 task… See the full description on the dataset page: https://huggingface.co/datasets/dyf2316/OmniFair.OmniAgentBench
OmniAgentBench Dataset
Overview
OmniAgentBench is a benchmark for evaluating multimodal agents under realistic "wild" conditions: speech input, acoustic noise, dense/scattered instructions, and multi-turn conversations. It wraps three existing agent benchmarks (MPCC, GUI Odyssey, EmbodiedBench) with speech audio, noise overlays, and wild text rewrites so that the same tasks can be evaluated under controlled input-modality variations.
Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/omniagentbenchspeech/OmniAgentBench.ONOTE
ONOTE: Omnimodal Notation Objective Topology Examination
ONOTE is a large-scale, omnimodal benchmark designed to evaluate music understanding across three major notation systems: Western Staff, Jianpu (Numbered Notation), and Guitar Tablature.
📂 Dataset Structure
The dataset is organized into two primary sub-directories based on the notation and instrument type:
1. pitch_Jianpu_dataset (Staff & Numbered Notation)
This sub-folder focuses on Western staff and… See the full description on the dataset page: https://huggingface.co/datasets/Weisiqing123/ONOTE.AURA-Chat-Edit
AURA-Chat-Edit: Conversational Music Editing Dataset
Overview
AURA-Chat-Edit is a large-scale conversational music editing dataset used to train AURA, a unified multimodal framework for conversational music editing. The dataset contains 66,539 multi-turn dialogues pairing natural-language edit instructions with structured edit-token outputs across 7 edit types.
Each dialogue simulates a user requesting a music edit (e.g., "remove the drums", "add a jazzy… See the full description on the dataset page: https://huggingface.co/datasets/OpenRB-Lab/AURA-Chat-Edit.ONOTE
ONOTE: Omnimodal Notation Objective Topology Examination
ONOTE is a large-scale, omnimodal benchmark designed to evaluate music understanding across three major notation systems: Western Staff, Jianpu (Numbered Notation), and Guitar Tablature.
📂 Dataset Structure
The dataset is organized into two primary sub-directories based on the notation and instrument type:
1. pitch_Jianpu_dataset (Staff & Numbered Notation)
This sub-folder focuses on Western staff and… See the full description on the dataset page: https://huggingface.co/datasets/chengxin666/ONOTE.Onomatopoeia_Dataset🎧 Onomatopoeia Dataset (Audio → Manga Expression)
音声解析結果をもとに、日本語のオノマトペ(擬音語・擬態語)を生成するためのデータセットです。
本データセットは、音そのものではなく、音から推定された特徴・空間・情景を入力とする構造化データであり、
漫画的な表現生成を目的としたマルチモーダルデータです。
📌 Dataset Summary
本データセットは以下のパイプラインから生成されています:
Audio
↓
Audio Features (04_features.json)
↓
Audio Events (05_audio_events.json)
↓
Space Judgement (06_space_judgement.json)
↓
Scene Interpretation (07_scene_interpretation.json)
↓
Onomatopoeia (08_onomatopoeia.json)
👉 音 → 空間 → 情景 → オノマトペ
という段階的生成構造を持ちます。
📊… See the full description on the dataset page: https://huggingface.co/datasets/yadorigi/Onomatopoeia_Dataset.acp-corpus-filtered
ACP Filtered Conversations (via Cortico)
This dataset is a filtered slice of conversation recordings and transcripts
from the American Conversation Project (ACP), retrieved via
Cortico's ACP integration. "Filtered" means every
fragment included here already passed an LLM-based salience pass (the
project's internal "wheat vs. chaff" filter) that removed filler, small talk,
and interjections — everything kept is a substantive, quote-anchored moment
someone actually said.
This is… See the full description on the dataset page: https://huggingface.co/datasets/omzugo/acp-corpus-filtered.
