datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ASMR-Archive-Processed
ASMR-Archive-Processed (WIP)
Update (2026-04-03): This dataset has reached the Hugging Face Public Storage Limit. After contacting support, we were informed that the only option is to pay for a storage expansion. Consequently, updates to this dataset are now suspended.
Work in Progress — expect breaking changes while the pipeline and data layout stabilize.
This dataset contains ASMR audio data sourced from DeliberatorArchiver/asmr-archive-data-01 and… See the full description on the dataset page: https://huggingface.co/datasets/OmniAICreator/ASMR-Archive-Processed.OmniReasoner-SFT
OmniReasoner-SFT
OmniReasoner-SFT is a mixed-source, research-only supervised fine-tuning dataset
for audio-visual and long-video reasoning. It contains two-stage cold-start SFT
trajectories with interval selection, zoom-in evidence, and final answers.
Contents
data/train.jsonl: HF-ready training JSONL with repo-relative media paths.
media/: raw and derived media referenced by train.jsonl.
manifests/media_manifest.jsonl: media inventory with repo paths, source
family… See the full description on the dataset page: https://huggingface.co/datasets/Rocky131/OmniReasoner-SFT.omnilingual-asr-corpus
Meta Omnilingual ASR Corpus
The Omnilingual ASR Corpus is a collection of spontaneous speech recordings and their transcriptions for 348 under-served languages. The corpus was collected as part of Meta FAIR’s Omnilingual ASR project (blog, model, paper) for the purposes of training automatic speech recognition (ASR) and spoken language identification models.
Data schema
{
`language`: "lij_Latn",
`iso_639_3`: "lij",
`iso_15924`: "Latn",
`glottocode`:… See the full description on the dataset page: https://huggingface.co/datasets/facebook/omnilingual-asr-corpus.Omnimodal-Agent-SFT-2K
OmniGAIA: Omni-Modal General AI Assistant Benchmark
📄 Paper
•
💻 Code & Demo
•
🤗 Dataset & Model
•
📈 Leaderboard
This dataset contains omni-modal agent supervised fine-tuning (SFT) trajectories in the LlamaFactory SFT data format. You can directly follow LlamaFactory's instructions to fine-tune your omni-modal LLMs.OmniGAIA is a benchmark for Omni-Modal General AI Assistants that jointly reason over vision, audio, and language with external tools. It is… See the full description on the dataset page: https://huggingface.co/datasets/RUC-NLPIR/Omnimodal-Agent-SFT-2K.VoiceAssistant-400KOmni-Fake-SET
Omni-Fake-SET
Omni-Fake-SET is the in-distribution split of Omni-Fake, a unified multimodal deepfake dataset for social-media forensics. It covers image, audio, video, and audio–video talking-head (AV-TH) modalities. Each modality uses the same three-way label space: real, fully synthetic, and tampered. Pair with the held-out benchmark Omni-Fake-OOD for out-of-distribution evaluation.
Paper: arXiv:2605.01638
Project page: Omni-Fake
License: CC-BY-4.0
Video (hybrid… See the full description on the dataset page: https://huggingface.co/datasets/JamalLee/Omni-Fake-SET.OmniVideoBench
OmniVideoBench: Towards Audio-Visual Understanding Evaluation for Omni MLLMs
✨ Overview
Recent advances in multimodal large language models (MLLMs) have brought remarkable progress in video understanding.However, most existing benchmarks fail to jointly evaluate both audio and visual reasoning — often focusing on one modality or overlooking their interaction.
🎬 OmniVideoBench fills this gap.It’s a large-scale, rigorously curated… See the full description on the dataset page: https://huggingface.co/datasets/NJU-LINK/OmniVideoBench.OmniAction-LIBERO-evalOmnibus
仓库信息
电报地址:https://t.me/vomebook ,有问题请在:https://huggingface.co/datasets/VoiceOfML/Omnibus/discussions 提出。
此仓库存储书纳百川专题:https://huggingface.co/datasets/VoiceOfML/Omnibus/tree/main 。
请使用:https://voiceofml-search.hf.space/Omnibus 进行文件检索(备用搜索站:https://vomebook.github.io/search/#/Omnibus )。
可使用:https://voiceofml-search.hf.space/Omnibus?wide=1 进行仓库内容查看(备用站:https://voiceofml-search.hf.space/Omnibus?wide=1 )。
你可以仅下载指针(只有文件名的信息)
If you want to clone without large files - just their… See the full description on the dataset page: https://huggingface.co/datasets/VoiceOfML/Omnibus.OmniGUI
OmniGUI: Benchmarking GUI Agents in Omni-Modal Smartphone Environments
[🍎 Project Page] [📖 arXiv Paper] [📊 Dataset] [💻 GitHub] [🏆 Leaderboard]
OmniGUI is a step-level GUI agent benchmark designed for omni-modal smartphone interaction. At each action step, the agent receives interleaved multimodal observations, including static screenshots, synchronous audio cues, short video clips, and action history, and must predict the next GUI action such as TAP or TYPE.
The benchmark… See the full description on the dataset page: https://huggingface.co/datasets/OmniGUI/OmniGUI.Omni-Fake-OOD
Omni-Fake-OOD
Omni-Fake-OOD is the out-of-distribution benchmark split of Omni-Fake. Samples come from held-out generators and platforms not included in training, for measuring cross-domain generalization. It covers image, audio, video, and audio–video talking-head (AV-TH) with the same three-class labels as Omni-Fake-SET: real, fully synthetic, and tampered. Use together with Omni-Fake-SET (in-distribution training data).
Paper: arXiv:2605.01638
Project page: Omni-Fake… See the full description on the dataset page: https://huggingface.co/datasets/JamalLee/Omni-Fake-OOD.Omni-Sets
Omni-Sets
A large-scale, multi-modal instruction-tuning dataset spanning six modalities (audio, speech, image, video, visual documents, and cross-modal omni) with both single-turn dense captions and multi-turn instruction-following conversations. Designed for training omni-modal language models that can perceive and reason across all modalities.
590,858 total samples | 5,635 hours of audio/video | 6 configs | 17 source datasets
Overview
Config
Modality… See the full description on the dataset page: https://huggingface.co/datasets/MBZUAI/Omni-Sets.OmnilingualASR-retrieval
Omnilingual ASR speech-text retrieval (MTEB)
Read speech paired with its human transcription, for languages that no existing
MTEB audio task covers.
Source: facebook/omnilingual-asr-corpus at revision 8648ba8, cc-by-4.0, official
test split. Recordings are re-encoded from FLAC to Opus at 16 kHz. Repeated
transcripts are dropped, since one would otherwise be relevant to several
recordings while only one is marked correct.
Built by… See the full description on the dataset page: https://huggingface.co/datasets/vnahata/OmnilingualASR-retrieval.OmniGAIA
OmniGAIA: Omni-Modal General AI Assistant Benchmark
📄 Paper
•
💻 Code & Demo
•
🤗 Dataset & Model
•
📈 Leaderboard
OmniGAIA is a benchmark for Omni-Modal General AI Assistants that jointly reason over vision, audio, and language with external tools. It is designed to evaluate long-horizon, multi-hop, open-form problem solving in realistic settings rather than short perception-only QA.
Benchmark Construction
The OmniGAIA construction… See the full description on the dataset page: https://huggingface.co/datasets/RUC-NLPIR/OmniGAIA.OmniBench
OmniBench
🌐 Homepage | 🏆 Leaderboard | 📖 Arxiv Paper | 🤗 Paper | 🤗 OmniBench Dataset | | 🤗 OmniInstruct_V1 Dataset | 🦜 Tweets
The project introduces OmniBench, a novel benchmark designed to rigorously evaluate models' ability to recognize, interpret, and reason across visual, acoustic, and textual inputs simultaneously. We define models capable of such tri-modal processing as omni-language models (OLMs).
Mini Leaderboard
This table shows the omni-language models in… See the full description on the dataset page: https://huggingface.co/datasets/m-a-p/OmniBench.OmniCVROmni-DuplexEval
Omni-DuplexEval
📖 arXiv | GitHub
Omni-DuplexEval is a benchmark for evaluating real-time duplex multimodal interaction. Unlike conventional offline video understanding benchmarks, Omni-DuplexEval focuses on streaming settings where models must continuously process evolving multimodal inputs and decide what to respond and when to respond.
The benchmark contains two scenarios:
Real-Time Description (RTD): evaluates continuous streaming description ability.
Proactive Reminder (PR):… See the full description on the dataset page: https://huggingface.co/datasets/Hothan/Omni-DuplexEval.AdvBench-omni
AdvBench-Omni
Paper | Code
AdvBench-Omni is a dataset constructed to evaluate and reveal safety vulnerabilities in Omni-modal Large Language Models (OLLMs). It is based on a modality-semantics decoupling principle to study how OLLMs handle cross-modal safety risks and conflicts.
Introduction
Omni-modal Large Language Models (OLLMs) expand multimodal capabilities but introduce new cross-modal safety risks. AdvBench-Omni reveals a significant vulnerability in these… See the full description on the dataset page: https://huggingface.co/datasets/shipnebula/AdvBench-omni.omnilingual-asr-corpus
Meta Omnilingual ASR Corpus
The Omnilingual ASR Corpus is a collection of spontaneous speech recordings and their transcriptions for 348 under-served languages. The corpus was collected as part of Meta FAIR’s Omnilingual ASR project (blog, model, paper) for the purposes of training automatic speech recognition (ASR) and spoken language identification models.
Data schema
{
`language`: "lij_Latn",
`iso_639_3`: "lij",
`iso_15924`: "Latn"… See the full description on the dataset page: https://huggingface.co/datasets/KathleenKunLiu/omnilingual-asr-corpus.Light-Omni-Training
Light-Omni Training Dataset
This repository contains the training data used by Light-Omni, a multimodal
agent framework for reflexive video understanding with long-term memory.
Light-Omni uses memory-augmented multimodal streams to train adapters for
memory construction, response generation, and reaction/action control.
Links
Project page: https://clare-nie.github.io/Light-Omni/
Code: https://github.com/Clare-Nie/Light-Omni
Dataset:… See the full description on the dataset page: https://huggingface.co/datasets/ClareNie/Light-Omni-Training.Omni_Bench_fixOmniReasoner-SFT
OmniReasoner-SFT
OmniReasoner-SFT is a mixed-source, research-only supervised fine-tuning dataset
for audio-visual and long-video reasoning. It contains two-stage cold-start SFT
trajectories with interval selection, zoom-in evidence, and final answers.
Contents
data/train.jsonl: HF-ready training JSONL with repo-relative media paths.
media/: raw and derived media referenced by train.jsonl.
manifests/media_manifest.jsonl: media inventory with repo paths, source… See the full description on the dataset page: https://huggingface.co/datasets/wyqsss/OmniReasoner-SFT.OmniScientist
OmniScientist Data
Raw scientific data behind the runs in OmniScientist: An Omni-Modal Omni-Discipline AI Scientist, one folder per case. The agent reads these files directly, so what is here is what it saw.
Paper: https://arxiv.org/abs/2608.13558
Project page: https://omni-scientist.github.io
Code: https://github.com/Omni-Scientist/OmniScientist
13 of the 36 cases in the paper are here. The rest are left out because their upstream terms grant no redistribution right, or… See the full description on the dataset page: https://huggingface.co/datasets/BradNLP/OmniScientist.omnievalkit-dataset
OmniEvalKit Evaluation Datasets
Evaluation datasets for OmniEvalKit,
a comprehensive evaluation framework for omni-modal (audio + video + image + text) models.
Overview
Total subsets: 65
Total samples: 315,264
Total size: 620.3 GB (Parquet with embedded audio/image/video)
Subsets with embedded video: 15
Subsets requiring external video download: 2
Usage
from datasets import load_dataset
ds = load_dataset("OmniEvalKit/omnievalkit-dataset", "aishell1_test")… See the full description on the dataset page: https://huggingface.co/datasets/OmniEvalKit/omnievalkit-dataset.omnivoice-best-of-n-training
🎙 Best-of-N Voice Cloning Training Data
Curated training dataset for fine-tuning OmniVoice for IWSLT 2026.
🎧 Listen to the Audio
This dataset has playable audio columns — click on any row in the dataset viewer
to listen to both the reference audio (original speaker) and the best synthesized audio
(selected by quality score).
Dataset Description
For each sentence in the dev split of ymoslem/acl-6060 (468 samples × 3 languages),
we synthesized audio with the… See the full description on the dataset page: https://huggingface.co/datasets/amanuelbyte/omnivoice-best-of-n-training.multi_en_qwen3_omni_sglang_regeneratedanonymous-storybench
Omni-StoryBench
Omni-StoryBench is a context-aware omnimodal story generation benchmark.Each sample provides a current story page and requires generating the next page's image, narration text, and speech utterance.
Dataset Structure
The dataset contains:
data/testset.jsonl: Main benchmark file.
images/: Page images.
texts/: Page text files.
speech/: Generated speech audio files.
instruction/: Source-level instruction metadata.
Data Fields
Each JSONL sample… See the full description on the dataset page: https://huggingface.co/datasets/omnibench/anonymous-storybench.bashkort_commands_omnivoice
Bashkort Commands OmniVoice
Partial eleven-label command snapshot generated with k2-fsa/OmniVoice
using the same cross-lingual voice-cloning recipe as
AigizK/homai_wake_word_omnivoice. Generation was stopped at the user's
request after 41,525 complete reference groups had been committed.
For every included reference row from the train split of:
bond005/sova_rudevices
the dataset contains one recording of every command:
Айвика — Russian
Айвикә — Bashkir
Айһылыу — Bashkir… See the full description on the dataset page: https://huggingface.co/datasets/AigizK/bashkort_commands_omnivoice.Daily-Omni
Daily-Omni
This repository provides the question-answering metadata of the Daily-Omni benchmark in a format compatible with lmms-eval.
The data is provided as a single parquet file containing only the QA annotations. Since raw videos are not included, please download them from the original release and match them with the QA annotations using video_id.
Task configurations and evaluation scripts are available in the SEATS repository: https://github.com/xxayt/SEATS.… See the full description on the dataset page: https://huggingface.co/datasets/xxayt/Daily-Omni.omnivoice-fr
Sample rate
24 kHz
Voice-designed
10,000
Voice-cloned
10,000
Total
20,000
