omni-modal
Omnimodal-Agent-SFT-2K
OmniGAIA: Omni-Modal General AI Assistant Benchmark
📄 Paper
•
💻 Code & Demo
•
🤗 Dataset & Model
•
📈 Leaderboard
This dataset contains omni-modal agent supervised fine-tuning (SFT) trajectories in the LlamaFactory SFT data format. You can directly follow LlamaFactory's instructions to fine-tune your omni-modal LLMs.OmniGAIA is a benchmark for Omni-Modal General AI Assistants that jointly reason over vision, audio, and language with external tools. It is… See the full description on the dataset page: https://huggingface.co/datasets/RUC-NLPIR/Omnimodal-Agent-SFT-2K.Omni-Bench
Omni-Bench
Overview
Omni-Bench is an evaluation benchmark for unified multimodal reasoning. It contains 800 samples spanning 4 Uni-Tasks:
Natural-Scene Perception: V*
Structured-Image: ArxivQA, ChartQA
Diagrammatic Math: Geometry3k, MathVista
Vision-Operational Scenes: ViC-Bench
Data Fields
Each example contains the following fields:
image (string): the image encoded as a Base64 string.The underlying bytes are typically common image formats (e.g.… See the full description on the dataset page: https://huggingface.co/datasets/ModalityDance/Omni-Bench.desi-sv1-omnimodal
DESI SV1 Omnimodal Dataset
Dataset Summary
This dataset contains 21,763 objects from the DESI Survey Validation 1 (SV1),
combining DESI optical spectra, Legacy Survey imaging, Gaia photometry, and derived parameters.
Split
Samples
train
17,410
validation
2,176
test
2,177
total
21,763
Modalities
Modality
Column
Shape
Notes
DESI Spectrum (raw flux)
spectrum_flux_raw
(7958,)
float32, normalize in training pipeline… See the full description on the dataset page: https://huggingface.co/datasets/OneAstronomy/desi-sv1-omnimodal.AR-Omni-Instruct-v0.1
AR-Omni-Instruct
Overview
AR-Omni-Instruct is a multimodal instruction-tuning dataset for training unified autoregressive any-to-any models.
All modalities are represented as discrete tokens in a single interleaved token stream, enabling standard next-token prediction training over multimodal sequences.
Dataset Summary
Type: multimodal instruction-tuning data
Format: discrete tokenized multimodal conversations / sequences
Use case: instruction tuning… See the full description on the dataset page: https://huggingface.co/datasets/ModalityDance/AR-Omni-Instruct-v0.1.AIDAS-Omni-Modal-Diffusion-assets
