datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
multi_task_multi_modal_knowledge_retrieval_benchmark_M2KR
PreFLMR M2KR Dataset Card
Dataset details
Dataset type:
M2KR is a benchmark dataset for multimodal knowledge retrieval. It contains a collection of tasks and datasets for training and evaluating multimodal knowledge retrieval models.
We pre-process the datasets into a uniform format and write several task-specific prompting instructions for each dataset. The details of the instruction can be found in the paper. The M2KR benchmark contains three types of tasks:… See the full description on the dataset page: https://huggingface.co/datasets/BByrneLab/multi_task_multi_modal_knowledge_retrieval_benchmark_M2KR.Multi-modal_dataset_named_SynthSoM
SynthSoM: A synthetic intelligent multi-modal sensing-communication dataset for Synesthesia of Machines (SoM)
📌 Overview
SynthSoM dataset covers eight rich and diverse application scenarios, including vehicle-road coordination, low-altitude economy, smart campus, as well as typical urban, suburban, rural, and campus environments. The urban scenario further includes intersections, ultra-wide lanes, elevated interchanges, and CBD areas; the suburban scenario… See the full description on the dataset page: https://huggingface.co/datasets/pku-pcni-lab/Multi-modal_dataset_named_SynthSoM.multi_task_multi_modal_knowledge_retrieval_benchmark_M2KR_CN
PreFLMR M2KR Dataset Card
Dataset details
Dataset type:
M2KR is a benchmark dataset for multimodal knowledge retrieval. It contains a collection of tasks and datasets for training and evaluating multimodal knowledge retrieval models.
We pre-process the datasets into a uniform format and write several task-specific prompting instructions for each dataset. The details of the instruction can be found in the paper. The M2KR benchmark contains three types of tasks:… See the full description on the dataset page: https://huggingface.co/datasets/BByrneLab/multi_task_multi_modal_knowledge_retrieval_benchmark_M2KR_CN.MixBench25MixBench is a benchmark for evaluating mixed-modality retrieval. It contains queries and corpora from four datasets: MSCOCO, Google_WIT, VisualNews, and OVEN. Each subset provides: query, corpus, mixed_corpus, and qrel splits.Multi-modal_dataset_for_Multi-Vehicle_Sensing_Aided_Precoding
Dataset for "SoM-Aided FDD Precoding with Sensing Heterogeneity: A Vertical Federated Learning Approach"
📌 Overview
This dataset is constructed based on the triangular block scenario in the M3SC dataset, including CSI data between the roadside unit and seven passing vehicles, downlink received signal data of vehicles, 64-line LiDAR point cloud data collected by vehicles, RGB image data collected by vehicles, and GPS positioning data of vehicles. The above-mentioned… See the full description on the dataset page: https://huggingface.co/datasets/pku-pcni-lab/Multi-modal_dataset_for_Multi-Vehicle_Sensing_Aided_Precoding.SAE_activations_modal_sentencesbimanual-table-cleanup-cross-embodiment-rich-modality-sample
Cross-Embodiment Bimanual Table Cleanup — Rich-Modality 10-Episode Inspection Sample
10 full-modality cross-embodiment bimanual table-cleanup episodes: 5 Franka Panda + 5 WidowXAI, 21,267 frames, 6 RGB views per robot, task-camera depth and segmentation, native robot state/action, end-effector trajectories, 6-DoF object poses, and QA annotations.
✅ Use it / ❌ Skip it
Use it for
Inspecting loaders, schemas, camera coverage, depth, segmentation, object poses… See the full description on the dataset page: https://huggingface.co/datasets/ExylosAi/bimanual-table-cleanup-cross-embodiment-rich-modality-sample.Multi-modal-Self-instruct
Dataset Description
Paper Information
Dataset Examples
Leaderboard
Dataset Usage
Data Downloading
Data Format
Evaluation
Citation
You can download the zip dataset directly, and both train and test subsets are collected in Multi-modal-Self-instruct.zip.
Dataset Description
Multi-Modal Self-Instruct dataset utilizes large language models and their code capabilities to synthesize massive abstract images and visual reasoning instructions across daily scenarios. This benchmark… See the full description on the dataset page: https://huggingface.co/datasets/zwq2018/Multi-modal-Self-instruct.ModalityFaultLines-SCEval
SCEval — Modality Fault Lines
Data for Modality Fault Lines: Structural Corruptions Reveal Fragile Omni-Modal Reasoning (Findings of EMNLP 2026).
SCEval is a human-verified benchmark for omni-modal robustness. Text, vision, and audio all remain
present, but controlled corruptions make the evidence inside a channel unreliable. Each corrupted
item is paired with its clean counterpart at the example level, so clean-to-corrupted comparisons are
made on the same underlying question… See the full description on the dataset page: https://huggingface.co/datasets/KZL96/ModalityFaultLines-SCEval.NEXUS-temporal_hierarchical_multi-modal
NEXUS: Neural Evolution for eXtensible Universal Semantics Dataset
(Temporal Multimodal Slices)
This dataset is a multi-modal, hierarchical, temporal representation derived from HuggingFaceFV/finevideo. It is designed for streaming training where the primary unit is a 10 ms "slice" that aggregates upward into moments (100 ms), seconds (1 s), experiences (10 s), and minutes (60 s).
It is meant to represent an extensible stream of "experience" as there are… See the full description on the dataset page: https://huggingface.co/datasets/Ardea/NEXUS-temporal_hierarchical_multi-modal.terminal-bench-2.1MixBench2026
MixBench: A Benchmark for Mixed Modality Retrieval
MixBench is a benchmark for evaluating retrieval across text, images, and multimodal documents. It is designed to test how well retrieval models handle queries and documents that span different modalities, such as pure text, pure images, and combined image+text inputs.
MixBench includes four subsets, each curated from a different data source:
MSCOCO
Google_WIT
VisualNews
OVEN
Each subset contains:
queries.jsonl: each entry… See the full description on the dataset page: https://huggingface.co/datasets/mixed-modality-search/MixBench2026.coffee-brew-cross-embodiment-rich-modality-sample
Cross-Embodiment Coffee Brew — Rich-Modality 10-Episode Inspection Sample
10 full-modality cross-embodiment coffee-brew episodes: 5 single-arm Franka Panda + 5 single-arm WidowXAI, 14,401 frames, 5 RGB views per robot, task-camera depth and segmentation, native robot state/action, end-effector trajectories, 6-DoF object poses, and QA annotations.
✅ Use it / ❌ Skip it
Use it for
Inspecting loaders, schemas, camera coverage, depth, segmentation, object poses… See the full description on the dataset page: https://huggingface.co/datasets/ExylosAi/coffee-brew-cross-embodiment-rich-modality-sample.multi-modal-derived-brain-network
PPMI Connectivity Graphs — HF Staging (Derivatives)
This dataset ships ready-to-use functional brain connectivity graphs derived from the PPMI cohort in a BIDS-ish derivatives layout. For each subject and parcellation, we include:
ROI time-series (*_desc-timeseries_parc-<name>.mat)
Pearson correlation connectivity matrix (*_desc-correlation_matrix_parc-<name>.mat)
JSON sidecars with summary fields (nodes, measure, symmetric/weighted flags)
Contents
data/… See the full description on the dataset page: https://huggingface.co/datasets/pakkinlau/multi-modal-derived-brain-network.multi_task_multi_modal_knowledge_retrieval_benchmark_M2KRfrontier-csautoinference-agentic-mix-v1
Autoinference Agentic Mix v1
This is a prompt set for the online_agentic serving benchmark. That profile stands
in for long-horizon agent traffic: a large context that grows turn over turn, with
short structured outputs at each step. The usual way to run it uses
generated-shared-prefix, which builds a synthetic shared prefix out of random tokens.
This dataset uses real agent trajectories instead, so the prefix reuse, the context
growth, and the token mix all match what an agent… See the full description on the dataset page: https://huggingface.co/datasets/modal-labs/autoinference-agentic-mix-v1.object-sorting-cross-embodiment-rich-modality-sample
Cross-Embodiment Object Sorting — Rich-Modality 20-Episode Inspection Sample
20 episodes total: 10 Franka Panda + 10 WidowXAI. A compact cross-embodiment inspection release for picking up an instructed object and placing it into an instructed target box. Each robot keeps its native LeRobot v2.1 state/action schema in a separate Viewer config. Five synchronized RGB views, metric depth and instance segmentation for every camera, robot state/action, end-effector trajectories… See the full description on the dataset page: https://huggingface.co/datasets/ExylosAi/object-sorting-cross-embodiment-rich-modality-sample.swebenchproMulti-modal_dataset_named_M3SC
M³SC: A generic dataset for mixed multi-modal (MMM) sensing and communication integration
📌 Overview
M³SC is the first simulation dataset for communication and multi-modal sensing for connected intelligent vehicles, covering multiple typical environments such as urban, suburban, and rural areas, and comprehensively considering scenario conditions including multi-weather, multi-time periods, multi-vehicle traffic densities, multi-frequency bands, and multi-antenna arrays.… See the full description on the dataset page: https://huggingface.co/datasets/pku-pcni-lab/Multi-modal_dataset_named_M3SC.Multi-modal_dataset_for_LiDAR-Aided_CSI_Estimation
Dataset for "Synesthesia of Machines-Enhanced Wideband Multi-User CSI Learning with LiDAR Sensing"
📌 Overview
This dataset is constructed based on urban crossroad scenario from the M3SC dataset and includes CSI data between a roadside unit and three passing vehicles, as well as 64-line LiDAR point cloud data collected by the roadside devices (processed using the lightweight data processing approach proposed in the paper). The dataset contains a total of 4,500 samples… See the full description on the dataset page: https://huggingface.co/datasets/pku-pcni-lab/Multi-modal_dataset_for_LiDAR-Aided_CSI_Estimation.modality-prefs-data
Modality biases in human preference judgments of LLM responses
Code, data, and stimuli for the paired-preference study comparing how
participants judge LLM responses delivered as text vs as audio
(TTS). The release lets you (a) reproduce every analysis in the paper
on the published anonymized data, and (b) deploy the same survey app to
collect new preferences on your own stimuli.
What's here
modality-preference-elicitation/
├── README.md ← you are… See the full description on the dataset page: https://huggingface.co/datasets/NeurIPS-Anon-2784/modality-prefs-data.Cross-Modal_Retrieval_BigEarthNet_14K_S1_and_S2dissolveCredits: Modal Labs.
Source: https://github.com/genmoai/mochi/blob/main/contrib/modal/main.py#L45
MiMIC_Multi-Modal_Indian_Earnings_Calls_DatasetMiMIC_analysis_code.ipynb : This is the actual code file with results
MiMIC_Multi-Modal_Indian_Earnings_Calls.xlsx : This is the main file having to be used for training. Descripton of columns of this file is given below.
getting_all_texts_together_embeddings_dim128_CPU.pkl : This is the file having embeddings of texts (extracted from transcripts, images) and tables (extracted from images) concatenated one after the other.… See the full description on the dataset page: https://huggingface.co/datasets/sohomghosh/MiMIC_Multi-Modal_Indian_Earnings_Calls_Dataset.vertical-modality-task-taxonomy-v1
vertical-modality-task-taxonomy-v1
Short Description
Synthetic taxonomy mapping GCC governance verticals to modalities, ML tasks, libraries, formats, and safety boundaries.
Purpose
vertical-modality-task-taxonomy-v1 is a governance-first planning taxonomy for the GCC Governance Intelligence Stack.
It organizes platform verticals, product/domain lines, business functions, modalities, ML task clusters, recommended libraries, recommended formats, safe outputs… See the full description on the dataset page: https://huggingface.co/datasets/BDR-AI/vertical-modality-task-taxonomy-v1.multi-modal-peg-in-square-hole-testThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "ur5",
"total_episodes": 51,
"total_frames": 8074,
"total_tasks": 1,
"total_videos": 153,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 10,
"splits": {
"train": "0:51"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/hainh22/multi-modal-peg-in-square-hole-test.autoinference-agentic-mix-v2
Autoinference Agentic Mix v2
300 real SWE-agent trajectories from
TIGER-Lab/SWE-Next-SFT-Trajectories,
expanded into one request per assistant turn. Each row carries the conversation
up to that turn and the model generates the turn. Replaying a trajectory in
turn_index order re-sends a growing prefix, which is how an agent loop
actually hits a prefix cache.
What changed from v1
v1 kept only requests with at least 34k prefix tokens. That cut trajectories
down to… See the full description on the dataset page: https://huggingface.co/datasets/modal-labs/autoinference-agentic-mix-v2.swebench-verifiedswegym-lite
