datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
SID_Set
Dataset Card for SID_Set
Dataset Summary
We provide Social media Image Detection dataSet (SID-Set), which offers three key advantages:
Extensive volume: Featuring 300K AI-generated/tampered and authentic images with comprehensive annotations.
Broad diversity: Encompassing fully synthetic and tampered images across various classes.
Elevated realism: Including images that are predominantly indistinguishable from genuine ones through mere visual inspection.
Please check… See the full description on the dataset page: https://huggingface.co/datasets/saberzl/SID_Set.colpali_train_set
Dataset Description
This dataset is the training set of ColPali it includes 127,460 query-image pairs from both openly available academic datasets (63%) and a synthetic dataset made up
of pages from web-crawled PDF documents and augmented with VLM-generated (Claude-3 Sonnet) pseudo-questions (37%).
Our training set is fully English by design, enabling us to study zero-shot generalization to non-English languages.
Dataset
#examples (query-page pairs)
Language
DocVQA
39… See the full description on the dataset page: https://huggingface.co/datasets/vidore/colpali_train_set.So-Fake-Set
Dataset Card for So-Fake-Set
Dataset Summary
We provide So-Fake-Set, A large-scale, diverse dataset tailored for social media image forgery detection!
Please check our website to explore more visual results.
Dataset Structure
"image" (Image): Input images, including real, full_synthetic, and tampered images.
"mask" (Image): Binary mask highlighting manipulated regions in tampered images.
"label" (str): Classification category.
"generator" (str): The… See the full description on the dataset page: https://huggingface.co/datasets/saberzl/So-Fake-Set.Omni-Fake-SET
Omni-Fake-SET
Omni-Fake-SET is the in-distribution split of Omni-Fake, a unified multimodal deepfake dataset for social-media forensics. It covers image, audio, video, and audio–video talking-head (AV-TH) modalities. Each modality uses the same three-way label space: real, fully synthetic, and tampered. Pair with the held-out benchmark Omni-Fake-OOD for out-of-distribution evaluation.
Paper: arXiv:2605.01638
Project page: Omni-Fake
License: CC-BY-4.0
Video (hybrid… See the full description on the dataset page: https://huggingface.co/datasets/JamalLee/Omni-Fake-SET.diffusion-pretrain-set-ft1
diffusion-pretrain-set-ft1
A multi-source image-caption pretraining dataset assembled from ten upstream
sources via a uniform ingest pipeline. Designed for a full pretrain or finetune
pipeline meant to curate for any major diffusion model preliminary, with the sole
intent to create a more powerful baseline preliminary train and a baseline
for synthesizing images to train the next generation of the VLM model.
This is a lot like the snake eating it's own tail, so it must be… See the full description on the dataset page: https://huggingface.co/datasets/AbstractPhil/diffusion-pretrain-set-ft1.Omni-Sets
Omni-Sets
A large-scale, multi-modal instruction-tuning dataset spanning six modalities (audio, speech, image, video, visual documents, and cross-modal omni) with both single-turn dense captions and multi-turn instruction-following conversations. Designed for training omni-modal language models that can perceive and reason across all modalities.
590,858 total samples | 5,635 hours of audio/video | 6 configs | 17 source datasets
Overview
Config
Modality… See the full description on the dataset page: https://huggingface.co/datasets/MBZUAI/Omni-Sets.diffbir-mixed-setsdiffusion-pretrain-set-ft1-1024
diffusion-pretrain-set-ft1-1024
1024px (2x) upscale of AbstractPhil/diffusion-pretrain-set-ft1.
WARNING
MUCH OF THIS DATA WAS MODEL UPSCALED USING RAPID UPSCALERS.
THIS IS NOT CONSISTENTLY HIGH FIDELITY NOR IS IT EVEN CLOSE TO FAIR FIDELITY AT TIMES.
PLEASE use this ONLY for pretraining, new concepts, and simple design purposes ONLY. HEAVILY PRUNE FOR FINETUNING.
Thank you, good luck my friends.
Details
Model: realesr-general-x4v3 (SRVGG Compact… See the full description on the dataset page: https://huggingface.co/datasets/AbstractPhil/diffusion-pretrain-set-ft1-1024.lebanese_aug_setSID_Set
Dataset Card for SID_Set
Dataset Summary
We provide Social media Image Detection dataSet (SID-Set), which offers three key advantages:
Extensive volume: Featuring 300K AI-generated/tampered and authentic images with comprehensive annotations.
Broad diversity: Encompassing fully synthetic and tampered images across various classes.
Elevated realism: Including images that are predominantly indistinguishable from genuine ones through mere visual inspection.
Please… See the full description on the dataset page: https://huggingface.co/datasets/RAID-techjam/SID_Set.So-Fake-Set-Resized-224emu_edit_test_set
Dataset Card for the Emu Edit Test Set
Dataset Summary
To create a benchmark for image editing we first define seven different categories of potential image editing operations: background alteration (background), comprehensive image changes (global), style alteration (style), object removal (remove), object addition (add), localized modifications (local), and color/texture alterations (texture).
Then, we utilize the diverse set of input images from the MagicBrush… See the full description on the dataset page: https://huggingface.co/datasets/facebook/emu_edit_test_set.SID_Set
Dataset Card for SID_Set
Dataset Summary
We provide Social media Image Detection dataSet (SID-Set), which offers three key advantages:
Extensive volume: Featuring 300K AI-generated/tampered and authentic images with comprehensive annotations.
Broad diversity: Encompassing fully synthetic and tampered images across various classes.
Elevated realism: Including images that are predominantly indistinguishable from genuine ones through mere visual inspection.
Please… See the full description on the dataset page: https://huggingface.co/datasets/Beastarz/SID_Set.colpali_train_set_split_by_sourcefast-food-floor-waste-grasping-training-set-next-pack-9f7b7681-1106dcde
Fast-Food Cleaning Robot — Floor Mess Dataset
Training dataset for a cleaning robot operating in fast-food-style food-service spaces (break areas / dining). Scenes are staged in break-area environments cluttered with food-service furnishings and food items (pizza, grocery food, cups, spoons) so the robot learns to perceive and act on mess. Covers detection, grasping, navigation, obstacle avoidance and pick-and-place. Renders are 1024x1024 with RGB plus albedo, metric depth and… See the full description on the dataset page: https://huggingface.co/datasets/physicl-community/fast-food-floor-waste-grasping-training-set-next-pack-9f7b7681-1106dcde.indexed-open-image-v4-test-set
Dataset Card for "indexed-open-image-v4-test-set"
More Information needed
cloth_full_dt0.05_rgb_v1_0This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "custom_mpm_robot",
"total_episodes": 359,
"total_frames": 212027,
"total_tasks": 1,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 20,
"splits": {
"train": "0:359"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Setsunainn/cloth_full_dt0.05_rgb_v1_0.concatenated-train-setemu_edit_test_set_generations
Dataset Card for the Emu Edit Generations on Emu Edit Test Set
Dataset Summary
This dataset contains Emu Edit's generations on the Emu Edit test set. For more information please read our paper or visit our homepage.
Licensing Information
Licensed with CC-BY-NC 4.0 License available here.
Citation Information
@inproceedings{Sheynin2023EmuEP,
title={Emu Edit: Precise Image Editing via Recognition and Generation Tasks},
author={Shelly Sheynin and… See the full description on the dataset page: https://huggingface.co/datasets/facebook/emu_edit_test_set_generations.colpali-train-set-splitted-translatedcube_transfer_green_temp_setupmpm_cloth-turn-128-rgb-v0-turnThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "custom_mpm_robot",
"total_episodes": 128,
"total_frames": 17408,
"total_tasks": 1,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 20,
"splits": {
"train": "0:128"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Setsunainn/mpm_cloth-turn-128-rgb-v0-turn.external_test_set_v1Dense-Set
Dense-Set
Dense-Set is a curated benchmark of visually dense scenes for text-to-image retrieval evaluation. It provides challenging subsets extracted from COCO and Flickr30K, focusing on crowded images with multiple object instances and underrepresented, low-attention classes.
This dataset is published alongside:
LARE: Low-Attention Region Encoding for Text–Image Retrieval
ICML 2026 Workshop on Efficient Multimodal Question Answering (EMM-QA) — Workshop Page
Project Page | Code… See the full description on the dataset page: https://huggingface.co/datasets/aalquwayfili/Dense-Set.mpm_cloth-drag_and_turn-128-rgb-v0-turnThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "custom_mpm_robot",
"total_episodes": 128,
"total_frames": 17408,
"total_tasks": 1,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 20,
"splits": {
"train": "0:128"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Setsunainn/mpm_cloth-drag_and_turn-128-rgb-v0-turn.mpm_cloth-drag_and_turn-128-depth-v2.0This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "custom_mpm_robot",
"total_episodes": 128,
"total_frames": 24192,
"total_tasks": 1,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 20,
"splits": {
"train": "0:128"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Setsunainn/mpm_cloth-drag_and_turn-128-depth-v2.0.gwhd2021_set0
Global Wheat Head Detection (GWHD) dataset
Cite as:
David, Etienne et al. (2020). Global Wheat Head Detection (GWHD) dataset: a large and diverse dataset of high-resolution RGB-labelled images to develop and benchmark wheat head detection methods. Plant Phenomics, 2020. Science Partner Journal.DOI: 10.5281/zenodo.5092309
License: CC-BY-4.0
mpm_cloth-drag_and_turn-128-rgb-v0-dragThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "custom_mpm_robot",
"total_episodes": 128,
"total_frames": 6784,
"total_tasks": 1,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 20,
"splits": {
"train": "0:128"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Setsunainn/mpm_cloth-drag_and_turn-128-rgb-v0-drag.mpm_cloth-turn-128-rgb-v0This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "custom_mpm_robot",
"total_episodes": 128,
"total_frames": 24192,
"total_tasks": 1,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 20,
"splits": {
"train": "0:128"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Setsunainn/mpm_cloth-turn-128-rgb-v0.medix-eval-sets
