datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Omni-Fake-SET
Omni-Fake-SET
Omni-Fake-SET is the in-distribution split of Omni-Fake, a unified multimodal deepfake dataset for social-media forensics. It covers image, audio, video, and audio–video talking-head (AV-TH) modalities. Each modality uses the same three-way label space: real, fully synthetic, and tampered. Pair with the held-out benchmark Omni-Fake-OOD for out-of-distribution evaluation.
Paper: arXiv:2605.01638
Project page: Omni-Fake
License: CC-BY-4.0
Video (hybrid… See the full description on the dataset page: https://huggingface.co/datasets/JamalLee/Omni-Fake-SET.OmniFake
OmniFake
OmniFake is a large-scale, well-categorized synthetic image dataset introduced in Few-Shot Synthetic Image Attribution: Identifying Unseen Generators with Limited Samples. It contains 1.17 million AI-generated images from 45 distinct generators, paired with 1.17 million real images, designed for research on AI-generated image (AIGI) detection and source attribution.
For usage instructions and experimental protocols, please refer to the OmniDFA GitHub repository.… See the full description on the dataset page: https://huggingface.co/datasets/MoeNew/OmniFake.Omni-Fake-OOD
Omni-Fake-OOD
Omni-Fake-OOD is the out-of-distribution benchmark split of Omni-Fake. Samples come from held-out generators and platforms not included in training, for measuring cross-domain generalization. It covers image, audio, video, and audio–video talking-head (AV-TH) with the same three-class labels as Omni-Fake-SET: real, fully synthetic, and tampered. Use together with Omni-Fake-SET (in-distribution training data).
Paper: arXiv:2605.01638
Project page: Omni-Fake… See the full description on the dataset page: https://huggingface.co/datasets/JamalLee/Omni-Fake-OOD.Omnitraffic_Dataset
🚗 OmniTraffic: A Large-scale Multi-view Spatiotemporal Dataset for Traffic Understanding
📌 Dataset Summary
Welcome to the OmniTraffic Dataset repository. This repository specifically hosts the complete OmniTraffic Dataset, containing the massive underlying pool of over 8 million generated VQA samples and ~280GB of multimodal data. It is designed for large-scale pre-training, fine-tuning, and pushing the scaling laws of multimodal large language models (MLLMs) and… See the full description on the dataset page: https://huggingface.co/datasets/CROHuang/Omnitraffic_Dataset.satellite-multitask-omni
🛰️ Satellite Multi-Task Omni Dataset
A unified, multi-task satellite/aerial imaging dataset designed for training omni-models that work with image+text as both input and output modalities. All data is converted to a consistent ChatML conversational format.
📊 Dataset Overview
Metric
Value
Total Samples
34,894
Train / Val / Test
31,404 / 1,744 / 1,746
Tasks
9 distinct task types
Sources
10 source datasets
Format
ChatML conversations + images… See the full description on the dataset page: https://huggingface.co/datasets/rahuldshetty/satellite-multitask-omni.
