datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
dragon
Dataset Card for DRAGON
🧾 ArXiv Preprint
DRAGON is a large-scale Dataset of Realistic imAges Generated by diffusiON models.
The dataset includes a total of 2.5 million training images and 100,000 test images generated using 25 diffusion models, spanning both recent advancements and older, well-established architectures.
Dataset Details
Dataset Description
The remarkable ease of use of diffusion models for image generation has led to a proliferation of… See the full description on the dataset page: https://huggingface.co/datasets/lesc-unifi/dragon.UniWorld-V1
The Geneval-style dataset is sourced from BLIP3o-60k.
This dataset is presented in the paper: UniWorld: High-Resolution Semantic Encoders for Unified Visual Understanding and Generation
More details can be found in UniWorld-V1
Data preparation
Download the data from LanguageBind/UniWorld-V1. The dataset consists of two parts: source images and annotation JSON files.
Prepare a data.txt file in the following format:
The first column is the root path to the image.
The second… See the full description on the dataset page: https://huggingface.co/datasets/LanguageBind/UniWorld-V1.majestrino-unified-detailed-captions
Majestrino Unified Detailed Captions
Filtered subset of laion/majestrino-data containing all samples with unified_detailed_caption.
Stats
4,658,407 samples
932 tar files (~1.1 GB each)
~1,017 GB total
Format
Each tar contains paired .flac + .json files.
JSON fields:
caption — the unified detailed caption
caption_type — always unified_detailed_caption
transcription — speech transcription (when available, normalized from multiple source keys)
duration — audio… See the full description on the dataset page: https://huggingface.co/datasets/TTS-AGI/majestrino-unified-detailed-captions.cc12m-wds-coco-recaptioned
CC12M WebDataset with COCO-style Recaptions
A large-scale image-text dataset containing 3 million images from Conceptual Captions 12M (CC12M) with COCO-style factual descriptions generated using NVIDIA Nemotron Nano 12B v2 VL.
Dataset Overview
Base Dataset: pixparse/cc12m-wds - Conceptual Captions 12M (CC12M)
Images: 3,000,000+ high-quality internet images
Recaption Model: NVIDIA Nemotron Nano 12B v2 VL
Recaption Style: COCO-style factual descriptions (20 words average)… See the full description on the dataset page: https://huggingface.co/datasets/undefined443/cc12m-wds-coco-recaptioned.UnifiedReward-2.0-T2X-score-data
Dataset Summary
UnifiedReward-2.0-T2X-score-data is added for our UnifiedReward-2.0-qwen-[3b/7b/32b/72b] training.
This dataset enables UnifiedReward-2.0 introducing several new capabilities:
Pairwise scoring for image and video generation assessment on Alignment, Coherence, Style dimensions.
Pointwise scoring for image and video generation assessment on Alignment, Coherence/Physics, Style dimensions.
Welcome to try the latest version, and the inference code is available at… See the full description on the dataset page: https://huggingface.co/datasets/CodeGoat24/UnifiedReward-2.0-T2X-score-data.UniSAR-7M
UniSAR-7M
A large-scale, multi-source synthetic aperture radar image corpus for self-supervised representation learning.
UniSAR-7M contains 7,047,666 single-channel SAR image samples assembled from public SAR datasets and openly available imagery from commercial satellite constellations. It provides the pretraining corpus for DINOSAR, a self-supervised learning framework that uses Content-Aware Multi-Crop (CAMC) to construct informative views of SAR imagery.
Associated… See the full description on the dataset page: https://huggingface.co/datasets/YTang/UniSAR-7M.unsupervised_peoples_speech_raw_voice_activity_detection_snippets_part_1mmtrailer-pe-av-unimodal-local-embeddingsyt-pe-av-unimodal-local-embeddingsUnifiedReward-Flex-SFT-90K
UnifiedReward-Flex-SFT-90K
This repository releases 90K SFT data of UnifiedReward-Flex.
For further details, please refer to the following resources:
📰 Paper: https://arxiv.org/abs/2602.02380
🪐 Project Page: https://codegoat24.github.io/UnifiedReward/flex
🤗 Model Collections: https://huggingface.co/collections/CodeGoat24/unifiedreward-flex
🤗 Dataset: https://huggingface.co/datasets/CodeGoat24/UnifiedReward-Flex-SFT-90K
👋 Point of Contact: Yibin Wang
Citation… See the full description on the dataset page: https://huggingface.co/datasets/CodeGoat24/UnifiedReward-Flex-SFT-90K.coco-karpathy-wds
COCO-2014 WebDataset Format (Karpathy Splits)
This dataset contains the COCO-2014 images and captions converted to WebDataset (WDS) format, using the Karpathy & Li (2015) dataset split for image captioning tasks.
Overview
Total Samples: 123,287 images with 5 reference captions each
Total Size: ~19 GB
Format: WebDataset (.tar shards)
Shard Size: 1,000 samples per tar file
License: CC-BY 4.0
Language: English
Structure
COCO-2014-WDS/
├── train/ (113… See the full description on the dataset page: https://huggingface.co/datasets/undefined443/coco-karpathy-wds.UNO1m-filtered-splitOBELICS_HQ_5M_UniFilter
UniFilter Synthetic Training Data (unifilter_train_data)
This repository contains UniFilter-Post-Train-Data, the synthetic training data used for the UniFilter model, as presented in the paper Train a Unified Multimodal Data Quality Classifier with Synthetic Data.
UniFilter is an efficient Multimodal Large Language Model (MLLM) designed as a Unified Multimodal Data Quality Classifier. It filters high-quality image-text caption and interleaved document data by generating quality… See the full description on the dataset page: https://huggingface.co/datasets/weizhiwang/OBELICS_HQ_5M_UniFilter.UniHand_PreviewThis data is a subset of the pretraining data for Being-H0.5.
Citation
Being-H0.5
@article{beingbeyond2026beingh05,
title={Being-H0.5: Scaling Human-Centric Robot Learning for Cross-Embodiment Generalization},
author={Luo, Hao and Wang, Ye and Zhang, Wanpeng and Zheng, Sipeng and Xi, Ziheng and Xu, Chaoyi and Xu, Haiweng and Yuan, Haoqi and Zhang, Chi and Wang, Yiqing and Feng, Yicheng and Lu, Zongqing},
journal={arXiv preprint arXiv:2601.12993},
year={2026}
}
Being-H0… See the full description on the dataset page: https://huggingface.co/datasets/BeingBeyond/UniHand_Preview.Uni-Edit-Train-Data
Uni-Edit Training Data: Uni-Edit-148k
Project Page | GitHub Repository | Paper
👀 Intro
We introduce Uni-Edit, an intelligent image editing task that serves as the first general task for Unified Multimodal Model (UMM) tuning. Unlike conventional mixed multi-task training that suffers from inherent task conflicts and requires complex multi-stage pipelines, Uni-Edit breaks this paradigm. It achieves true mutual reinforcement by improving image… See the full description on the dataset page: https://huggingface.co/datasets/Uni-Edit/Uni-Edit-Train-Data.Unified_Road_Defect_Dataset
Unified Road Defect Dataset
A merged, YOLO-format road-defect detection dataset that combines RDD-2022
(primary, ground-level, 6 countries) with two supplementary aerial/drone
datasets — UAV-PDD2023 (China) and RoadDamageVision (China + Spain) —
into a single 4-class CRDDC schema.
This is a derived dataset. It re-packages and re-labels images from three
independently published sources. All credit for the underlying images and
original annotations belongs to their respective… See the full description on the dataset page: https://huggingface.co/datasets/TamAko783/Unified_Road_Defect_Dataset.mydataundefined2movielens-pe-av-unimodal-local-embeddingsunidisc_hqThis repository contains the dataset used in the paper Unified Multimodal Discrete Diffusion.
Code: https://github.com/AlexSwerdlow/unidisc
Additionally, we release a synthetic dataset available here and the corresponding generation scripts as well as the raw data.
Uni-Januspdm_carlamajestrino-unified-detailed-captions-temporal
Majestrino Unified Detailed Captions with Temporal Aspects
Filtered subset of laion/majestrino-data containing only samples with unified_detailed_caption_with_temporal_aspects.
Stats
4,128,665 samples
826 tar files (~1.1 GB each)
~878 GB total
Format
Each tar contains paired .flac + .json files.
JSON fields:
caption — the unified detailed caption with temporal aspects
caption_type — always unified_detailed_caption_with_temporal_aspects
transcription — speech… See the full description on the dataset page: https://huggingface.co/datasets/TTS-AGI/majestrino-unified-detailed-captions-temporal.Uni10K
Uni10K from
A LoD of Gaussians: Unified Training and Rendering for Ultra-Large-Scale Reconstruction with External Memory
Felix Windisch1
Thomas Köhler1
Lukas Radl1
Mattia D'Urso1
Michael Steiner1
Dieter Schmalstieg1,2
Markus Steinberger1,3
1Graz University of Technology,
2University of Stuttgart,
3Huawei Technologies… See the full description on the dataset page: https://huggingface.co/datasets/mattia-durso/Uni10K.parczech4speech-unsegmented
ParCzech4Speech (Unsegmented Variant)
Dataset Summary
ParCzech4Speech (Unsegmented Variant) is a large-scale Czech speech dataset derived from parliamentary recordings and official transcripts.
This variant captures continuous speech segments without enforcing sentence boundaries, making it well-suited for real-world streaming ASR scenarios
and speech modeling tasks that benefit from natural discourse flow.
The dataset is created using a combination of WhisperX and… See the full description on the dataset page: https://huggingface.co/datasets/ufal/parczech4speech-unsegmented.Brats2021_slicedUniHand_PreviewThis data is a subset of the pretraining data for Being-H0.5.
Citation
Being-H0.5
@article{beingbeyond2026beingh05,
title={Being-H0.5: Scaling Human-Centric Robot Learning for Cross-Embodiment Generalization},
author={Luo, Hao and Wang, Ye and Zhang, Wanpeng and Zheng, Sipeng and Xi, Ziheng and Xu, Chaoyi and Xu, Haiweng and Yuan, Haoqi and Zhang, Chi and Wang, Yiqing and Feng, Yicheng and Lu, Zongqing},
journal={arXiv preprint arXiv:2601.12993},
year={2026}
}
Being-H0… See the full description on the dataset page: https://huggingface.co/datasets/zdx123222/UniHand_Preview.UniDet3D@misc{kolodiazhnyi2024unidet3dmultidatasetindoor3d,
title={UniDet3D: Multi-dataset Indoor 3D Object Detection},
author={Maksim Kolodiazhnyi and Anna Vorontsova and Matvey Skripkin and Danila Rukhovich and Anton Konushin},
year={2024},
eprint={2409.04234},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2409.04234},
}
unarxive-2024-processedSpotify-unsorted-300kSpotify-unsorted-300k is a dataset made from 307920 images of albums/singles, without any sorting.
Please feel free to modify and utilize the scraper.
