datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
DanQing100M
100M Chinese image-text pairs | 12TB dataset | 2024-2025 web data
DanQing: An Up-to-Date Large-Scale Chinese Vision-Language Pre-training Dataset
Project Page | Paper | Code
Hengyu Shen∗, Tiancheng Gu∗, Bin Qin, Lan Wu, Yuling Wu, Shuo Tan, Zelong Sun, Jun Wang, Nan Wu, Xiang An, Weidong Cai, Ziyong Feng‡, Kaicheng Yang†
∗ Equal Contribution | ‡ Team Leader | † Project Leader
📣 News
[2026/01/16] ✨ We release the paper of DanQing.
[2026/01/15] 🔥 We release the… See the full description on the dataset page: https://huggingface.co/datasets/DeepGlint-AI/DanQing100M.RealSyn100M
[ACM MM25] RealSyn: An Effective and Scalable Multimodal Interleaved Document Transformation Paradigm
Tiancheng Gu,
Kaicheng Yang,
Chaoyi Zhang,
Yin Xie,
Xiang An,
Ziyong Feng,
Dongnan Liu,
Weidong Cai,
Jiankang Deng
💡 Introduction
Contrastive Language-Image Pre-training (CLIP) demonstrates promising performance on a wide variety of benchmarks. However, a substantial volume of non-paired data, such as multimodal interleaved documents, remains… See the full description on the dataset page: https://huggingface.co/datasets/Kaichengalex/RealSyn100M.lunarsim-shoemaker-traverse-100m-v1
LunarSim Rover Traverse Synthetic Dataset V1
Overview
Synthetic lunar surface dataset generated using
LunarSim, an NVIDIA Isaac
Sim-based lunar environment simulator. Images simulate a forward-facing rover
camera traversing terrain near the Shoemaker crater at the lunar south pole.
The dataset includes RGB images, depth maps, semantic/instance segmentation
masks, shadow masks, and COCO-format rock detection annotations.
Generation
Simulator: LunarSim (NVIDIA… See the full description on the dataset page: https://huggingface.co/datasets/elementrobotics/lunarsim-shoemaker-traverse-100m-v1.laion_text_debiased_100M
100M Text Debiased Subset from LAION 2B
Captions in LAION-2B have a significant bias towards describing visual text content embedded in the images.
Released CLIP models have strong text spotting bias in almost every style of web images, resulting in the CLIP-filtering datasets inherently biased towards visual text dominant data.
CLIP models easily learn text spotting capacity from parrot captions while failing to connect the vision-language semantics, just like a text spotting… See the full description on the dataset page: https://huggingface.co/datasets/linyq/laion_text_debiased_100M.Arabic-Image-Captioning_100M
Arabic Image Captioning Dataset (100M Sample)
The first large-scale Arabic multimodal dataset.
This groundbreaking dataset contains 100 million Arabic image captions, representing the first comprehensive Arabic multimodal resource of this scale and quality. Generated using our Mutarjim translation model, this dataset addresses the critical gap in Arabic multimodal AI resources and enables researchers to develop sophisticated Arabic vision-language systems for the first time.… See the full description on the dataset page: https://huggingface.co/datasets/Misraj/Arabic-Image-Captioning_100M.yoco_moon_100mpx100m_kmweibull_k_100m_cogwind_speed_100m_cog
