datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
RenderedTextThis dataset has been created by Stability AI and LAION.
This dataset contains 12 million 1024x1024 images of handwritten text written on a digital 3D sheet of paper generated using Blender geometry nodes and rendered using Blender Cycles. The text has varying font size, color, and rotation, and the paper was rendered under random lighting conditions.
Note that, the first 10 million examples are in the root folder of this dataset repository and the remaining 2 million are in ./remaining (due… See the full description on the dataset page: https://huggingface.co/datasets/wendlerc/RenderedText.wds_imagenet-rrobopoint-data
RoboPoint Dataset Card
Dataset details
This dataset contains 1432K image-QA instances used to fine-tune RoboPoint, a VLM for spatial affordance prediction. It consists of the following parts:
347K object reference instances from a synthetic data pipeline;
320K free space reference instances from a synthetic data pipeline;
100K object detection instaces from LVIS;
150K GPT-generated instruction-following instances from liuhaotian/LLaVA-Instruct-150K;
515K general-purpose… See the full description on the dataset page: https://huggingface.co/datasets/wentao-yuan/robopoint-data.RevealLayer-100K
RevealLayer Open Dataset
RevealLayer Open is the open-source dataset accompanying RevealLayer: Disentangling Hidden and Visible Layers via Occlusion-Aware Image Decomposition.
Paper: https://arxiv.org/html/2605.11818v1 Accepted by ICML 2026
RevealLayer studies box-guided layered image decomposition for natural images. Given an RGB image and instance bounding boxes, the task is to decompose the scene into a clean background and object-level foreground layers, where each… See the full description on the dataset page: https://huggingface.co/datasets/qihoo360/RevealLayer-100K.cc12m-recaptionedMolmo2-ER-RoboPoint
Molmo2-ER · wentao-yuan/robopoint-data
1.43M robotics affordance instruction-tuning examples (pointing + detection + VQA).
This is a re-hosted, loader-ready subset of the upstream dataset, used to train allenai/Molmo2-ER-4B. Files mirror the upstream layout; nothing in the data has been modified.
Upstream source
Original dataset: wentao-yuan/robopoint-data
Paper: RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics (arXiv:2406.10721)
License:… See the full description on the dataset page: https://huggingface.co/datasets/allenai/Molmo2-ER-RoboPoint.cc12m-wds-coco-recaptioned
CC12M WebDataset with COCO-style Recaptions
A large-scale image-text dataset containing 3 million images from Conceptual Captions 12M (CC12M) with COCO-style factual descriptions generated using NVIDIA Nemotron Nano 12B v2 VL.
Dataset Overview
Base Dataset: pixparse/cc12m-wds - Conceptual Captions 12M (CC12M)
Images: 3,000,000+ high-quality internet images
Recaption Model: NVIDIA Nemotron Nano 12B v2 VL
Recaption Style: COCO-style factual descriptions (20 words average)… See the full description on the dataset page: https://huggingface.co/datasets/undefined443/cc12m-wds-coco-recaptioned.SPI-2M
SPI-2M
We introduce Stylized Pathology Images SPI-2M for stain normalisation via neural style transfer in histopathology.
For full details on dataset sourcing, creation etc please see our paper
Dataset download
The data repo of this repository is organised as follows:
sources: contains the 4096 curated source images zipped together
targets: contains the 512 target images zipped together
stylized: contains 512 .npy files, each has the same index as a corresponding target… See the full description on the dataset page: https://huggingface.co/datasets/R-J/SPI-2M.redcaps5m_resizedobjaverse_rendering_setImageNet-RenditiontestdataObjaverseXL_github_rendersOneMillionFaces
million-faces
Welcome to "million-faces", one of the largest facesets available to the public. Comprising a staggering one million faces, all images in this dataset are entirely AI-generated.
Due to the nature of AI-generated images, please be aware that some artifacts may be present in the dataset.
The dataset is currently being uploaded to Hugging Face, a renowned platform for hosting datasets and models for the machine learning community.
Usage
Feel free to use… See the full description on the dataset page: https://huggingface.co/datasets/RichardErkhov/OneMillionFaces.RORem_datasetre_trMolmo2-ER-RefSpatial
Molmo2-ER · JingkunAn/RefSpatial
2.5M spatial-referring corpus (web + indoor + simulated) covering 31 spatial relations.
This is a re-hosted, loader-ready subset of the upstream dataset, used to train allenai/Molmo2-ER-4B. Files mirror the upstream layout; nothing in the data has been modified.
Upstream source
Original dataset: JingkunAn/RefSpatial
Paper: RoboRefer: Towards Spatial Referring with Reasoning in Vision-Language Models for Robotics (arXiv:2506.04308)… See the full description on the dataset page: https://huggingface.co/datasets/allenai/Molmo2-ER-RefSpatial.wds_vtab-resisc45Recap-Datacomp-1B_tars_part7Final size: 7,236,721, samples per tar: 10000
realcc3m-recap-wdsmSOP-765k
mSOP-765k: A Benchmark For Multi-Modal Structured Output Predictions
The mSOP-765k dataset serves as a benchmark for Multi-Modal Structured Output Predictions.
The dataset contains approximately 765k data, comprising both images and textual data.
Data
Image Data:
The images are cropped from scanned advertisement leaflets.
The image data is divided into train and test splits.
The image dataset is available in two versions: one with images resized so that the longer edge… See the full description on the dataset page: https://huggingface.co/datasets/retail-product-promotion/mSOP-765k.The_Million_Song_DatasetRealBlurRealEstate10KPubTables-v2
PubTables-v2
PubTables-v2 is a new large-scale dataset for full-page and multi-page table extraction.
See also: Hugging Face Paper Page
News
2026 Mar 17: New paper draft with more experiments, especially for multi-page table extraction2026 Feb 11: PubTables-v2 has been officially released on Hugging Face!2025 Dec 11: Our paper is now available on arXiv
Collections
PubTables-v2 comes in 3 collections.
Each collection contains tables in a specific context:… See the full description on the dataset page: https://huggingface.co/datasets/rohanSingh969/PubTables-v2.RSTeller
⚠️ Usage Warning
This is the latest version of RSTeller, updated on 2025-01-28. Users who accessed this dataset before this date can find the legacy version, which is preserved for reference. Additionally, we have released the metadata for this dataset.
For the details and the usage of the dataset, please refer to our github repository page.
Citation
If you find the dataset and our paper useful, please consider citing our paper:
@article{ge2025rsteller… See the full description on the dataset page: https://huggingface.co/datasets/SlytherinGe/RSTeller.Motion324OriAnyV2_Train_Render
Orient Anything V2 Dataset
Project Page | Paper | GitHub
Orient Anything V2 is an enhanced foundation model for unified understanding of object 3D orientation and rotation from single or paired images. This repository contains the training data (final rendering data) used for the model.
Sample Usage
Below is a snippet to run inference using the model and data logic, as found in the official GitHub repository:
import numpy as np
from PIL importImage
import torch
import… See the full description on the dataset page: https://huggingface.co/datasets/Viglong/OriAnyV2_Train_Render.wds_vtab-diabetic_retinopathy
