datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
gs-images-v2pickapic_v2_webdatasetwebdataset archive of yuvalkirstain/pickapic_v2.
Dataloading code can be found here.
PubTables-v2
PubTables-v2
PubTables-v2 is a new large-scale dataset for full-page and multi-page table extraction.
Official dataset evaluation scripts and leaderboard coming soon!
In the meantime, you can create your own evaluation using GriTS with our open-source package, pip install grits-metric.
Report any issues here: https://github.com/kensho-technologies/grits.
See also: Hugging Face Paper Page
News
2026 Apr 15: Code for the GriTS metric released… See the full description on the dataset page: https://huggingface.co/datasets/kensho/PubTables-v2.POEM-v2font-square-v2
Accessing the font-square-v2 Dataset on Hugging Face
The font-square-v2 dataset is hosted on Hugging Face at blowing-up-groundhogs/font-square-v2. It is stored in WebDataset format, with tar files organized as follows:
tars/train/: Contains {000..499}.tar shards for the main training split.
tars/fine_tune/: Contains {000..049}.tar shards for fine-tuning.
Each tar file contains multiple samples, where each sample includes:
An RGB image (.rgb.png)
A black-and-white image (.bw.png)… See the full description on the dataset page: https://huggingface.co/datasets/blowing-up-groundhogs/font-square-v2.Aesthetic-Train-V2
Aesthetic-Train-V2 Dataset
We introduce Aesthetic-Train-V2, a high-quality traing set for ultra-high-resolution image generation.
For more details, please refer to our paper:
Diffusion-4K: Ultra-High-Resolution Image Synthesis with Latent Diffusion Models (CVPR 2025)
Ultra-High-Resolution Image Synthesis: Data, Method and Evaluation
Source code is available at https://github.com/zhang0jhon/diffusion-4k.
Citation
If you find our paper or dataset is helpful in your… See the full description on the dataset page: https://huggingface.co/datasets/zhang0jhon/Aesthetic-Train-V2.PubTables-v2
PubTables-v2
PubTables-v2 is a new large-scale dataset for full-page and multi-page table extraction.
See also: Hugging Face Paper Page
News
2026 Mar 17: New paper draft with more experiments, especially for multi-page table extraction2026 Feb 11: PubTables-v2 has been officially released on Hugging Face!2025 Dec 11: Our paper is now available on arXiv
Collections
PubTables-v2 comes in 3 collections.
Each collection contains tables in a specific context:… See the full description on the dataset page: https://huggingface.co/datasets/rohanSingh969/PubTables-v2.ImageNet-V23d-wm-atomic-v2
3d-wm-atomic-v2 — Synthetic CAD construction videos
Per-op animated frame sequences for CadQuery construction programs, with
per-frame atomic op labels in continuous raw mm. Designed as training
data for video → action (IDM) and image → next frame
(video-gen) models.
Format version: v2-continuous-mm (2026-05-29 onward).
One clip = one (case, view) pair
data_*/{bNNNN}/train/shard-NNNNNN.tar
└── {case_uid}_v{NN}/
├── 0000.png .. NNNN.png # 256x256 rendered… See the full description on the dataset page: https://huggingface.co/datasets/hz6666/3d-wm-atomic-v2.VIPS-v2xreal-assets
VIPS — V2X-Real Evaluation Assets
Generated assets for the VIPS benchmark — a NAVSIM-style two-stage PDM
(Predictive Driver Model) evaluation on the V2X-Real
dataset. Use them together with the code at
github.com/mickeykang16/VIPS.
These are generated assets only. The official V2X-Real raw sensor images
are not redistributed here — download them from UCLA (step 1 below).
Contents
Path
Description
meta/spd_infos_temporal_test.pkl
Temporal test-split… See the full description on the dataset page: https://huggingface.co/datasets/mickeykang/VIPS-v2xreal-assets.StreamGaze_v2
StreamGaze Dataset
StreamGaze is a comprehensive streaming video benchmark for evaluating MLLMs on gaze-based QA tasks across past, present, and future contexts.
Companion dataset: The EgoGazeVQA dataset is hosted separately at Peanuttoad/gaze_dataset.
📁 Dataset Structure
streamgaze/
├── metadata/
│ ├── egtea.csv # EGTEA fixation metadata
│ ├── egoexolearn.csv # EgoExoLearn fixation metadata
│ └── holoassist.csv # HoloAssist… See the full description on the dataset page: https://huggingface.co/datasets/Peanuttoad/StreamGaze_v2.Sleep-EDF-V2sd2hd_images_v2lfhre-images-v2rukopys-curated-mvp-v2
RUKOPYS Curated MVP: Ukrainian Handwriting Recognition Dataset
RUKOPYS Curated MVP is a cleaned, task-ready derivative of
UkrainianCatholicUniversity/rukopys for Ukrainian handwritten
document AI. It turns the raw RUKOPYS release into reproducible artifacts for page-level
vision-language fine-tuning, crop-level transcription, and layout detection.
This dataset is designed for practical HTR work: train a model, inspect the normalized records,
evaluate layout/text extraction, and… See the full description on the dataset page: https://huggingface.co/datasets/AlexandreSheva/rukopys-curated-mvp-v2.mike-no-hito
Mike No Hito
This dataset contains nearly all the images from artist Mike No Hito, huge thanks to him for the cute catgirls :3.
Cleaned, deduped, and tagged using 9001/copyparty, LagPixelLOL/mitgw, and SmilingWolf/wd-eva02-large-tagger-v3.
>:3 me when on my way to steal everything and throw them into gradient descent >:P
font-square-v2-pairs3d-wm-atomic-v2vtktext-image-finetune-1k-v2drifting-vla-v2-vu-robopoint
