datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Puffin-4M
Thinking with Camera: A Unified Multimodal Model for Camera-Centric Understanding and Generation
📖 Project Page | 🖥️ GitHub | 🤗 Hugging Face | 📑 Paper
Dataset Details
Datasets and benchmarks that span vision, language, and camera modalities remain scarce in the domain of spatial multimodal intelligence.
To address this gap, we introduce Puffin-4M, a large-scale, high-quality dataset comprising 4 million vision-language-camera… See the full description on the dataset page: https://huggingface.co/datasets/KangLiao/Puffin-4M.PubTables-v2
PubTables-v2
PubTables-v2 is a new large-scale dataset for full-page and multi-page table extraction.
Official dataset evaluation scripts and leaderboard coming soon!
In the meantime, you can create your own evaluation using GriTS with our open-source package, pip install grits-metric.
Report any issues here: https://github.com/kensho-technologies/grits.
See also: Hugging Face Paper Page
News
2026 Apr 15: Code for the GriTS metric released… See the full description on the dataset page: https://huggingface.co/datasets/kensho/PubTables-v2.imagenet-ckiwi_subVirtual_KITTI2VIVID-10M
VIVID-10M
[project page] | [Paper] | [arXiv]
VIVID-10M is the first large-scale hybrid image-video local editing dataset aimed at reducing data construction and model training costs, comprising 9.7M samples that encompass a wide range of video editing tasks.
Data Index
The data index is located at four .csv files:
vivid-image-change.csv
vivid-image-remove.csv
vivid-video-change.csv
vivid-video-remove.csv
VIVID-Video splits contains the columns:
local_caption, #… See the full description on the dataset page: https://huggingface.co/datasets/KlingTeam/VIVID-10M.ImageNet-Renditionllava-1.5-665k-instructionsThis dataset repository, LLaVA-1.5-665K-Instructions, is notably utilized in the paper Zero-Shot Vision Encoder Grafting via LLM Surrogates.
The official code repository for the paper can be found here: https://github.com/kaiyuyue/zero
LLaVA-1.5-665K-Instructions
This dataset repo contains the entire LLaVA-1.5-665K-Instructions in one place, including images and text sequences.
The images are in train_split/*.tars and the text sequences are in jsons:
llava_v1_5_mix665k.json is the… See the full description on the dataset page: https://huggingface.co/datasets/kaiyuyue/llava-1.5-665k-instructions.Objaverse_zero123_wdscoco-karpathy-wds
COCO-2014 WebDataset Format (Karpathy Splits)
This dataset contains the COCO-2014 images and captions converted to WebDataset (WDS) format, using the Karpathy & Li (2015) dataset split for image captioning tasks.
Overview
Total Samples: 123,287 images with 5 reference captions each
Total Size: ~19 GB
Format: WebDataset (.tar shards)
Shard Size: 1,000 samples per tar file
License: CC-BY 4.0
Language: English
Structure
COCO-2014-WDS/
├── train/ (113… See the full description on the dataset page: https://huggingface.co/datasets/undefined443/coco-karpathy-wds.ImageNet-V2Korean.OCR.Img.text.pairstylebench-sSIGGRAPH 2026 / ACM TOG Journal Track
imagenetstormer-40yrstrain: 1979~2018 val: 2019 test: 2020
HF_HUB_ENABLE_HF_TRANSFER=1 hf download —repo-type dataset —local-dir /workspace/stormer/ KyleBae1017/stormer-40yrs
cat wb2_h5df.tar.part-* | tar -xvf -
danbooru2023-webp-4Mpixel
Danbooru 2023 webp: A space-efficient version of Danbooru 2023
This dataset is a resized/re-encoded version of danbooru2023.
Which removed the non-image/truncated files and resize all of them into smaller size.
This dataset already be updated to latest_id = 7,832,883.
Thx to DeepGHS!
Notice: content of updates folder and deepghs/danbooru_newest-webp-4Mpixel have been merged to 2000~2999.tar, You can ignore all the content in updates folder safely!
Details
This… See the full description on the dataset page: https://huggingface.co/datasets/KBlueLeaf/danbooru2023-webp-4Mpixel.kittikitti-cThis dataset is created by MonoTTA: Fully Test-Time Adaptation for Monocular 3D Object Detection, based on KITTI.
You can check this link for more details: https://arxiv.org/abs/2405.19682v1
And access the code: https://github.com/Hongbin98/MonoTTA
Please double-check the demands of KITTI when you try to download this dataset and obey their rules.
cc12m-sam2-parse-treekitti-depth-completion
KITTI Depth Completion
This repository contains a tar.zst archive of the kitti_depth_completion directory, split into 5 GiB parts.
Archive parts: 4
Total size: 21,374,819,904 bytes
Source file count: 193,568
Restore
cat kitti_depth_completion.tar.zst.part-* | tar --zstd -xf -
sphere-encoder-fid-artifacts
Sphere Encoder FID Evaluation Artifacts
This repository contains the evaluation artifacts for the paper Image Generation with a Sphere Encoder.
Project Page | GitHub Repository
These artifacts include data statistic files (fid_stats) and reference images (fid_refs) used to calculate Fréchet Inception Distance (FID) for generative models across several datasets, including CIFAR-10, ImageNet, Animal Faces, and Oxford Flowers.
Workspace Setup
Download the evaluation… See the full description on the dataset page: https://huggingface.co/datasets/kaiyuyue/sphere-encoder-fid-artifacts.Describable-Textures-DatasetWikipedia-Knowledge-2M
📃 Paper | 🤗 Hugging Face | ⭐ Github
Dataset Overview
In the table below, we provide a brief summary of the dataset statistics.
Category
Size
Total Sample
2019163
Total Image
2019163
Average Answer Length
84
Maximum Answer Length
5851
JSON Overview
Each dictionary in the JSON file contains three keys: 'id', 'image', and 'conversations'.
The 'id' is the unique identifier for the current data in the entire dataset.
The 'image' stores… See the full description on the dataset page: https://huggingface.co/datasets/Ghaser/Wikipedia-Knowledge-2M.TWDS16all2024speedtest_1kitti-depth-gtzangei-dit-stage-1-250k-256px-imgall2024speedtest_0VLMBench_datasetall2024speedtest_7
