datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
CALVIN_ABC_tarPuffin-4M
Thinking with Camera: A Unified Multimodal Model for Camera-Centric Understanding and Generation
📖 Project Page | 🖥️ GitHub | 🤗 Hugging Face | 📑 Paper
Dataset Details
Datasets and benchmarks that span vision, language, and camera modalities remain scarce in the domain of spatial multimodal intelligence.
To address this gap, we introduce Puffin-4M, a large-scale, high-quality dataset comprising 4 million vision-language-camera… See the full description on the dataset page: https://huggingface.co/datasets/KangLiao/Puffin-4M.mls_hq_urgent_track1PubTables-v2
PubTables-v2
PubTables-v2 is a new large-scale dataset for full-page and multi-page table extraction.
Official dataset evaluation scripts and leaderboard coming soon!
In the meantime, you can create your own evaluation using GriTS with our open-source package, pip install grits-metric.
Report any issues here: https://github.com/kensho-technologies/grits.
See also: Hugging Face Paper Page
News
2026 Apr 15: Code for the GriTS metric released… See the full description on the dataset page: https://huggingface.co/datasets/kensho/PubTables-v2.emo_webds_2emo_parleremo_webdsnpm3d-kitti-carlaimagenet-cDeepJEB
DeepJEB: 3D Deep Learning-Based Synthetic Jet Engine Bracket Dataset
This is the Hugging Face distribution of DeepJEB, a synthetic 3D jet engine
bracket dataset of 2,138 designs with paired geometry and finite-element
analysis (FEA) results, generated from the SimJEB seed set via a DeepSDF-based
generative model and an automated simulation pipeline.
This repository mirrors the official DeepJEB v1.0 release. To keep the ~68k
component files practical to host, each component… See the full description on the dataset page: https://huggingface.co/datasets/KAIST-SmartDesignLab/DeepJEB.kiwi_subVirtual_KITTI2VIVID-10M
VIVID-10M
[project page] | [Paper] | [arXiv]
VIVID-10M is the first large-scale hybrid image-video local editing dataset aimed at reducing data construction and model training costs, comprising 9.7M samples that encompass a wide range of video editing tasks.
Data Index
The data index is located at four .csv files:
vivid-image-change.csv
vivid-image-remove.csv
vivid-video-change.csv
vivid-video-remove.csv
VIVID-Video splits contains the columns:
local_caption, #… See the full description on the dataset page: https://huggingface.co/datasets/KlingTeam/VIVID-10M.ImageNet-Renditionstereo4d-lefteye-perspective
Dataset Summary
This dataset contains the left-eye rectified perspective views from the Stereo4D dataset (Paper). Each video is generated using the rectify.py script, which processes VR180 stereo videos to produce 512×512 video clips with a 60° field of view perspective camera. This dataset is intended to be used alongside the Stereo4D dataset annotations which can be found here.
This dataset is provided as-is for non-commercial research purposes only.
Download
git clone… See the full description on the dataset page: https://huggingface.co/datasets/KevinMathew/stereo4d-lefteye-perspective.llava-1.5-665k-instructionsThis dataset repository, LLaVA-1.5-665K-Instructions, is notably utilized in the paper Zero-Shot Vision Encoder Grafting via LLM Surrogates.
The official code repository for the paper can be found here: https://github.com/kaiyuyue/zero
LLaVA-1.5-665K-Instructions
This dataset repo contains the entire LLaVA-1.5-665K-Instructions in one place, including images and text sequences.
The images are in train_split/*.tars and the text sequences are in jsons:
llava_v1_5_mix665k.json is the… See the full description on the dataset page: https://huggingface.co/datasets/kaiyuyue/llava-1.5-665k-instructions.Objaverse_zero123_wdsOpenDialog
OpenDialog
OpenDialog is a 6.8k hours spoken dialogue dataset, introduced in the paper ZipVoice-Dialog: Non-Autoregressive Spoken Dialogue Generation with Flow Matching.
Paper: https://arxiv.org/abs/2507.09318
GitHub: https://github.com/k2-fsa/ZipVoice
Project Page: https://zipvoice-dialog.github.io
OpenDialog is the first large-scale (6.8k hours) open-source spoken dialogue dataset derived from in-the-wild speech data. It consists of:
English data: 5074 hours
Chinese data: 1759… See the full description on the dataset page: https://huggingface.co/datasets/k2-fsa/OpenDialog.AllTheBacteria-FCGR-7merVideoTemp-o3 VideoTemp-o3: Harmonizing Temporal Grounding and Video Understanding in Agentic Thinking-with-Videos
Illustration of the agentic pipeline in VideoTemp-o3. Given a video QA pair, the model performs on-demand temporal grounding to locate the most relevant segment, then refines it iteratively. Finally, it produces a reliable answer grounded in the pertinent visual evidence.
Data Source
The question and answer pairs used for training VideoTemp-o3 are sourced from… See the full description on the dataset page: https://huggingface.co/datasets/Kwai-Keye/VideoTemp-o3.laion-coco-13m-tarcoco-karpathy-wds
COCO-2014 WebDataset Format (Karpathy Splits)
This dataset contains the COCO-2014 images and captions converted to WebDataset (WDS) format, using the Karpathy & Li (2015) dataset split for image captioning tasks.
Overview
Total Samples: 123,287 images with 5 reference captions each
Total Size: ~19 GB
Format: WebDataset (.tar shards)
Shard Size: 1,000 samples per tar file
License: CC-BY 4.0
Language: English
Structure
COCO-2014-WDS/
├── train/ (113… See the full description on the dataset page: https://huggingface.co/datasets/undefined443/coco-karpathy-wds.ImageNet-V2Korean.OCR.Img.text.pairmls_hqstylebench-sSIGGRAPH 2026 / ACM TOG Journal Track
CausalDynamics
CausalDynamics: A large-scale benchmark for structural discovery of dynamical causal models
NeurIPS 2025
A comprehensive benchmark framework designed to rigorously evaluate state-of-the-art causal discovery algorithms for dynamical systems.
Key Features
1️⃣ Large-Scale Benchmark. Systematically evaluate state-of-the-art causal discovery algorithms on thousands of graph challenges with increasing difficulty.
2️⃣ Customizable Data Generation. Scalable… See the full description on the dataset page: https://huggingface.co/datasets/kausable/CausalDynamics.imagenetemo_speech_filtered_v12 second filtered emotional speech in webdataset format
https://huggingface.co/datasets/EQ4You/Emotional_Speech
stormer-40yrstrain: 1979~2018 val: 2019 test: 2020
HF_HUB_ENABLE_HF_TRANSFER=1 hf download —repo-type dataset —local-dir /workspace/stormer/ KyleBae1017/stormer-40yrs
cat wb2_h5df.tar.part-* | tar -xvf -
