datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
FastUMI_100k_lerobot
FastUMI-100K: Advancing Data-Driven Robotic Manipulation with a Large-Scale UMI-Style Dataset
[paper] [dataset]
## Overview
FastUMI-100K is a large-scale, high-quality UMI-style dataset designed for data-driven robotic manipulation learning. Featuring over **100K+ demonstration trajectories** across **54 diverse tasks** and hundreds of object types, the dataset provides multi-view wrist-mounted fisheye images and high-frequency end-effector states. To… See the full description on the dataset page: https://huggingface.co/datasets/IPEC-COMMUNITY/FastUMI_100k_lerobot.Egocentric-100K
Egocentric-100K is the largest dataset of manual labor. You can visualize the dataset here.
Egocentric-100K is state-of-the-art in hand visibility and active manipulation density compared to previous in-the-wild egocentric datasets. The complete 30,000 frame evaluation set is available at Egocentric-100K-Evaluation.
Dataset Statistics
Attribute
Value
Total Hours
100,405
Total Frames
10.8 billion
Video Clips
2,010,759
Median Clip Length
180.0 seconds
Mean… See the full description on the dataset page: https://huggingface.co/datasets/builddotai/Egocentric-100K.FineTome-100k
FineTome-100k
The FineTome dataset is a subset of arcee-ai/The-Tome (without arcee-ai/qwen2-72b-magpie-en), re-filtered using HuggingFaceFW/fineweb-edu-classifier.
It was made for my article "Fine-tune Llama 3.1 Ultra-Efficiently with Unsloth".
TimeLens-100K
TimeLens-100K
📑 Paper | 💻 Code | 🏠 Project Page | 🤗 Model & Data
✨ Dataset Description
TimeLens-100K is a large-scale, diverse, and high-quality training dataset for video temporal grounding. It was proposed in our paper TimeLens: Rethinking Video Temporal Grounding with Multimodal LLMs and used for training TimeLens models. The annotation process was conducted using an automated pipeline powered by Gemini-2.5-Pro.
📊 Dataset Statistics
Total Videos:… See the full description on the dataset page: https://huggingface.co/datasets/TencentARC/TimeLens-100K.Complete_Data_Source_100K_HOURS
Multi-Language Audio Collection (100K Hours)
This repository is physically reorganized for Absolute 100% Data Visibility.
🏗️ Global Consolidator
Select your language subset to listen to high-quality waveform audio. All shards from legacy and modern pipelines are automatically routed here.
arxiv-ai-ml-100k-papers
license: other
tags:
- arxiv
- ocr
- machine-learning
---
# obswork/arxiv-ai-ml-100k
A 99,999-paper stratified subset of
[`Rendra8631/arxiv-papers`](https://huggingface.co/datasets/Rendra8631/arxiv-papers)
at revision `a2c6afb51332d2744b46308df6917697582f8cd4`, filtered to the primary
subjects `cs.AI`, `cs.CV`, `cs.LG`, and `stat.ML`. Only papers submitted in `2023`-`2025` are included.
This dataset is a build artifact of the OCR… See the full description on the dataset page: https://huggingface.co/datasets/obswork/arxiv-ai-ml-100k-papers.ChatDoctor-HealthCareMagic-100k
Dataset Card for "ChatDoctor-HealthCareMagic-100k"
More Information needed
arxiv-ai-ml-100k-pages
license: other
tags:
- arxiv
- ocr
- machine-learning
---
# obswork/arxiv-ai-ml-100k-pages
A **page-bounded** stratified subset of the raw pool dataset
[`obswork/arxiv-ai-ml-100k`](https://huggingface.co/datasets/obswork/arxiv-ai-ml-100k),
filtered to primary subjects `cs.AI`, `cs.CV`, `cs.LG`, and `stat.ML`.
The raw pool is itself a 100k-paper stratified sample from… See the full description on the dataset page: https://huggingface.co/datasets/obswork/arxiv-ai-ml-100k-pages.Q-Eval-100K
Q-Eval-100K Dataset (CVPR 2025 Oral)
📝 Introduction
The Q-Eval-100K dataset encompasses both text-to-image and text-to-video models, with 960K human annotations specifically focused on visual quality and alignment for 100K instances (60K images and 40K videos).
We utilize multiple popular text-to- image and text-to-video models to ensure diversity, which include FLUX, Lumina-T2X, PixArt, Stable Diffusion 3, Stable Diffusion XL, DALL·E 3, Wanx, Midjourney, Hunyuan-DiT… See the full description on the dataset page: https://huggingface.co/datasets/meituan-longcat/Q-Eval-100K.Zero-To-CAD-100k
Zero-to-CAD 100K
A curated subset of 100,000 geometrically diverse CAD construction sequences selected from Zero-to-CAD 1M.
Zero-to-CAD: Agentic Synthesis of Interpretable CAD Programs at Million-Scale Without Real Data
Mohammadmehdi Ataei, Farzaneh Askari, Kamal Rahimi Malekshan, Pradeep Kumar Jayaraman
Autodesk Research
Related Resources
Resource
Link
📄 Paper
Zero-to-CAD: Agentic Synthesis of Interpretable CAD Programs at Million-Scale Without… See the full description on the dataset page: https://huggingface.co/datasets/ADSKAILab/Zero-To-CAD-100k.UltraHR-100K
UltraHR-100K: Enhancing UHR Image Synthesis with A Large-Scale High-Quality Dataset
UltraHR-100K
Ultra-high-resolution (UHR) text-to-image (T2I) generation has seen notable progress. However, two key challenges remain : 1) the absence of a large-scale high-quality UHR T2I dataset, and (2) the neglect of tailored training strategies for fine-grained detail synthesis in UHR scenarios. To tackle the first challenge, we introduce… See the full description on the dataset page: https://huggingface.co/datasets/zhihefang/UltraHR-100K.ai-images-100ksweden_100K_difficultMLIC-Train-100K
Abstract
The latent representation in learned image compression encompasses channel-wise, local spatial, and global spatial correlations, which are essential for the entropy model to capture for conditional entropy minimization. Efficiently capturing these contexts within a single entropy model, especially in high-resolution image coding, presents a challenge due to the computational complexity of existing global context modules. To address this challenge, we propose the Linear… See the full description on the dataset page: https://huggingface.co/datasets/Whiteboat/MLIC-Train-100K.ManiTwin-100K
ManiTwin-100K: Manipulation-Ready Digital Object Twins
Project Page |
Paper
ManiTwin-100K is a large-scale dataset of manipulation-ready digital object twins designed for robotic manipulation research. Each object includes simulation-ready 3D meshes, physical properties, functional point annotations, grasp configurations, and rich language descriptions—all validated through physics-based simulation.
Note: We are currently releasing approximately 1K sample objects with a… See the full description on the dataset page: https://huggingface.co/datasets/ManiTwin/ManiTwin-100K.RevealLayer-100K
RevealLayer Open Dataset
RevealLayer Open is the open-source dataset accompanying RevealLayer: Disentangling Hidden and Visible Layers via Occlusion-Aware Image Decomposition.
Paper: https://arxiv.org/html/2605.11818v1 Accepted by ICML 2026
RevealLayer studies box-guided layered image decomposition for natural images. Given an RGB image and instance bounding boxes, the task is to decompose the scene into a clean background and object-level foreground layers, where each… See the full description on the dataset page: https://huggingface.co/datasets/qihoo360/RevealLayer-100K.ManiTwin-100K
ManiTwin-100K: Manipulation-Ready Digital Object Twins
Project Page |
Paper
ManiTwin-100K is a large-scale dataset of manipulation-ready digital object twins designed for robotic manipulation research. Each object includes simulation-ready 3D meshes, physical properties, functional point annotations, grasp configurations, and rich language descriptions—all validated through physics-based simulation.
Note: We are currently releasing approximately 1K sample objects with a… See the full description on the dataset page: https://huggingface.co/datasets/Dhanush944/ManiTwin-100K.JAVEdit-100k
Yinan Chen 1★ · Chuming Lin 2★ · Zhennan Chen 3 · Yuxiang Zeng 4 · Junwei Zhu 2 ·
Yali Bi 1 · Xijie Huang 5 · Chengming Xu 2 · Donghao Luo 2 · Zhucun Xue 1 ·
Xiaobin Hu 6 · Chengjie Wang 2 · Yong Liu 1 · Jiangning Zhang 1,2 📧 · Shuicheng Yan 6
1 Zhejiang University 2 YouTu Lab, Tencent 3 Nanjing University
4 University of Auckland 5 Fudan University 6 National University of Singapore
😊 Dataset Introduction
JAVEdit-100k is the official dataset… See the full description on the dataset page: https://huggingface.co/datasets/Coraxor/JAVEdit-100k.LLaVA-CoT-100k
Dataset Card for LLaVA-CoT
The LLaVA-CoT-100k dataset is introduced in the paper LLaVA-CoT: Let Vision Language Models Reason Step-by-Step. This dataset is designed to enable Vision-Language Models (VLMs) to perform autonomous multistage reasoning, integrating samples from various visual question-answering sources with structured reasoning annotations. It aims to address the challenges VLMs face in systematic and structured reasoning for complex visual question-answering tasks.… See the full description on the dataset page: https://huggingface.co/datasets/Xkev/LLaVA-CoT-100k.political-leaning-tweets-100k
🗳️ political-leaning-tweets-100k
Châtelet AI presents a 100,000+ dataset of tweets labelled for political leaning: neutral, liberal, conservative.Labels are machine-generated using a SOTA thinking-enabled LLM. The dataset is intended for research on political language modelling, ideology detection, robustness, and safety evaluation.
📦 Dataset Card
Name: chatelet/political-leaning-tweets-100k
Publisher: Châtelet AI
Licence: MIT with additional restrctions against… See the full description on the dataset page: https://huggingface.co/datasets/chatelet/political-leaning-tweets-100k.cosmopedia-100k
Dataset description
This is a 100k subset of Cosmopedia dataset. A synthetic dataset of textbooks, blogposts, stories, posts and WikiHow articles generated by Mixtral-8x7B-Instruct-v0.1.
Here's how you can load the dataset
from datasets import load_dataset
ds = load_dataset("HuggingFaceTB/cosmopedia-100k", split="train")
libgen-10k-100kVisualGenome_VG_100K_1_and_2
openwebtext-100k
Dataset Card for "openwebtext-100k"
More Information needed
Q-Eval-100K
Q-Eval-100K Dataset (CVPR 2025 Oral)
📝 Introduction
The Q-Eval-100K dataset encompasses both text-to-image and text-to-video models, with 960K human annotations specifically focused on visual quality and alignment for 100K instances (60K images and 40K videos).
We utilize multiple popular text-to- image and text-to-video models to ensure diversity, which include FLUX, Lumina-T2X, PixArt, Stable Diffusion 3, Stable Diffusion XL, DALL·E 3, Wanx, Midjourney, Hunyuan-DiT… See the full description on the dataset page: https://huggingface.co/datasets/kucklily/Q-Eval-100K.vipe-dynpose-100kpp
ViPE Dataset Release
This dataset contains the camera pose, depth, and intrinsics estimated using ViPE. For more details of the dataset, please refer to the Github link.
Please consider citing the following paper if you found this dataset helpful:
@article{huang2025vipe,
title={Vipe: Video pose engine for 3d geometric perception},
author={Huang, Jiahui and Zhou, Qunjie and Rabeti, Hesam and Korovko, Aleksandr and Ling, Huan and Ren, Xuanchi and Shen, Tianchang and Gao, Jun and… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/vipe-dynpose-100kpp.PICA-100K
PICABench: How Far Are We from Physically Realistic Image Editing?
Paper | Project Page | Code
Overview
PICABench probes how far current editing models are from physically realistic image manipulation. It ties together:
PICABench benchmark – physics-aware editing cases spanning eight laws across Optics, Mechanics, and State Transition, each labeled with superficial/intermediate/explicit difficulty tiers.
PICAEval metric – region-grounded, QA-based verification with… See the full description on the dataset page: https://huggingface.co/datasets/Andrew613/PICA-100K.fitvto-100k
FIT: A Large-Scale Dataset for Fit-Aware Virtual Try-On
The official preview dataset from the paper "FIT: A Large-Scale Dataset for Fit-Aware Virtual Try-On".
This dataset supports garment-centric virtual try-on and try-off research, containing 100,000 training and 5,000 evaluation triplets. Each sample pairs a person image with a layflat garment image and body/garment measurements.
Dataset Structure
Each split (train / eval) contains four aligned modalities — all… See the full description on the dataset page: https://huggingface.co/datasets/Yuanhao-Harry-Wang/fitvto-100k.vids-100karena-human-preference-100k
Overview
This dataset contains leaderboard conversation data collected between June 2024 and August 2024.
It includes English human preference evaluations used to develop Arena Explorer.
Additionally, we provide an embedding file, which contains precomputed embeddings for the English conversations.
These embeddings are used in the topic modeling pipeline to categorize and analyze these conversations.
For a detailed exploration of the dataset and analysis methods, refer to the… See the full description on the dataset page: https://huggingface.co/datasets/lmarena-ai/arena-human-preference-100k.
