datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
RekaDaily-10k-raw
RekaDaily-10k (raw)
Raw, unscripted, first-person daily-life video, collected through
Claru, Reka's data collection marketplace — recorded by
paid collectors in their own homes and workplaces on head-mounted and handheld
phones, across multiple regions.
Videos are delivered as recorded — no cuts, no trimming, no editing, no
filtering beyond basic integrity checks. A processed tier (short clips with
machine captions) is released separately under the same RekaDaily-10k prefix.… See the full description on the dataset page: https://huggingface.co/datasets/RekaAI/RekaDaily-10k-raw.RekaDaily-10k-processed
RekaDaily-10k (processed)
Short first-person clips cut from the RekaDaily-10k
recordings —
unscripted daily-life video collected through Claru, Reka's
data collection marketplace, recorded by paid collectors in their own homes and
workplaces on head-mounted and handheld phones, across multiple regions.
Every clip carries one dense caption and a multi-question Q&A exchange
written in the second person ("What am I doing in this video?"), so the corpus
drops straight into… See the full description on the dataset page: https://huggingface.co/datasets/RekaAI/RekaDaily-10k-processed.amara-spatial-10k
AmaraSpatial-10K
A Semantically Anchored, Metric-Scale 3D Dataset for Embodied AI and Spatial Computing
10,071 AI-generated 3D meshes across 10 top-level categories and 476 subcategories — from basilisks to bassoons, cottages to cosmic stations — curated by Zero One Creative to close the spatial alignment gap that makes most generative 3D repositories unusable for zero-shot deployment in game engines, robotics simulators, and AR/VR pipelines.
Every asset is… See the full description on the dataset page: https://huggingface.co/datasets/ZeroOneCreative/amara-spatial-10k.OpenMath-Vision-CoT-10kAirGoal-10k
AirGoal-10k
AirGoal-10k is an aerial image-goal navigation dataset released with
UA-NWM: Uncertainty-Aware World Model for Aerial Image-Goal Navigation.
Project page: https://duryi.github.io/UA-NWM-Project-Page/Code: https://github.com/DurYi/UA-NWMPaper: https://arxiv.org/abs/2608.05597
Dataset Summary
AirGoal-10k contains 11,000 aerial navigation trajectories for image-goal navigation. Each trajectory
contains 12 RGB observations and trajectory metadata. The test… See the full description on the dataset page: https://huggingface.co/datasets/DurYi/AirGoal-10k.Critic-10K
Critic-10K Dataset
This repository hosts the Critic-10K dataset, introduced in the paper The Consistency Critic: Correcting Inconsistencies in Generated Images via Reference-Guided Attentive Alignment.
The Critic-10K dataset is specifically constructed to address and rectify inconsistencies in generated images. It comprises reference-degraded-target triplets, obtained through VLM-based selection and explicit degradation. This dataset effectively simulates common inaccuracies or… See the full description on the dataset page: https://huggingface.co/datasets/ziheng1234/Critic-10K.Egocentric_10K_Evaluation
Dataset Card for Egocentric_10K_Evaluation
This is a FiftyOne dataset with 30000 samples.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import fiftyone as fo
from fiftyone.utils.huggingface import load_from_hub
# Load the dataset
# Note: other available arguments include 'max_samples', etc
dataset = load_from_hub("Voxel51/Egocentric_10K_Evaluation")
# Launch the App
session = fo.launch_app(dataset)
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/Egocentric_10K_Evaluation.IntegraCAR-LULC-10K
IntegraCAR-LULC-10K: A High-Resolution Optical Satellite Dataset for LULC Segmentation in the Brazilian Rural Environmental Registry
[!WARNING]
⚠️ High Volume & Storage Advisory (+300 GB)
This repository hosts the complete 10,000-tile collection (IntegraCAR-LULC-10K), consisting of over 300 GB of high-resolution satellite imagery ( 2048×20482048 \times 20482048×2048 px at 0.5 m/px0.5\text{ m/px}0.5 m/px ) and pixel-level segmentation masks stored in… See the full description on the dataset page: https://huggingface.co/datasets/laicsiifes/IntegraCAR-LULC-10K.ai2thor-perspective-qa-10kpred_llava_next_10ksplash-art-gacha-collection-10k
Splash Art Collection 10K
This collection features 11,755 character splash arts or 角色立绘 sourced from 47 gacha games, meticulously gathered from Fandom and Biligame WIKI.
The dataset is suitable for fine-tuning T2I models on splash art generation domain, utilizing the image and prompt fields. It includes a mix of both high- and low-quality splash arts of various styles, allowing you to curate and select the images that best suit your training needs.
Data Structure… See the full description on the dataset page: https://huggingface.co/datasets/mrzjy/splash-art-gacha-collection-10k.TripVVT-10K
TripVVT-10K Dataset
News
2026.06: TripVVT has been accepted by ECCV 2026.
2026.04: The TripVVT paper is available on arXiv.
The project page is available at https://shaodingbao.github.io/TripVVT/.
TripVVT-10K is a large-scale dataset for in-the-wild Video Virtual Try-On (VVT). It contains 10,031 high-quality video samples with triplet supervision, covering upper-body garments, lower-body garments, and dresses.
TripVVT-10K is released together with the… See the full description on the dataset page: https://huggingface.co/datasets/TripVVT/TripVVT-10K.pseudo-camera-10k
pseudo-camera-10k dataset
Contents
This dataset contains 10k free images from world class photographers. The images have been resized using Lanczos antialiasing, with their smaller edge shifted to 1024px.
The aim of this dataset is a highly variable but high quality and high resolution set of images containing difficult concepts, with about half of the images being numbered group shots and family portraits with the number of subjects labeled.
No images were upsampled in… See the full description on the dataset page: https://huggingface.co/datasets/bghira/pseudo-camera-10k.S1-Omni-Corpus-10K
S1-Omni-Corpus-10K
An open-source scientific multimodal reasoning dataset subset for S1-Omni
🧬 Model Introduction
S1-Omni is a unified scientific multimodal reasoning model for scientific understanding, prediction, and generation. It is developed by the ScienceOne AI team of the Chinese Academy of Sciences.
S1-Omni addresses fragmented scientific AI capabilities with a shared backbone for cross-disciplinary, cross-modal, and cross-task understanding and reasoning… See the full description on the dataset page: https://huggingface.co/datasets/ScienceOne-AI/S1-Omni-Corpus-10K.MaskEdit-10k
🌊 MaskEdit-10k: Dataset for paper "MaskFlow: Precise, Consistent and Seamless Regional Image Editing"
MaskEdit-10k is a mask-guided regional image-editing dataset containing 11,213 source-mask-target triplets. Each sample pairs a source image and a spatial mask with a target image and two complementary editing prompts.
Dataset structure
The dataset provides three configurations and two fixed splits per configuration.
Configuration… See the full description on the dataset page: https://huggingface.co/datasets/ReyChiaro/MaskEdit-10k.DigiCam-Mirflickr-MultiMask-10KArtiMuse-10K
ArtiMuse:
Fine-Grained Image Aesthetics Assessment with Joint Scoring and Expert-Level Understanding
[🌐 Project Page]
[🚀 Online Demo]
[💻 Code]
[📄 Paper]
[[🧩 Checkpoints: 🤗 Hugging Face | 🤖 ModelScope]]
🌟 Building upon on ArtiMuse, we introduce UniPercept, a comprehensive follow-up work that provides a meticulous study on perceptual-level image understanding. It spans Image Aesthetics Assessment (IAA), Image Quality Assessment (IQA), and Image Structure & Texture… See the full description on the dataset page: https://huggingface.co/datasets/Thunderbolt215215/ArtiMuse-10K.SciGenEdit-10K
SciGenEdit-10K
An Open Dataset for Scientific Image Generation and Editing
English | 简体中文
📖 Introduction
SciGenEdit-10K is a public subset released with the S1-Omni-Image project. It is designed for research on scientific image generation, scientific image editing, and multi-turn scientific image generation and editing.
S1-Omni-Image is a unified multimodal model developed by the ScienceOne team at the Chinese Academy of Sciences for scientific… See the full description on the dataset page: https://huggingface.co/datasets/ScienceOne-AI/SciGenEdit-10K.Argimi-Ardian-Finance-10k-text-image
The ArGiMI Ardian datasets : text and images
The ArGiMi project is committed to open-source principles and data sharing.
Thanks to our generous partners, we are releasing several valuable datasets to the public.
Dataset description
This dataset comprises 34,000 financial annual reports, written in English, meticulously
extracted from their original PDF format to provide a valuable resource for researchers and developers in financial
analysis and natural language… See the full description on the dataset page: https://huggingface.co/datasets/artefactory/Argimi-Ardian-Finance-10k-text-image.LogoBrief-10K-assets
LogoBrief-10K — preview assets
Public, ungated image assets for the LogoBrief-10K dataset card.
This repo exists only because the main dataset is gated, and Hugging Face's access gating applies to every file in a gated repo — including images referenced by the dataset card itself. Without this split, the card's own preview images (sample logos, methodology charts) would be invisible to anyone who hasn't already been granted access, defeating their purpose. These files are… See the full description on the dataset page: https://huggingface.co/datasets/Logolabs/LogoBrief-10K-assets.pred_qwen3vl_ins_10k_originai2thor-perspective-qa-10k-object-category-namesPersona-Fluxed-10k-2608
Persona Fluxed 10k
Synthetic persona portraits rendered with FLUX.2-klein-4b (8-step, 1024x1024) from the NVIDIA Nemotron-Personas-* datasets.
Each persona is grounded in real-world demographic, geographic and
personality-trait distributions for its country (CC BY 4.0 source; no real people).
Currently Nemotron-Personas exist for:
USA — English
Japan — Japanese
India — English, Hindi
Brazil — Portuguese
Singapore — English
France — French
Korea — Korean
El Salvador — Spanish… See the full description on the dataset page: https://huggingface.co/datasets/retowyss/Persona-Fluxed-10k-2608.P3M-10K
P3M-10K
P3M-10K (Privacy-Preserving Portrait Matting) is a large-scale portrait matting
benchmark. It is redistributed here from the original release by
JizhiziLi/P3M.
If you use this dataset, please cite the original paper:
Jizhizi Li, Sihan Ma, Xin Zhang, Dacheng Tao.
"Privacy-Preserving Portrait Matting." ACM International Conference on
Multimedia (ACM MM), 2021.
Contents
Each example is a portrait RGB image and its corresponding alpha matte:
Column… See the full description on the dataset page: https://huggingface.co/datasets/nobg/P3M-10K.Egocentric-10K-Evaluation
Egocentric10K
Ego4D
Epic-Kitchens
To evaluate the three in-the-wild egocentric datasets Egocentric-10K, Ego4D, and EPIC-KITCHENS-100 on hand visibility and active manipulation density as a proxy for data efficiency, we randomly sample 10k frames from each dataset and run them through a gemini-2.5-flash.
Hand Visibility
Prompt:
You are labeling an egocentric first-person image. Your task is to count… See the full description on the dataset page: https://huggingface.co/datasets/builddotai/Egocentric-10K-Evaluation.TAP360-10kai2thor-perspective-qa-10k-fix-x-markVisionFoundry-10K
VisionFoundry-10K
VisionFoundry-10K is a synthetic visual question answering (VQA) dataset with 10,000 image-question-answer triples spanning 10 vision-centric tasks. The data is produced by the VisionFoundry pipeline: an LLM generates task-aware questions, answers, and detailed text-to-image prompts; a text-to-image model synthesizes images; and a strong multimodal verifier filters samples for alignment.
VisionFoundry: Teaching VLMs Visual Perception with… See the full description on the dataset page: https://huggingface.co/datasets/zlab-princeton/VisionFoundry-10K.subject_dataset_10kaxsy_ps-inspection_10k
