datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
10Kh-RealOmin-OpenData
Boasting over 13,000 hours of cumulative data and 5 million+ clips, it ranks as the largest open-source embodied intelligence dataset in the industry.
Update Notes:Stage 3 data upload completed.
13,000+ hours of pure dual-hand data with frame-level alignment latency < 1ms
Full high-precision trajectory reconstruction, breaking the limit of superficial open source, fully ready-to-use
3,000+ contributors and 10,000+ real household scenarios with exceptional diversity… See the full description on the dataset page: https://huggingface.co/datasets/genrobot2025/10Kh-RealOmin-OpenData.SAGE-10k
SAGE-10k
SAGE-10k is a large-scale interactive indoor scene dataset featuring realistic layouts, generated by the agentic-driven pipeline introduced in "SAGE: Scalable Agentic 3D Scene Generation for Embodied AI". The dataset contains 10,000 diverse scenes spanning 50 room types and styles, along with 565K uniquely generated 3D objects.
🔑 Key Features
SAGE-10k integrates a wide variety of scenes, and particularly, preserves small items… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/SAGE-10k.RekaDaily-10k-raw
RekaDaily-10k (raw)
Raw, unscripted, first-person daily-life video, collected through
Claru, Reka's data collection marketplace — recorded by
paid collectors in their own homes and workplaces on head-mounted and handheld
phones, across multiple regions.
Videos are delivered as recorded — no cuts, no trimming, no editing, no
filtering beyond basic integrity checks. A processed tier (short clips with
machine captions) is released separately under the same RekaDaily-10k prefix.… See the full description on the dataset page: https://huggingface.co/datasets/RekaAI/RekaDaily-10k-raw.Egocentric-10K
Egocentric-10K is the largest egocentric dataset. It is the first dataset collected exclusively in real factories.
Your browser does not support the video tag.
Egocentric-10K is state-of-the-art in hand visibility and active manipulation density compared to previous in-the-wild egocentric datasets. The complete 30,000 frame evaluation set is available at Egocentric-10K-Evaluation.
Dataset Statistics
Attribute
Value
Total Hours
10,000
Total Frames
1.08 billion… See the full description on the dataset page: https://huggingface.co/datasets/builddotai/Egocentric-10K.pile-10kThe first 10K elements of The Pile, useful for debugging models trained on it. See the HuggingFace page for the full Pile for more info. Inspired by stas' great resource doing the same for OpenWebText
IG-10K-Dataset
The Imitator Game - IG-10K Dataset
This dataset accompanies the paper The Imitator Game: Benchmarking Robot Imitative Ability Beyond Action Prediction.
It contains paired human-robot demonstrations for the Imitator Game benchmark, spanning four levels of imitation difficulty (L0–L3) across real and simulated settings. The IG-10K dataset includes over 20,000 paired episodes across 50+ tasks and 6 domains, and is provided in LeRobot-0.5.0 format.
For more details, see the project… See the full description on the dataset page: https://huggingface.co/datasets/imitator-game/IG-10K-Dataset.dsir-pile-10kDL3DV-10K-Meshed
DL3DV-10K-Meshed
A derivative of DL3DV-10K providing undistorted 480P views
together with ground-truth surface geometry, used to train
Surflo.
This dataset is not a replacement for DL3DV-10K. It redistributes some of
DL3DV-10K imagery, and remains subject to the DL3DV-10K Terms of Use.
See Licensing and terms before using it.
What we changed
Relative to DL3DV-ALL-480P:
Undistorted images. All 480P frames were reprocessed with COLMAP to remove lens
distortion.… See the full description on the dataset page: https://huggingface.co/datasets/AntoineGuedon/DL3DV-10K-Meshed.RekaDaily-10k-processed
RekaDaily-10k (processed)
Short first-person clips cut from the RekaDaily-10k
recordings —
unscripted daily-life video collected through Claru, Reka's
data collection marketplace, recorded by paid collectors in their own homes and
workplaces on head-mounted and handheld phones, across multiple regions.
Every clip carries one dense caption and a multi-question Q&A exchange
written in the second person ("What am I doing in this video?"), so the corpus
drops straight into… See the full description on the dataset page: https://huggingface.co/datasets/RekaAI/RekaDaily-10k-processed.10Kh-RealOmin-OpenDataBoasting over 10,000 hours of cumulative data and 1 million+ clips, it ranks as the largest open-source embodied intelligence dataset in the industry.
Compared with other datasets, it has the following advantages:
Ample Data Volume & Strong Generalization
Each skill is supported by sufficient data, collected from over 3,000 households and nearly 10,000 distinct fine-grained targets. It avoids simple repetitions and ensures robust generalization.
Authentic Scenarios & Focused… See the full description on the dataset page: https://huggingface.co/datasets/ad1t7a/10Kh-RealOmin-OpenData.amara-spatial-10k
AmaraSpatial-10K
A Semantically Anchored, Metric-Scale 3D Dataset for Embodied AI and Spatial Computing
10,071 AI-generated 3D meshes across 10 top-level categories and 476 subcategories — from basilisks to bassoons, cottages to cosmic stations — curated by Zero One Creative to close the spatial alignment gap that makes most generative 3D repositories unusable for zero-shot deployment in game engines, robotics simulators, and AR/VR pipelines.
Every asset is… See the full description on the dataset page: https://huggingface.co/datasets/ZeroOneCreative/amara-spatial-10k.Argimi-Ardian-Finance-10k-text
The ArGiMI Ardian datasets : Text-only version
The ArGiMi project is committed to open-source principles and data sharing.
Thanks to our generous partners, we are releasing several valuable datasets to the public.
Dataset description
This text-only dataset comprises 34,000 financial annual reports, written in English, meticulously
extracted from their original PDF format to provide a valuable resource for researchers and developers in financial
analysis and natural… See the full description on the dataset page: https://huggingface.co/datasets/artefactory/Argimi-Ardian-Finance-10k-text.database-10k-val
database-10k-val
最终 40K 数据集的 10,000 样本验证子集(2026-09-17)。
选择约束
不含 inset:所有样本 layout_type != inset。
不含冻结/删除族:样本主 panel 与全部副 panel 的 chart_family 均不属于 area、matrix、set_relation。
从最终修复后的 40K 中按 (family, subtype, layout, difficulty) 分层、确定性抽取 10,000 条。
文件
figure2data_10k_val.sqlite3:子集 SQLite,10,000 samples / 70,000 documents。
images/:10,000 PNG。
documents/:10,000 JSON。
arrays/:有原始数组的样本 NPZ。
plans/generation_plan_10k_val.jsonl:10,000 行子集计划。… See the full description on the dataset page: https://huggingface.co/datasets/ZZoutian/database-10k-val.sp500-edgar-10k
Dataset Card for SP500-EDGAR-10K
Dataset Summary
This dataset contains the annual reports for all SP500 historical constituents from 2010-2022 from SEC EDGAR Form 10-K filings.
It also contains n-day future returns of each firm's stock price from each filing date.
Dataset Structure
Data Fields
[More Information Needed]
Data Splits
[More Information Needed]
Dataset Creation
Source Data
Initial Data Collection… See the full description on the dataset page: https://huggingface.co/datasets/jlohding/sp500-edgar-10k.STARK_10k
STARK: Spatial-Temporal reAsoning benchmaRK
STARK is a comprehensive benchmark designed to systematically evaluate large language models (LLMs) and large reasoning models (LRMs) on spatial-temporal reasoning tasks, particularly for applications in cyber-physical systems (CPS) such as robotics, autonomous vehicles, and smart city infrastructure.
Dataset Summary
Hierarchical Benchmark: Tasks are structured across three levels of reasoning complexity:
State Estimation:… See the full description on the dataset page: https://huggingface.co/datasets/prquan/STARK_10k.LongAlign-10k
LongAlign-10k
🤗 [LongAlign Dataset] • 💻 [Github Repo] • 📃 [LongAlign Paper]
LongAlign is the first full recipe for LLM alignment on long context. We propose the LongAlign-10k dataset, containing 10,000 long instruction data of 8k-64k in length. We investigate on trianing strategies, namely packing (with loss weighting) and sorted batching, which are all implemented in our code. For real-world long context evaluation, we introduce LongBench-Chat that evaluate the… See the full description on the dataset page: https://huggingface.co/datasets/zai-org/LongAlign-10k.c4-10k-mini-tokenized-16-ctx-gelu-1l-testsfineweb-edu-pretokenized-10K
Marin/Levanter Subsampled Pretokenized Dataset
Dataset
Train Urls:
gs://marin-us-central2/raw/fineweb-edu-c2beb4/3c452cb/huggingface.co/datasets/HuggingFaceFW/fineweb-edu/resolve/3c452cb
Factsheet
Original cache: gs://marin-us-central2/tokenized/fineweb-edu-24698d
Tokenizer: stanford-crfm/marin-tokenizer
Seed 42
Number of tokens: 10,390
(This readme is automatically generated by Marin.)
ultrachat-10k-chatmlKodCode-Light-RL-10K
🐱 KodCode: A Diverse, Challenging, and Verifiable Synthetic Dataset for Coding
KodCode is the largest fully-synthetic open-source dataset providing verifiable solutions and tests for coding tasks. It contains 12 distinct subsets spanning various domains (from algorithmic to package-specific knowledge) and difficulty levels (from basic coding exercises to interview and competitive programming challenges). KodCode is designed for both supervised fine-tuning (SFT) and RL tuning.
🕸️… See the full description on the dataset page: https://huggingface.co/datasets/KodCode/KodCode-Light-RL-10K.OpenMath-Vision-CoT-10kopenwebtext-10kAn open-source replication of the WebText dataset from OpenAI.
This is a small subset representing the first 10K records from the original dataset - created for testing.
The full 8M-record dataset is at https://huggingface.co/datasets/openwebtextc4-10k
Dataset Card for "c4-10k"
More Information needed
AirGoal-10k
AirGoal-10k
AirGoal-10k is an aerial image-goal navigation dataset released with
UA-NWM: Uncertainty-Aware World Model for Aerial Image-Goal Navigation.
Project page: https://duryi.github.io/UA-NWM-Project-Page/Code: https://github.com/DurYi/UA-NWMPaper: https://arxiv.org/abs/2608.05597
Dataset Summary
AirGoal-10k contains 11,000 aerial navigation trajectories for image-goal navigation. Each trajectory
contains 12 RGB observations and trajectory metadata. The test… See the full description on the dataset page: https://huggingface.co/datasets/DurYi/AirGoal-10k.Pytorch-Code-10K
Hot Coco Training Dataset
A curated collection of 10,625 high-quality PyTorch and Transformers code examples with AI-generated captions. This dataset was specifically built for fine-tuning code-specialized language models like Qimi (Coming soon!)
Dataset Description
This dataset contains Python code snippets sourced from open-source repositories that utilize PyTorch or Hugging Face Transformers. Each sample includes:
code: The raw Python source code (typically… See the full description on the dataset page: https://huggingface.co/datasets/Monster-Code/Pytorch-Code-10K.Articraft-10KThis repository contains the 10k articulated 3D objects (in URDF format) from Articraft-10K.
Articraft-10K is a large-scale articulated 3D dataset generated by the Articraft agent.
sec-10k-markdown-uncompressed
📄 SEC 10-K Full Uncompressed Markdown Filings (12.3k Documents)
Dataset Summary
This dataset contains 12,361 full-length, uncompressed SEC Form 10-K annual reports converted from EDGAR HTML to clean Markdown format across 1,379 companies (spanning 2004 to 2025, core 2014–2025).
The dataset is organized as uncompressed Markdown files structured by company ticker subdirectories (AAPL/10-K_2024.md, NVDA/10-K_2024.md, etc.), complete with company metadata manifests… See the full description on the dataset page: https://huggingface.co/datasets/astr010/sec-10k-markdown-uncompressed.VoCo-10kDataset for CVPR 2024 paper, "VoCo: A Simple-yet-Effective Volume Contrastive Learning Framework for 3D Medical Image Analysis"
https://arxiv.org/abs/2402.17300
Authors: Linshan Wu, Jiaxin Zhuang, and Hao Chen
Download Dataset
cd VoCo
mkdir data
huggingface-cli download Luffy503/VoCo-10k --repo-type dataset --local-dir . --cache-dir ./cache
libgen-10k-100kEgocentric_10K_Evaluation
Dataset Card for Egocentric_10K_Evaluation
This is a FiftyOne dataset with 30000 samples.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import fiftyone as fo
from fiftyone.utils.huggingface import load_from_hub
# Load the dataset
# Note: other available arguments include 'max_samples', etc
dataset = load_from_hub("Voxel51/Egocentric_10K_Evaluation")
# Launch the App
session = fo.launch_app(dataset)
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/Egocentric_10K_Evaluation.
