CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01genrobot2025 /10Kh-RealOmin-OpenDatagated Boasting over 13,000 hours of cumulative data and 5 million+ clips, it ranks as the largest open-source embodied intelligence dataset in the industry. Update Notes:Stage 3 data upload completed. 13,000+ hours of pure dual-hand data with frame-level alignment latency < 1ms Full high-precision trajectory reconstruction, breaking the limit of superficial open source, fully ready-to-use 3,000+ contributors and 10,000+ real household scenarios with exceptional diversity… See the full description on the dataset page: https://huggingface.co/datasets/genrobot2025/10Kh-RealOmin-OpenData.videoroboticsn>1T272 likes525k downloads5mo agoHugging Face02nvidia /SAGE-10k SAGE-10k SAGE-10k is a large-scale interactive indoor scene dataset featuring realistic layouts, generated by the agentic-driven pipeline introduced in "SAGE: Scalable Agentic 3D Scene Generation for Embodied AI". The dataset contains 10,000 diverse scenes spanning 50 room types and styles, along with 565K uniquely generated 3D objects. 🔑 Key Features SAGE-10k integrates a wide variety of scenes, and particularly, preserves small items… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/SAGE-10k.text-to-3d10K<n<100K85 likes391k downloads7mo agoHugging Face03RekaAI /RekaDaily-10k-raw RekaDaily-10k (raw) Raw, unscripted, first-person daily-life video, collected through Claru, Reka's data collection marketplace — recorded by paid collectors in their own homes and workplaces on head-mounted and handheld phones, across multiple regions. Videos are delivered as recorded — no cuts, no trimming, no editing, no filtering beyond basic integrity checks. A processed tier (short clips with machine captions) is released separately under the same RekaDaily-10k prefix.… See the full description on the dataset page: https://huggingface.co/datasets/RekaAI/RekaDaily-10k-raw.imagevideo-classification100K<n<1M22 likes226k downloads10d agoHugging Face04builddotai /Egocentric-10Kgated Egocentric-10K is the largest egocentric dataset. It is the first dataset collected exclusively in real factories. Your browser does not support the video tag. Egocentric-10K is state-of-the-art in hand visibility and active manipulation density compared to previous in-the-wild egocentric datasets. The complete 30,000 frame evaluation set is available at Egocentric-10K-Evaluation. Dataset Statistics Attribute Value Total Hours 10,000 Total Frames 1.08 billion… See the full description on the dataset page: https://huggingface.co/datasets/builddotai/Egocentric-10K.350 likes70k downloads7mo agoHugging Face05NeelNanda /pile-10kThe first 10K elements of The Pile, useful for debugging models trained on it. See the HuggingFace page for the full Pile for more info. Inspired by stas' great resource doing the same for OpenWebText text10K<n<100K34 likes38k downloads4y agoHugging Face06imitator-game /IG-10K-Dataset The Imitator Game - IG-10K Dataset This dataset accompanies the paper The Imitator Game: Benchmarking Robot Imitative Ability Beyond Action Prediction. It contains paired human-robot demonstrations for the Imitator Game benchmark, spanning four levels of imitation difficulty (L0–L3) across real and simulated settings. The IG-10K dataset includes over 20,000 paired episodes across 50+ tasks and 6 domains, and is provided in LeRobot-0.5.0 format. For more details, see the project… See the full description on the dataset page: https://huggingface.co/datasets/imitator-game/IG-10K-Dataset.videorobotics1K<n<10K4 likes30k downloads21d agoHugging Face07timaeus /dsir-pile-10ktext10K<n<100K0 likes24k downloads2y agoHugging Face08AntoineGuedon /DL3DV-10K-Meshedgated DL3DV-10K-Meshed A derivative of DL3DV-10K providing undistorted 480P views together with ground-truth surface geometry, used to train Surflo. This dataset is not a replacement for DL3DV-10K. It redistributes some of DL3DV-10K imagery, and remains subject to the DL3DV-10K Terms of Use. See Licensing and terms before using it. What we changed Relative to DL3DV-ALL-480P: Undistorted images. All 480P frames were reprocessed with COLMAP to remove lens distortion.… See the full description on the dataset page: https://huggingface.co/datasets/AntoineGuedon/DL3DV-10K-Meshed.depth-estimation10K<n<100K6 likes23k downloads29d agoHugging Face09RekaAI /RekaDaily-10k-processed RekaDaily-10k (processed) Short first-person clips cut from the RekaDaily-10k recordings — unscripted daily-life video collected through Claru, Reka's data collection marketplace, recorded by paid collectors in their own homes and workplaces on head-mounted and handheld phones, across multiple regions. Every clip carries one dense caption and a multi-question Q&A exchange written in the second person ("What am I doing in this video?"), so the corpus drops straight into… See the full description on the dataset page: https://huggingface.co/datasets/RekaAI/RekaDaily-10k-processed.imagevideo-text-to-text1M<n<10M2 likes13k downloads10d agoHugging Face10ad1t7a /10Kh-RealOmin-OpenDataBoasting over 10,000 hours of cumulative data and 1 million+ clips, it ranks as the largest open-source embodied intelligence dataset in the industry. Compared with other datasets, it has the following advantages: Ample Data Volume & Strong Generalization Each skill is supported by sufficient data, collected from over 3,000 households and nearly 10,000 distinct fine-grained targets. It avoids simple repetitions and ensures robust generalization. Authentic Scenarios & Focused… See the full description on the dataset page: https://huggingface.co/datasets/ad1t7a/10Kh-RealOmin-OpenData.videoroboticsn>1T12 likes12k downloads9mo agoHugging Face11ZeroOneCreative /amara-spatial-10k AmaraSpatial-10K A Semantically Anchored, Metric-Scale 3D Dataset for Embodied AI and Spatial Computing 10,071 AI-generated 3D meshes across 10 top-level categories and 476 subcategories — from basilisks to bassoons, cottages to cosmic stations — curated by Zero One Creative to close the spatial alignment gap that makes most generative 3D repositories unusable for zero-shot deployment in game engines, robotics simulators, and AR/VR pipelines. Every asset is… See the full description on the dataset page: https://huggingface.co/datasets/ZeroOneCreative/amara-spatial-10k.imagetext-to-3d10K<n<100K11 likes11k downloads5mo agoHugging Face12artefactory /Argimi-Ardian-Finance-10k-text The ArGiMI Ardian datasets : Text-only version The ArGiMi project is committed to open-source principles and data sharing. Thanks to our generous partners, we are releasing several valuable datasets to the public. Dataset description This text-only dataset comprises 34,000 financial annual reports, written in English, meticulously extracted from their original PDF format to provide a valuable resource for researchers and developers in financial analysis and natural… See the full description on the dataset page: https://huggingface.co/datasets/artefactory/Argimi-Ardian-Finance-10k-text.texttext-retrieval1M<n<10M19 likes9.2k downloads7mo agoHugging Face13ZZoutian /database-10k-val database-10k-val 最终 40K 数据集的 10,000 样本验证子集(2026-09-17)。 选择约束 不含 inset:所有样本 layout_type != inset。 不含冻结/删除族:样本主 panel 与全部副 panel 的 chart_family 均不属于 area、matrix、set_relation。 从最终修复后的 40K 中按 (family, subtype, layout, difficulty) 分层、确定性抽取 10,000 条。 文件 figure2data_10k_val.sqlite3:子集 SQLite,10,000 samples / 70,000 documents。 images/:10,000 PNG。 documents/:10,000 JSON。 arrays/:有原始数组的样本 NPZ。 plans/generation_plan_10k_val.jsonl:10,000 行子集计划。… See the full description on the dataset page: https://huggingface.co/datasets/ZZoutian/database-10k-val.0 likes7.4k downloads4d agoHugging Face14jlohding /sp500-edgar-10k Dataset Card for SP500-EDGAR-10K Dataset Summary This dataset contains the annual reports for all SP500 historical constituents from 2010-2022 from SEC EDGAR Form 10-K filings. It also contains n-day future returns of each firm's stock price from each filing date. Dataset Structure Data Fields [More Information Needed] Data Splits [More Information Needed] Dataset Creation Source Data Initial Data Collection… See the full description on the dataset page: https://huggingface.co/datasets/jlohding/sp500-edgar-10k.tabular1K<n<10K22 likes7k downloads3y agoHugging Face15prquan /STARK_10k STARK: Spatial-Temporal reAsoning benchmaRK STARK is a comprehensive benchmark designed to systematically evaluate large language models (LLMs) and large reasoning models (LRMs) on spatial-temporal reasoning tasks, particularly for applications in cyber-physical systems (CPS) such as robotics, autonomous vehicles, and smart city infrastructure. Dataset Summary Hierarchical Benchmark: Tasks are structured across three levels of reasoning complexity: State Estimation:… See the full description on the dataset page: https://huggingface.co/datasets/prquan/STARK_10k.textquestion-answering10K<n<100K1 likes6.8k downloads11mo agoHugging Face16zai-org /LongAlign-10k LongAlign-10k 🤗 [LongAlign Dataset] • 💻 [Github Repo] • 📃 [LongAlign Paper] LongAlign is the first full recipe for LLM alignment on long context. We propose the LongAlign-10k dataset, containing 10,000 long instruction data of 8k-64k in length. We investigate on trianing strategies, namely packing (with loss weighting) and sorted batching, which are all implemented in our code. For real-world long context evaluation, we introduce LongBench-Chat that evaluate the… See the full description on the dataset page: https://huggingface.co/datasets/zai-org/LongAlign-10k.textquestion-answering1K<n<10K100 likes6.1k downloads3y agoHugging Face17chanind /c4-10k-mini-tokenized-16-ctx-gelu-1l-tests1K<n<10K0 likes5.6k downloads2y agoHugging Face18marin-community /fineweb-edu-pretokenized-10K Marin/Levanter Subsampled Pretokenized Dataset Dataset Train Urls: gs://marin-us-central2/raw/fineweb-edu-c2beb4/3c452cb/huggingface.co/datasets/HuggingFaceFW/fineweb-edu/resolve/3c452cb Factsheet Original cache: gs://marin-us-central2/tokenized/fineweb-edu-24698d Tokenizer: stanford-crfm/marin-tokenizer Seed 42 Number of tokens: 10,390 (This readme is automatically generated by Marin.) 0 likes5k downloads1y agoHugging Face19smangrul /ultrachat-10k-chatmltext10K<n<100K6 likes5k downloads3y agoHugging Face20KodCode /KodCode-Light-RL-10K 🐱 KodCode: A Diverse, Challenging, and Verifiable Synthetic Dataset for Coding KodCode is the largest fully-synthetic open-source dataset providing verifiable solutions and tests for coding tasks. It contains 12 distinct subsets spanning various domains (from algorithmic to package-specific knowledge) and difficulty levels (from basic coding exercises to interview and competitive programming challenges). KodCode is designed for both supervised fine-tuning (SFT) and RL tuning. 🕸️… See the full description on the dataset page: https://huggingface.co/datasets/KodCode/KodCode-Light-RL-10K.tabularquestion-answering10K<n<100K9 likes4.6k downloads1y agoHugging Face21Asklv /OpenMath-Vision-CoT-10kimage10K<n<100K1 likes4.6k downloads9mo agoHugging Face22stas /openwebtext-10kAn open-source replication of the WebText dataset from OpenAI. This is a small subset representing the first 10K records from the original dataset - created for testing. The full 8M-record dataset is at https://huggingface.co/datasets/openwebtext32 likes4.5k downloads5y agoHugging Face23NeelNanda /c4-10k Dataset Card for "c4-10k" More Information needed text10K<n<100K0 likes4.4k downloads4y agoHugging Face24DurYi /AirGoal-10k AirGoal-10k AirGoal-10k is an aerial image-goal navigation dataset released with UA-NWM: Uncertainty-Aware World Model for Aerial Image-Goal Navigation. Project page: https://duryi.github.io/UA-NWM-Project-Page/Code: https://github.com/DurYi/UA-NWMPaper: https://arxiv.org/abs/2608.05597 Dataset Summary AirGoal-10k contains 11,000 aerial navigation trajectories for image-goal navigation. Each trajectory contains 12 RGB observations and trajectory metadata. The test… See the full description on the dataset page: https://huggingface.co/datasets/DurYi/AirGoal-10k.imagerobotics100K<n<1M1 likes4k downloads28d agoHugging Face25Monster-Code /Pytorch-Code-10K Hot Coco Training Dataset A curated collection of 10,625 high-quality PyTorch and Transformers code examples with AI-generated captions. This dataset was specifically built for fine-tuning code-specialized language models like Qimi (Coming soon!) Dataset Description This dataset contains Python code snippets sourced from open-source repositories that utilize PyTorch or Hugging Face Transformers. Each sample includes: code: The raw Python source code (typically… See the full description on the dataset page: https://huggingface.co/datasets/Monster-Code/Pytorch-Code-10K.texttext-generation1K<n<10K1 likes3.8k downloads2mo agoHugging Face26camvsl /Articraft-10KThis repository contains the 10k articulated 3D objects (in URDF format) from Articraft-10K. Articraft-10K is a large-scale articulated 3D dataset generated by the Articraft agent. 38 likes3.8k downloads4mo agoHugging Face27astr010 /sec-10k-markdown-uncompressed 📄 SEC 10-K Full Uncompressed Markdown Filings (12.3k Documents) Dataset Summary This dataset contains 12,361 full-length, uncompressed SEC Form 10-K annual reports converted from EDGAR HTML to clean Markdown format across 1,379 companies (spanning 2004 to 2025, core 2014–2025). The dataset is organized as uncompressed Markdown files structured by company ticker subdirectories (AAPL/10-K_2024.md, NVDA/10-K_2024.md, etc.), complete with company metadata manifests… See the full description on the dataset page: https://huggingface.co/datasets/astr010/sec-10k-markdown-uncompressed.texttext-generation1K<n<10K0 likes3.3k downloads2mo agoHugging Face28Luffy503 /VoCo-10kDataset for CVPR 2024 paper, "VoCo: A Simple-yet-Effective Volume Contrastive Learning Framework for 3D Medical Image Analysis" https://arxiv.org/abs/2402.17300 Authors: Linshan Wu, Jiaxin Zhuang, and Hao Chen Download Dataset cd VoCo mkdir data huggingface-cli download Luffy503/VoCo-10k --repo-type dataset --local-dir . --cache-dir ./cache 10 likes2.7k downloads2y agoHugging Face29atom-in-the-universe /libgen-10k-100k0 likes2.6k downloads3y agoHugging Face30Voxel51 /Egocentric_10K_Evaluation Dataset Card for Egocentric_10K_Evaluation This is a FiftyOne dataset with 30000 samples. Installation If you haven't already, install FiftyOne: pip install -U fiftyone Usage import fiftyone as fo from fiftyone.utils.huggingface import load_from_hub # Load the dataset # Note: other available arguments include 'max_samples', etc dataset = load_from_hub("Voxel51/Egocentric_10K_Evaluation") # Launch the App session = fo.launch_app(dataset) Dataset… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/Egocentric_10K_Evaluation.imageimage-classification10K<n<100K1 likes2.5k downloads10mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.