datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
RekaDaily-10k-raw
RekaDaily-10k (raw)
Raw, unscripted, first-person daily-life video, collected through
Claru, Reka's data collection marketplace — recorded by
paid collectors in their own homes and workplaces on head-mounted and handheld
phones, across multiple regions.
Videos are delivered as recorded — no cuts, no trimming, no editing, no
filtering beyond basic integrity checks. A processed tier (short clips with
machine captions) is released separately under the same RekaDaily-10k prefix.… See the full description on the dataset page: https://huggingface.co/datasets/RekaAI/RekaDaily-10k-raw.pile-10kThe first 10K elements of The Pile, useful for debugging models trained on it. See the HuggingFace page for the full Pile for more info. Inspired by stas' great resource doing the same for OpenWebText
dsir-pile-10kRekaDaily-10k-processed
RekaDaily-10k (processed)
Short first-person clips cut from the RekaDaily-10k
recordings —
unscripted daily-life video collected through Claru, Reka's
data collection marketplace, recorded by paid collectors in their own homes and
workplaces on head-mounted and handheld phones, across multiple regions.
Every clip carries one dense caption and a multi-question Q&A exchange
written in the second person ("What am I doing in this video?"), so the corpus
drops straight into… See the full description on the dataset page: https://huggingface.co/datasets/RekaAI/RekaDaily-10k-processed.amara-spatial-10k
AmaraSpatial-10K
A Semantically Anchored, Metric-Scale 3D Dataset for Embodied AI and Spatial Computing
10,071 AI-generated 3D meshes across 10 top-level categories and 476 subcategories — from basilisks to bassoons, cottages to cosmic stations — curated by Zero One Creative to close the spatial alignment gap that makes most generative 3D repositories unusable for zero-shot deployment in game engines, robotics simulators, and AR/VR pipelines.
Every asset is… See the full description on the dataset page: https://huggingface.co/datasets/ZeroOneCreative/amara-spatial-10k.Argimi-Ardian-Finance-10k-text
The ArGiMI Ardian datasets : Text-only version
The ArGiMi project is committed to open-source principles and data sharing.
Thanks to our generous partners, we are releasing several valuable datasets to the public.
Dataset description
This text-only dataset comprises 34,000 financial annual reports, written in English, meticulously
extracted from their original PDF format to provide a valuable resource for researchers and developers in financial
analysis and natural… See the full description on the dataset page: https://huggingface.co/datasets/artefactory/Argimi-Ardian-Finance-10k-text.sp500-edgar-10k
Dataset Card for SP500-EDGAR-10K
Dataset Summary
This dataset contains the annual reports for all SP500 historical constituents from 2010-2022 from SEC EDGAR Form 10-K filings.
It also contains n-day future returns of each firm's stock price from each filing date.
Dataset Structure
Data Fields
[More Information Needed]
Data Splits
[More Information Needed]
Dataset Creation
Source Data
Initial Data Collection… See the full description on the dataset page: https://huggingface.co/datasets/jlohding/sp500-edgar-10k.STARK_10k
STARK: Spatial-Temporal reAsoning benchmaRK
STARK is a comprehensive benchmark designed to systematically evaluate large language models (LLMs) and large reasoning models (LRMs) on spatial-temporal reasoning tasks, particularly for applications in cyber-physical systems (CPS) such as robotics, autonomous vehicles, and smart city infrastructure.
Dataset Summary
Hierarchical Benchmark: Tasks are structured across three levels of reasoning complexity:
State Estimation:… See the full description on the dataset page: https://huggingface.co/datasets/prquan/STARK_10k.LongAlign-10k
LongAlign-10k
🤗 [LongAlign Dataset] • 💻 [Github Repo] • 📃 [LongAlign Paper]
LongAlign is the first full recipe for LLM alignment on long context. We propose the LongAlign-10k dataset, containing 10,000 long instruction data of 8k-64k in length. We investigate on trianing strategies, namely packing (with loss weighting) and sorted batching, which are all implemented in our code. For real-world long context evaluation, we introduce LongBench-Chat that evaluate the… See the full description on the dataset page: https://huggingface.co/datasets/zai-org/LongAlign-10k.ultrachat-10k-chatmlc4-10k
Dataset Card for "c4-10k"
More Information needed
KodCode-Light-RL-10K
🐱 KodCode: A Diverse, Challenging, and Verifiable Synthetic Dataset for Coding
KodCode is the largest fully-synthetic open-source dataset providing verifiable solutions and tests for coding tasks. It contains 12 distinct subsets spanning various domains (from algorithmic to package-specific knowledge) and difficulty levels (from basic coding exercises to interview and competitive programming challenges). KodCode is designed for both supervised fine-tuning (SFT) and RL tuning.
🕸️… See the full description on the dataset page: https://huggingface.co/datasets/KodCode/KodCode-Light-RL-10K.OpenMath-Vision-CoT-10kPytorch-Code-10K
Hot Coco Training Dataset
A curated collection of 10,625 high-quality PyTorch and Transformers code examples with AI-generated captions. This dataset was specifically built for fine-tuning code-specialized language models like Qimi (Coming soon!)
Dataset Description
This dataset contains Python code snippets sourced from open-source repositories that utilize PyTorch or Hugging Face Transformers. Each sample includes:
code: The raw Python source code (typically… See the full description on the dataset page: https://huggingface.co/datasets/Monster-Code/Pytorch-Code-10K.Critic-10K
Critic-10K Dataset
This repository hosts the Critic-10K dataset, introduced in the paper The Consistency Critic: Correcting Inconsistencies in Generated Images via Reference-Guided Attentive Alignment.
The Critic-10K dataset is specifically constructed to address and rectify inconsistencies in generated images. It comprises reference-degraded-target triplets, obtained through VLM-based selection and explicit degradation. This dataset effectively simulates common inaccuracies or… See the full description on the dataset page: https://huggingface.co/datasets/ziheng1234/Critic-10K.sec-10k-markdown-uncompressed
📄 SEC 10-K Full Uncompressed Markdown Filings (12.3k Documents)
Dataset Summary
This dataset contains 12,361 full-length, uncompressed SEC Form 10-K annual reports converted from EDGAR HTML to clean Markdown format across 1,379 companies (spanning 2004 to 2025, core 2014–2025).
The dataset is organized as uncompressed Markdown files structured by company ticker subdirectories (AAPL/10-K_2024.md, NVDA/10-K_2024.md, etc.), complete with company metadata manifests… See the full description on the dataset page: https://huggingface.co/datasets/astr010/sec-10k-markdown-uncompressed.wildchat_creative_writing_annotated_10k10k_prompts_ranked
Dataset Card for 10k_prompts_ranked
10k_prompts_ranked is a dataset of prompts with quality rankings created by 314 members of the open-source ML community using Argilla, an open-source tool to label data. The prompts in this dataset include both synthetic and human-generated prompts sourced from a variety of heavily used datasets that include prompts.
The dataset contains 10,331 examples and can be used for training and evaluating language models on prompt ranking tasks. The… See the full description on the dataset page: https://huggingface.co/datasets/data-is-better-together/10k_prompts_ranked.IntegraCAR-LULC-10K
IntegraCAR-LULC-10K: A High-Resolution Optical Satellite Dataset for LULC Segmentation in the Brazilian Rural Environmental Registry
[!WARNING]
⚠️ High Volume & Storage Advisory (+300 GB)
This repository hosts the complete 10,000-tile collection (IntegraCAR-LULC-10K), consisting of over 300 GB of high-resolution satellite imagery ( 2048×20482048 \times 20482048×2048 px at 0.5 m/px0.5\text{ m/px}0.5 m/px ) and pixel-level segmentation masks stored in… See the full description on the dataset page: https://huggingface.co/datasets/laicsiifes/IntegraCAR-LULC-10K.ai2thor-perspective-qa-10kPKU-SafeRLHF-10K
Paper
You can find more information in our paper.
Dataset Paper: https://arxiv.org/abs/2307.04657
financial-qa-10K10-K_sec_filings
Dataset Card for "10-K_sec_filings"
Dataset of 93.5K 10K SEC EDGAR filings since 1999 year. This dataset contains a lot of bad parsed filings and also empty rows
More Information needed
meow-10k
Dataset Card for Meow-10K
Meow-10K is a high-fidelity, synchronized quad-modal dataset comprising 10,000 feline samples. It is the primary training corpus for Meow-Omni 1, designed to facilitate deep intention reasoning in computational ethology.
Dataset Summary
Meow-10K provides the first large-scale training foundation for Multimodal Large Language Models (MLLMs) to learn the causal relationships between external behaviours and internal physiological states. By… See the full description on the dataset page: https://huggingface.co/datasets/smgjch/meow-10k.splash-art-gacha-collection-10k
Splash Art Collection 10K
This collection features 11,755 character splash arts or 角色立绘 sourced from 47 gacha games, meticulously gathered from Fandom and Biligame WIKI.
The dataset is suitable for fine-tuning T2I models on splash art generation domain, utilizing the image and prompt fields. It includes a mix of both high- and low-quality splash arts of various styles, allowing you to curate and select the images that best suit your training needs.
Data Structure… See the full description on the dataset page: https://huggingface.co/datasets/mrzjy/splash-art-gacha-collection-10k.ai-sec-10k-filingsTripVVT-10K
TripVVT-10K Dataset
News
2026.06: TripVVT has been accepted by ECCV 2026.
2026.04: The TripVVT paper is available on arXiv.
The project page is available at https://shaodingbao.github.io/TripVVT/.
TripVVT-10K is a large-scale dataset for in-the-wild Video Virtual Try-On (VVT). It contains 10,031 high-quality video samples with triplet supervision, covering upper-body garments, lower-body garments, and dresses.
TripVVT-10K is released together with the… See the full description on the dataset page: https://huggingface.co/datasets/TripVVT/TripVVT-10K.pseudo-camera-10k
pseudo-camera-10k dataset
Contents
This dataset contains 10k free images from world class photographers. The images have been resized using Lanczos antialiasing, with their smaller edge shifted to 1024px.
The aim of this dataset is a highly variable but high quality and high resolution set of images containing difficult concepts, with about half of the images being numbered group shots and family portraits with the number of subjects labeled.
No images were upsampled in… See the full description on the dataset page: https://huggingface.co/datasets/bghira/pseudo-camera-10k.mls_eng_10k
Dataset Summary
This is a 10K hours subset of English version of the Multilingual LibriSpeech (MLS) dataset.
The data archives were restructured from the original ones from OpenSLR to make it easier to stream.
MLS dataset is a large multilingual corpus suitable for speech research. The dataset is derived from read audiobooks from LibriVox and consists of
8 languages - English, German, Dutch, Spanish, French, Italian, Portuguese, Polish. It includes about 44.5K hours of English and… See the full description on the dataset page: https://huggingface.co/datasets/parler-tts/mls_eng_10k.c4-en-10kThis is a small subset representing the first 10K records of the original C4 dataset, "en" subset - created for testing. The records were extracted after having been shuffled.
The full 1TB+ dataset is at https://huggingface.co/datasets/c4.
