datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
RekaDaily-10k-raw
RekaDaily-10k (raw)
Raw, unscripted, first-person daily-life video, collected through
Claru, Reka's data collection marketplace — recorded by
paid collectors in their own homes and workplaces on head-mounted and handheld
phones, across multiple regions.
Videos are delivered as recorded — no cuts, no trimming, no editing, no
filtering beyond basic integrity checks. A processed tier (short clips with
machine captions) is released separately under the same RekaDaily-10k prefix.… See the full description on the dataset page: https://huggingface.co/datasets/RekaAI/RekaDaily-10k-raw.RekaDaily-10k-processed
RekaDaily-10k (processed)
Short first-person clips cut from the RekaDaily-10k
recordings —
unscripted daily-life video collected through Claru, Reka's
data collection marketplace, recorded by paid collectors in their own homes and
workplaces on head-mounted and handheld phones, across multiple regions.
Every clip carries one dense caption and a multi-question Q&A exchange
written in the second person ("What am I doing in this video?"), so the corpus
drops straight into… See the full description on the dataset page: https://huggingface.co/datasets/RekaAI/RekaDaily-10k-processed.sp500-edgar-10k
Dataset Card for SP500-EDGAR-10K
Dataset Summary
This dataset contains the annual reports for all SP500 historical constituents from 2010-2022 from SEC EDGAR Form 10-K filings.
It also contains n-day future returns of each firm's stock price from each filing date.
Dataset Structure
Data Fields
[More Information Needed]
Data Splits
[More Information Needed]
Dataset Creation
Source Data
Initial Data Collection… See the full description on the dataset page: https://huggingface.co/datasets/jlohding/sp500-edgar-10k.KodCode-Light-RL-10K
🐱 KodCode: A Diverse, Challenging, and Verifiable Synthetic Dataset for Coding
KodCode is the largest fully-synthetic open-source dataset providing verifiable solutions and tests for coding tasks. It contains 12 distinct subsets spanning various domains (from algorithmic to package-specific knowledge) and difficulty levels (from basic coding exercises to interview and competitive programming challenges). KodCode is designed for both supervised fine-tuning (SFT) and RL tuning.
🕸️… See the full description on the dataset page: https://huggingface.co/datasets/KodCode/KodCode-Light-RL-10K.wildchat_creative_writing_annotated_10k10k_prompts_ranked
Dataset Card for 10k_prompts_ranked
10k_prompts_ranked is a dataset of prompts with quality rankings created by 314 members of the open-source ML community using Argilla, an open-source tool to label data. The prompts in this dataset include both synthetic and human-generated prompts sourced from a variety of heavily used datasets that include prompts.
The dataset contains 10,331 examples and can be used for training and evaluating language models on prompt ranking tasks. The… See the full description on the dataset page: https://huggingface.co/datasets/data-is-better-together/10k_prompts_ranked.PKU-SafeRLHF-10K
Paper
You can find more information in our paper.
Dataset Paper: https://arxiv.org/abs/2307.04657
TripVVT-10K
TripVVT-10K Dataset
News
2026.06: TripVVT has been accepted by ECCV 2026.
2026.04: The TripVVT paper is available on arXiv.
The project page is available at https://shaodingbao.github.io/TripVVT/.
TripVVT-10K is a large-scale dataset for in-the-wild Video Virtual Try-On (VVT). It contains 10,031 high-quality video samples with triplet supervision, covering upper-body garments, lower-body garments, and dresses.
TripVVT-10K is released together with the… See the full description on the dataset page: https://huggingface.co/datasets/TripVVT/TripVVT-10K.sec-10k-lsh-chunks
📈 SEC 10-K Cleaned Text Chunks & LSH Boilerplate Dataset
Dataset Summary
This dataset contains 13,562,130 cleaned text chunks extracted from 12,361 SEC Form 10-K annual filings across 1,380 companies (spanning 2004 to 2025, core 2014–2025).
Every chunk across all 1,380 companies (including mega-cap leaders such as AAPL, MSFT, NVDA, AMZN, GOOGL, META, TSLA, JPM, WMT, XOM, AVGO, LLY) is annotated with metadata, token counts, table indicators, and a pre-computed… See the full description on the dataset page: https://huggingface.co/datasets/astr010/sec-10k-lsh-chunks.parler-tts_mls_eng_10k_snac_token_old
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/blanchon/parler-tts_mls_eng_10k_snac_token_old.anki-revlogs-10k
Introduction
Anki Revlogs 10K is a dataset of 10k collections from Anki for FSRS project. It is a random sample of collections with 5000+ revlog entries, so it should contain a mix of older (still active) users, and newer users.
The dataset contains three parts: revlogs, cards, and decks.
Revlogs
This dataset contains flashcard review records with the following fields:
card_id: Unique identifier for each flashcard
day_offset: Number of days since the start of… See the full description on the dataset page: https://huggingface.co/datasets/open-spaced-repetition/anki-revlogs-10k.VENUS-10K
Dataset Card for VENUS
Dataset Summary
Data from: Speaking Beyond Language: A Large-Scale Multimodal Dataset for Learning Nonverbal Cues from Video-Grounded Dialogues
@article{kim2025speaking,
title={Speaking Beyond Language: A Large-Scale Multimodal Dataset for Learning Nonverbal Cues from Video-Grounded Dialogues},
author={Kim, Youngmin and Chung, Jiwan and Kim, Jisoo and Lee, Sunghyun and Lee, Sangkyu and Kim, Junhyeok and Yang, Cheoljong and Yu, Youngjae}… See the full description on the dataset page: https://huggingface.co/datasets/winston1214/VENUS-10K.China-K12-STEM-10K-CoT-Reasoning
K12-STEM-CoT-Chinese
1.54M Chinese K12 STEM problems with chain-of-thought solutions, 48% with diagrams.
The largest structured Chinese math/physics/chemistry reasoning dataset.
This is a curated sample (10,000 problems) of the full 1.54M dataset available via API.
Full Dataset Access
Access the full 1,540,000+ problems via API →
This Sample
Full API
Total problems
10,025
1,540,000+
With CoT solutions
10,025
1,490,000+
With diagrams
6,093
740,000+… See the full description on the dataset page: https://huggingface.co/datasets/lfaviate/China-K12-STEM-10K-CoT-Reasoning.cinepile_10kedit_10k_charomx_f_16ten_10kai_2camera_blue_block_pickThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "omx_f",
"total_episodes": 5,
"total_frames": 1503,
"total_tasks": 1,
"total_videos": 10,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 15,
"splits": {
"train": "0:5"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/hicohico1/omx_f_16ten_10kai_2camera_blue_block_pick.PhysDPO-10kTo use PhysDPO, run the following command to combine the parts into a single ZIP file:
cat PhysDPO_part_* > PhysDPO.zip
Apple-10K-2022bangla-10k
Bangla-10K: A Challenging, Metadata-Rich Corpus of Read and Conversational Bengali Speech from India and Bangladesh
Bangla-10K is a 10,816-hour Bengali speech corpus with
624,951 recordings from India and Bangladesh: a 10,070.8-hour core corpus
(567,323 recordings) and a separately collected 745.1-hour evaluation set
(57,628 recordings). It combines scripted single-speaker read speech with
natural multi-speaker conversations for Bengali automatic speech recognition
(ASR).
The… See the full description on the dataset page: https://huggingface.co/datasets/psdn-ai/bangla-10k.anki-revlogs-10k-id
Anki Revlogs 10K (-id variant)
This is a derived, analysis-ready view of the same 10,000 Anki collections as
Anki Revlogs 10K.
It is generated from the corresponding
raw export,
so it does not contain information that is not present in that source.
The three tables and user partitioning are unchanged. The -id variant differs in two ways:
entity IDs in cards and decks, and card_id in revlogs, are preserved as the original
signed 64-bit Anki values instead of being factorized to… See the full description on the dataset page: https://huggingface.co/datasets/open-spaced-repetition/anki-revlogs-10k-id.RAIL-HH-10K RAIL-HH-10K: Multi-Dimensional Safety Alignment Dataset
The first large-scale safety dataset with 99.5% multi-dimensional annotation coverage across 8 ethical dimensions.
📖 Read Blog •
📖 Paper (Coming Soon) •
🚀 Quick Start •
🔌 RAIL API •
💻 Examples
🌟 What Makes RAIL-HH-10K Special?
🎯 Near-Complete Coverage
99.5% dimension coverage across all 8 ethical dimensions
Most existing datasets: 40-70% coverage
RAIL-HH-10K: 98-100%… See the full description on the dataset page: https://huggingface.co/datasets/responsible-ai-labs/RAIL-HH-10K.zelo-scores-10kx100-gemma-4-31b-itzelo-scores-10kx100-gpt-oss-20bspotify-top-10k-songsthis list has been extracted from anna's archive : https://annas-archive.li/blog/spotify/spotify-top-10k-songs-table.html
the script used to scrape can be found here : https://gist.github.com/the-code-rider/96838f5d6ff538377776b6ddbb1c633d
code-10k
Dataset Card for "code-10k"
More Information needed
dynamic_robot_bench_dr_scripted_10k
dynamic_robot_bench_dr_scripted_10k
10,000 scripted-expert demonstrations across all 100 dynamic task families of
dynamic-robot-bench — a conveyor-belt dynamic-manipulation benchmark (Franka
Panda + wrist camera, ManiSkill 3 / SAPIEN GPU sim). One LeRobot v2.1 dataset:
100 episodes per family, success-filtered, language-prompted per episode.
Collection configuration (identical for every family)
Scripted expert with per-step auto-derived speed caps, recorded as… See the full description on the dataset page: https://huggingface.co/datasets/Damin3927/dynamic_robot_bench_dr_scripted_10k.lol-basic-matches-challenger-10k
GPTilt: 10K League of Legends Challenger Matches
This dataset is part of the GPTilt open-source initiative, aimed at democratizing access to high-quality LoL data for research and analysis, fostering public exploration, and advancing the community's understanding of League of Legends through data science and AI. It provides detailed data from high-elo matches.
By using this dataset, users accept full responsibility for any consequences arising from its use. GPTilt assumes no… See the full description on the dataset page: https://huggingface.co/datasets/gptilt/lol-basic-matches-challenger-10k.2026-09-03_shake4it_bench_5sensors_v3_dragonfly_10kHz_nfft_512This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/jogarulfop/2026-09-03_shake4it_bench_5sensors_v3_dragonfly_10kHz_nfft_512.2026-09-03_shake4it_bench_5sensors_v3_dgf_passif_10kHz_nfft_512This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/jogarulfop/2026-09-03_shake4it_bench_5sensors_v3_dgf_passif_10kHz_nfft_512.threat-detection-responses-10k
