CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01RekaAI /RekaDaily-10k-raw RekaDaily-10k (raw) Raw, unscripted, first-person daily-life video, collected through Claru, Reka's data collection marketplace — recorded by paid collectors in their own homes and workplaces on head-mounted and handheld phones, across multiple regions. Videos are delivered as recorded — no cuts, no trimming, no editing, no filtering beyond basic integrity checks. A processed tier (short clips with machine captions) is released separately under the same RekaDaily-10k prefix.… See the full description on the dataset page: https://huggingface.co/datasets/RekaAI/RekaDaily-10k-raw.imagevideo-classification100K<n<1M22 likes223k downloads11d agoHugging Face02NeelNanda /pile-10kThe first 10K elements of The Pile, useful for debugging models trained on it. See the HuggingFace page for the full Pile for more info. Inspired by stas' great resource doing the same for OpenWebText text10K<n<100K34 likes38k downloads4y agoHugging Face03timaeus /dsir-pile-10ktext10K<n<100K0 likes18k downloads2y agoHugging Face04RekaAI /RekaDaily-10k-processed RekaDaily-10k (processed) Short first-person clips cut from the RekaDaily-10k recordings — unscripted daily-life video collected through Claru, Reka's data collection marketplace, recorded by paid collectors in their own homes and workplaces on head-mounted and handheld phones, across multiple regions. Every clip carries one dense caption and a multi-question Q&A exchange written in the second person ("What am I doing in this video?"), so the corpus drops straight into… See the full description on the dataset page: https://huggingface.co/datasets/RekaAI/RekaDaily-10k-processed.imagevideo-text-to-text1M<n<10M2 likes17k downloads11d agoHugging Face05ZeroOneCreative /amara-spatial-10k AmaraSpatial-10K A Semantically Anchored, Metric-Scale 3D Dataset for Embodied AI and Spatial Computing 10,071 AI-generated 3D meshes across 10 top-level categories and 476 subcategories — from basilisks to bassoons, cottages to cosmic stations — curated by Zero One Creative to close the spatial alignment gap that makes most generative 3D repositories unusable for zero-shot deployment in game engines, robotics simulators, and AR/VR pipelines. Every asset is… See the full description on the dataset page: https://huggingface.co/datasets/ZeroOneCreative/amara-spatial-10k.imagetext-to-3d10K<n<100K11 likes11k downloads5mo agoHugging Face06artefactory /Argimi-Ardian-Finance-10k-text The ArGiMI Ardian datasets : Text-only version The ArGiMi project is committed to open-source principles and data sharing. Thanks to our generous partners, we are releasing several valuable datasets to the public. Dataset description This text-only dataset comprises 34,000 financial annual reports, written in English, meticulously extracted from their original PDF format to provide a valuable resource for researchers and developers in financial analysis and natural… See the full description on the dataset page: https://huggingface.co/datasets/artefactory/Argimi-Ardian-Finance-10k-text.texttext-retrieval1M<n<10M19 likes9.2k downloads7mo agoHugging Face07jlohding /sp500-edgar-10k Dataset Card for SP500-EDGAR-10K Dataset Summary This dataset contains the annual reports for all SP500 historical constituents from 2010-2022 from SEC EDGAR Form 10-K filings. It also contains n-day future returns of each firm's stock price from each filing date. Dataset Structure Data Fields [More Information Needed] Data Splits [More Information Needed] Dataset Creation Source Data Initial Data Collection… See the full description on the dataset page: https://huggingface.co/datasets/jlohding/sp500-edgar-10k.tabular1K<n<10K22 likes7k downloads3y agoHugging Face08prquan /STARK_10k STARK: Spatial-Temporal reAsoning benchmaRK STARK is a comprehensive benchmark designed to systematically evaluate large language models (LLMs) and large reasoning models (LRMs) on spatial-temporal reasoning tasks, particularly for applications in cyber-physical systems (CPS) such as robotics, autonomous vehicles, and smart city infrastructure. Dataset Summary Hierarchical Benchmark: Tasks are structured across three levels of reasoning complexity: State Estimation:… See the full description on the dataset page: https://huggingface.co/datasets/prquan/STARK_10k.textquestion-answering10K<n<100K1 likes6.7k downloads11mo agoHugging Face09zai-org /LongAlign-10k LongAlign-10k 🤗 [LongAlign Dataset] • 💻 [Github Repo] • 📃 [LongAlign Paper] LongAlign is the first full recipe for LLM alignment on long context. We propose the LongAlign-10k dataset, containing 10,000 long instruction data of 8k-64k in length. We investigate on trianing strategies, namely packing (with loss weighting) and sorted batching, which are all implemented in our code. For real-world long context evaluation, we introduce LongBench-Chat that evaluate the… See the full description on the dataset page: https://huggingface.co/datasets/zai-org/LongAlign-10k.textquestion-answering1K<n<10K100 likes5.9k downloads3y agoHugging Face10smangrul /ultrachat-10k-chatmltext10K<n<100K6 likes5.1k downloads3y agoHugging Face11NeelNanda /c4-10k Dataset Card for "c4-10k" More Information needed text10K<n<100K0 likes5.1k downloads4y agoHugging Face12KodCode /KodCode-Light-RL-10K 🐱 KodCode: A Diverse, Challenging, and Verifiable Synthetic Dataset for Coding KodCode is the largest fully-synthetic open-source dataset providing verifiable solutions and tests for coding tasks. It contains 12 distinct subsets spanning various domains (from algorithmic to package-specific knowledge) and difficulty levels (from basic coding exercises to interview and competitive programming challenges). KodCode is designed for both supervised fine-tuning (SFT) and RL tuning. 🕸️… See the full description on the dataset page: https://huggingface.co/datasets/KodCode/KodCode-Light-RL-10K.tabularquestion-answering10K<n<100K9 likes4.8k downloads1y agoHugging Face13Asklv /OpenMath-Vision-CoT-10kimage10K<n<100K1 likes4.6k downloads9mo agoHugging Face14Monster-Code /Pytorch-Code-10K Hot Coco Training Dataset A curated collection of 10,625 high-quality PyTorch and Transformers code examples with AI-generated captions. This dataset was specifically built for fine-tuning code-specialized language models like Qimi (Coming soon!) Dataset Description This dataset contains Python code snippets sourced from open-source repositories that utilize PyTorch or Hugging Face Transformers. Each sample includes: code: The raw Python source code (typically… See the full description on the dataset page: https://huggingface.co/datasets/Monster-Code/Pytorch-Code-10K.texttext-generation1K<n<10K1 likes3.8k downloads2mo agoHugging Face15ziheng1234 /Critic-10K Critic-10K Dataset This repository hosts the Critic-10K dataset, introduced in the paper The Consistency Critic: Correcting Inconsistencies in Generated Images via Reference-Guided Attentive Alignment. The Critic-10K dataset is specifically constructed to address and rectify inconsistencies in generated images. It comprises reference-degraded-target triplets, obtained through VLM-based selection and explicit degradation. This dataset effectively simulates common inaccuracies or… See the full description on the dataset page: https://huggingface.co/datasets/ziheng1234/Critic-10K.imageimage-to-image1K<n<10K8 likes3.3k downloads10mo agoHugging Face16astr010 /sec-10k-markdown-uncompressed 📄 SEC 10-K Full Uncompressed Markdown Filings (12.3k Documents) Dataset Summary This dataset contains 12,361 full-length, uncompressed SEC Form 10-K annual reports converted from EDGAR HTML to clean Markdown format across 1,379 companies (spanning 2004 to 2025, core 2014–2025). The dataset is organized as uncompressed Markdown files structured by company ticker subdirectories (AAPL/10-K_2024.md, NVDA/10-K_2024.md, etc.), complete with company metadata manifests… See the full description on the dataset page: https://huggingface.co/datasets/astr010/sec-10k-markdown-uncompressed.texttext-generation1K<n<10K0 likes2.8k downloads2mo agoHugging Face17sam-paech /wildchat_creative_writing_annotated_10ktabular10K<n<100K2 likes2.4k downloads9mo agoHugging Face18data-is-better-together /10k_prompts_ranked Dataset Card for 10k_prompts_ranked 10k_prompts_ranked is a dataset of prompts with quality rankings created by 314 members of the open-source ML community using Argilla, an open-source tool to label data. The prompts in this dataset include both synthetic and human-generated prompts sourced from a variety of heavily used datasets that include prompts. The dataset contains 10,331 examples and can be used for training and evaluating language models on prompt ranking tasks. The… See the full description on the dataset page: https://huggingface.co/datasets/data-is-better-together/10k_prompts_ranked.tabulartext-classification10K<n<100K170 likes2k downloads3y agoHugging Face19laicsiifes /IntegraCAR-LULC-10K IntegraCAR-LULC-10K: A High-Resolution Optical Satellite Dataset for LULC Segmentation in the Brazilian Rural Environmental Registry [!WARNING] ⚠️ High Volume & Storage Advisory (+300 GB) This repository hosts the complete 10,000-tile collection (IntegraCAR-LULC-10K), consisting of over 300 GB of high-resolution satellite imagery ( 2048×20482048 \times 20482048×2048 px at 0.5&nbsp;m/px0.5\text{ m/px}0.5&nbsp;m/px ) and pixel-level segmentation masks stored in… See the full description on the dataset page: https://huggingface.co/datasets/laicsiifes/IntegraCAR-LULC-10K.imageimage-segmentation10K<n<100K3 likes1.9k downloads20d agoHugging Face20weikaih /ai2thor-perspective-qa-10kimage10K<n<100K0 likes1.9k downloads1y agoHugging Face21PKU-Alignment /PKU-SafeRLHF-10K Paper You can find more information in our paper. Dataset Paper: https://arxiv.org/abs/2307.04657 tabulartext-generation10K<n<100K62 likes1.7k downloads3y agoHugging Face22virattt /financial-qa-10Ktext1K<n<10K102 likes1.7k downloads2y agoHugging Face23winterForestStump /10-K_sec_filings Dataset Card for "10-K_sec_filings" Dataset of 93.5K 10K SEC EDGAR filings since 1999 year. This dataset contains a lot of bad parsed filings and also empty rows More Information needed text10K<n<100K3 likes1.5k downloads3y agoHugging Face24smgjch /meow-10k Dataset Card for Meow-10K Meow-10K is a high-fidelity, synchronized quad-modal dataset comprising 10,000 feline samples. It is the primary training corpus for Meow-Omni 1, designed to facilitate deep intention reasoning in computational ethology. Dataset Summary Meow-10K provides the first large-scale training foundation for Multimodal Large Language Models (MLLMs) to learn the causal relationships between external behaviours and internal physiological states. By… See the full description on the dataset page: https://huggingface.co/datasets/smgjch/meow-10k.audio10K<n<100K3 likes1.5k downloads4mo agoHugging Face25mrzjy /splash-art-gacha-collection-10k Splash Art Collection 10K This collection features 11,755 character splash arts or 角色立绘 sourced from 47 gacha games, meticulously gathered from Fandom and Biligame WIKI. The dataset is suitable for fine-tuning T2I models on splash art generation domain, utilizing the image and prompt fields. It includes a mix of both high- and low-quality splash arts of various styles, allowing you to curate and select the images that best suit your training needs. Data Structure… See the full description on the dataset page: https://huggingface.co/datasets/mrzjy/splash-art-gacha-collection-10k.image1K<n<10K6 likes1.2k downloads2y agoHugging Face26nichiriu /ai-sec-10k-filingstext0 likes1.2k downloads3mo agoHugging Face27TripVVT /TripVVT-10Kgated TripVVT-10K Dataset News 2026.06: TripVVT has been accepted by ECCV 2026. 2026.04: The TripVVT paper is available on arXiv. The project page is available at https://shaodingbao.github.io/TripVVT/. TripVVT-10K is a large-scale dataset for in-the-wild Video Virtual Try-On (VVT). It contains 10,031 high-quality video samples with triplet supervision, covering upper-body garments, lower-body garments, and dresses. TripVVT-10K is released together with the… See the full description on the dataset page: https://huggingface.co/datasets/TripVVT/TripVVT-10K.imageimage-to-video10K<n<100K6 likes1.1k downloads3mo agoHugging Face28bghira /pseudo-camera-10k pseudo-camera-10k dataset Contents This dataset contains 10k free images from world class photographers. The images have been resized using Lanczos antialiasing, with their smaller edge shifted to 1024px. The aim of this dataset is a highly variable but high quality and high resolution set of images containing difficult concepts, with about half of the images being numbered group shots and family portraits with the number of subjects labeled. No images were upsampled in… See the full description on the dataset page: https://huggingface.co/datasets/bghira/pseudo-camera-10k.image1K<n<10K52 likes1.1k downloads2y agoHugging Face29parler-tts /mls_eng_10k Dataset Summary This is a 10K hours subset of English version of the Multilingual LibriSpeech (MLS) dataset. The data archives were restructured from the original ones from OpenSLR to make it easier to stream. MLS dataset is a large multilingual corpus suitable for speech research. The dataset is derived from read audiobooks from LibriVox and consists of 8 languages - English, German, Dutch, Spanish, French, Italian, Portuguese, Polish. It includes about 44.5K hours of English and… See the full description on the dataset page: https://huggingface.co/datasets/parler-tts/mls_eng_10k.audioautomatic-speech-recognition1M<n<10M31 likes1.1k downloads2y agoHugging Face30stas /c4-en-10kThis is a small subset representing the first 10K records of the original C4 dataset, "en" subset - created for testing. The records were extracted after having been shuffled. The full 1TB+ dataset is at https://huggingface.co/datasets/c4.text10K<n<100K5 likes997 downloads4y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.