CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01zai-org /LongAlign-10k LongAlign-10k 🤗 [LongAlign Dataset] • 💻 [Github Repo] • 📃 [LongAlign Paper] LongAlign is the first full recipe for LLM alignment on long context. We propose the LongAlign-10k dataset, containing 10,000 long instruction data of 8k-64k in length. We investigate on trianing strategies, namely packing (with loss weighting) and sorted batching, which are all implemented in our code. For real-world long context evaluation, we introduce LongBench-Chat that evaluate the… See the full description on the dataset page: https://huggingface.co/datasets/zai-org/LongAlign-10k.textquestion-answering1K<n<10K100 likes5.7k downloads3y agoHugging Face02Asklv /OpenMath-Vision-CoT-10kimage10K<n<100K1 likes5k downloads9mo agoHugging Face03timaeus /dsir-pile-10ktext10K<n<100K0 likes4.6k downloads2y agoHugging Face04PKU-Alignment /PKU-SafeRLHF-10K Paper You can find more information in our paper. Dataset Paper: https://arxiv.org/abs/2307.04657 tabulartext-generation10K<n<100K62 likes1.7k downloads3y agoHugging Face05smgjch /meow-10k Dataset Card for Meow-10K Meow-10K is a high-fidelity, synchronized quad-modal dataset comprising 10,000 feline samples. It is the primary training corpus for Meow-Omni 1, designed to facilitate deep intention reasoning in computational ethology. Dataset Summary Meow-10K provides the first large-scale training foundation for Multimodal Large Language Models (MLLMs) to learn the causal relationships between external behaviours and internal physiological states. By… See the full description on the dataset page: https://huggingface.co/datasets/smgjch/meow-10k.audio10K<n<100K3 likes1.5k downloads5mo agoHugging Face06TripVVT /TripVVT-10Kgated TripVVT-10K Dataset News 2026.06: TripVVT has been accepted by ECCV 2026. 2026.04: The TripVVT paper is available on arXiv. The project page is available at https://shaodingbao.github.io/TripVVT/. TripVVT-10K is a large-scale dataset for in-the-wild Video Virtual Try-On (VVT). It contains 10,031 high-quality video samples with triplet supervision, covering upper-body garments, lower-body garments, and dresses. TripVVT-10K is released together with the… See the full description on the dataset page: https://huggingface.co/datasets/TripVVT/TripVVT-10K.imageimage-to-video10K<n<100K6 likes1.1k downloads3mo agoHugging Face07Thunderbolt215215 /ArtiMuse-10K ArtiMuse: Fine-Grained Image Aesthetics Assessment with Joint Scoring and Expert-Level Understanding [🌐 Project Page] [🚀 Online Demo] [💻 Code] [📄 Paper] [[🧩 Checkpoints: 🤗 Hugging Face | 🤖 ModelScope]] 🌟 Building upon on ArtiMuse, we introduce UniPercept, a comprehensive follow-up work that provides a meticulous study on perceptual-level image understanding. It spans Image Aesthetics Assessment (IAA), Image Quality Assessment (IQA), and Image Structure & Texture… See the full description on the dataset page: https://huggingface.co/datasets/Thunderbolt215215/ArtiMuse-10K.imagequestion-answering1K<n<10K4 likes822 downloads8mo agoHugging Face08TeichAI /Ox-Alpha-10k Ox Alpha - 10k 10,005 single-turn prompts for text-response teacher generation Each row carries id, category, subcategory All data was gathered using stealth/ox-alpha via OpenRouter (reasoning effort high) Topic distribution Category Rows Share Coding (incl. Go/Rust, C++/Java/C#, shell/CLI) 944 9.5% Knowledge QA 891 9.0% Logical reasoning & decisions 734 7.4% Web development 720 7.2% Game development 720 7.2% Three.js / browser 3D 620 6.2%… See the full description on the dataset page: https://huggingface.co/datasets/TeichAI/Ox-Alpha-10k.texttext-generation1K<n<10K34 likes752 downloads1mo agoHugging Face09ScienceOne-AI /S1-Omni-Corpus-10K S1-Omni-Corpus-10K An open-source scientific multimodal reasoning dataset subset for S1-Omni 🧬 Model Introduction S1-Omni is a unified scientific multimodal reasoning model for scientific understanding, prediction, and generation. It is developed by the ScienceOne AI team of the Chinese Academy of Sciences. S1-Omni addresses fragmented scientific AI capabilities with a shared backbone for cross-disciplinary, cross-modal, and cross-task understanding and reasoning… See the full description on the dataset page: https://huggingface.co/datasets/ScienceOne-AI/S1-Omni-Corpus-10K.image10K<n<100K1 likes724 downloads2mo agoHugging Face10ScienceOne-AI /SciGenEdit-10K SciGenEdit-10K An Open Dataset for Scientific Image Generation and Editing English | 简体中文 📖 Introduction SciGenEdit-10K is a public subset released with the S1-Omni-Image project. It is designed for research on scientific image generation, scientific image editing, and multi-turn scientific image generation and editing. S1-Omni-Image is a unified multimodal model developed by the ScienceOne team at the Chinese Academy of Sciences for scientific… See the full description on the dataset page: https://huggingface.co/datasets/ScienceOne-AI/SciGenEdit-10K.imagetext-to-image10K<n<100K3 likes604 downloads3mo agoHugging Face11styletts2-community /multilingual-phonemes-10k-alpha Multilingual Phonemes 10K Alpha This dataset contains approximately 10,000 pairs of text and phonemes from each supported language. We support 15 languages in this dataset, so we have a total of ~150K pairs. This does not include the English-XL dataset, which includes another 100K unique rows. Languages We support 15 languages, which means we have around 150,000 pairs of text and phonemes in multiple languages. This excludes the English-XL dataset, which has 100K unique… See the full description on the dataset page: https://huggingface.co/datasets/styletts2-community/multilingual-phonemes-10k-alpha.text100K<n<1M39 likes594 downloads3y agoHugging Face12KerryMe /cinepile_10ktabular10K<n<100K0 likes520 downloads3mo agoHugging Face13lmarena-ai /webdev-arena-preference-10k WebDev Arena Preference Dataset This dataset contains 10K real-world Webdev Arena battle with 10 state-of-the-art LLMs. More details in the blog post. Dataset License Agreement This Agreement contains the terms and conditions that govern your access and use of the WebDev Arena Dataset (Arena Dataset). You may not use the Arena Dataset if you do not accept this Agreement. By clicking to accept, accessing the Arena Dataset, or both, you hereby agree to the terms of the… See the full description on the dataset page: https://huggingface.co/datasets/lmarena-ai/webdev-arena-preference-10k.text10K<n<100K20 likes401 downloads2y agoHugging Face14PhysGame /PhysDPO-10kTo use PhysDPO, run the following command to combine the parts into a single ZIP file: cat PhysDPO_part_* > PhysDPO.zip tabular10K<n<100K2 likes379 downloads2y agoHugging Face15zai-org /LongReward-10k LongReward-10k 💻 [Github Repo] • 📃 [LongReward Paper] LongReward-10k dataset contains 10,000 long-context QA instances (both English and Chinese, up to 64,000 words). The sft split contains SFT data generated by GLM-4-0520, following the self-instruct method in LongAlign. Using this split, we supervised fine-tune two models: LongReward-glm4-9b-SFT and LongReward-llama3.1-8b-SFT, which are based on GLM-4-9B and Meta-Llama-3.1-8B, respectively. The dpo_glm4_9b and… See the full description on the dataset page: https://huggingface.co/datasets/zai-org/LongReward-10k.texttext-generation10K<n<100K8 likes373 downloads2y agoHugging Face16Hakureirm /bird-species-10k Bird-10K SigLIP Training Dataset Taxonomy-aware image-text dataset for fine-tuning SigLIP on 10,753 bird species. Dataset Structure taxonomy.json # eBird v2025 taxonomy (species/genus/family/order hierarchy) hard_negatives.json # Same-genus & same-family negative sampling index siglip_train.jsonl # 692,236 training image-text pairs siglip_val.jsonl # 32,125 validation pairs images_224_*.tar.gz # 224x224 bird images organized by species… See the full description on the dataset page: https://huggingface.co/datasets/Hakureirm/bird-species-10k.text100K<n<1M0 likes349 downloads6mo agoHugging Face17qinglinhou /sokoban-10k-vjepa2-tokenizedtext10K<n<100K0 likes319 downloads6mo agoHugging Face18hkust-nlp /deita-10k-v0 Dataset Card for Deita 10K V0 GitHub | Paper Deita is an open-sourced project designed to facilitate Automatic Data Selection for instruction tuning in Large Language Models (LLMs). This dataset includes 10k of lightweight, high-quality alignment SFT data, mainly automatically selected from the following datasets: ShareGPT (Apache 2.0 listed, no official repo found): Use the 58 K ShareGPT dataset for selection. UltraChat (MIT): Sample 105 K UltraChat dataset for selection.… See the full description on the dataset page: https://huggingface.co/datasets/hkust-nlp/deita-10k-v0.text10K<n<100K30 likes313 downloads3y agoHugging Face19wangrongsheng /icliniq-10k-entext1K<n<10K4 likes243 downloads3y agoHugging Face20keplersystems /UrduShers-10ktext10K<n<100K1 likes228 downloads2y agoHugging Face21samhog /psychology-10ktext1K<n<10K54 likes187 downloads3y agoHugging Face22liumindmind /NekoQA-10K Dataset Card for NekoQA-10K 🐱 Dataset Summary NekoQA-10K 是一个面向大语言模型的 猫娘对话数据集,共包含 10,000 条 QA 对话。 所有回答均遵循统一的 猫娘人设: 称呼用户为“主人” 在句尾添加特定口癖(如“喵~”、“no desu”、“的说喵”) 保持可爱、撒娇、二次元风格 该数据集的主要用途是研究 大语言模型的“猫娘味”塑造能力,为微调、对话风格迁移、拟人化交互研究提供素材。 Supported Tasks and Benchmarks 风格微调 (Style Finetuning): 提升模型的“猫娘化”特征。 角色对话生成 (Persona-based Dialogue): 研究 LLM 的角色一致性建模。 情感陪伴研究 (Affective Computing): 探索模型与用户的情感交互能力。 评测基准 NekoBench: 可结合本数据集,评估模型在“猫娘味感知指数 (NPS)”上的表现。 Languages… See the full description on the dataset page: https://huggingface.co/datasets/liumindmind/NekoQA-10K.text10K<n<100K69 likes186 downloads1y agoHugging Face23byliutao /coco2014val_10k COCO-2014 Val 10K (256×256) A curated subset of the COCO 2014 Validation set containing 9,986 image-caption pairs, designed as a standard reference benchmark for FID (Fréchet Inception Distance) evaluation in text-to-image generation research. Dataset Summary Attribute Value Source COCO 2014 Validation Split Num Samples 9,986 Resolution 256 × 256 (center-cropped & resized) Image Format PNG, RGB Total Size ~1.2 GB Random Seed 42 Original Pool 40,504… See the full description on the dataset page: https://huggingface.co/datasets/byliutao/coco2014val_10k.imageimage-to-text1K<n<10K1 likes182 downloads5mo agoHugging Face24mishl /StreetVision-10K StreetVision-10K Each sample contains: A system prompt instructing the model to act as an OSINT/geospatial expert A user message with a street-level photo and the instruction to determine coordinates An assistant response with ground-truth coordinates in <direct_lon_lat_output>longitude,latitude</direct_lon_lat_output> format Format Each line is a JSON array of ChatML messages: [ {"role": "system", "content": "..."}, {"role": "user", "content": [… See the full description on the dataset page: https://huggingface.co/datasets/mishl/StreetVision-10K.imagevisual-question-answering10K<n<100K1 likes156 downloads6mo agoHugging Face25keenanpepper /whestbench-relu-mlp-moments-10k WhestBench Random ReLU MLPs — Monte-Carlo Activation Cumulants (10k) A dataset of 10,500 random ReLU MLPs (10,000 train + 500 held-out test) together with Monte-Carlo–estimated per-layer activation cumulants (mean, variance, skewness, kurtosis) for every layer, for both the pre-activation and post-ReLU signals. Built for the WhestBench Estimation Challenge 2026 and for research on analytic moment / uncertainty propagation through deep networks. The generative process… See the full description on the dataset page: https://huggingface.co/datasets/keenanpepper/whestbench-relu-mlp-moments-10k.tabularothern<1K0 likes156 downloads3mo agoHugging Face26wahyurejeki /dapurmu-conversational-commerce-10k 🍳 Dapurmu Conversational AI-Commerce (10.499 Dialogs) Dataset multi-turn conversational e-commerce bergaya pedagang pasar lokal Indonesia ("Kang Dapur") yang dilengkapi dengan fitur: Tawar-menawar produk segar (margin lebar) Penolakan sembako margin tipis dan pengalihan bundling Konsultasi menu resep masakan Nusantara Cek stok & kesegaran Pertahanan anti-jailbreak (floor price protection) Format: ChatML / OpenAI Tool Calling format. text-generation10K<n<100K0 likes154 downloads16d agoHugging Face27TigerResearch /tigerbot-zhihu-zh-10kTigerbot 基于开源搜集的知乎数据生成的sft问答对 Usage import datasets ds_sft = datasets.load_dataset('TigerResearch/tigerbot-zhihu-zh-10k') text10K<n<100K26 likes142 downloads3y agoHugging Face28SAMPOQ /supply-chain-10k 10-K Supply-Chain Graph Customer and supplier relationships that US-listed companies name in their own SEC 10-K filings, extracted into a directed graph. 16,172 rows across 2,659 filers, of which 3,568 resolve to a counterparty we also cover. Every row carries the accession number of the filing it came from, so any single claim can be checked against the original document on SEC EDGAR. This data is sold by financial data vendors and has no free equivalent we know of.… See the full description on the dataset page: https://huggingface.co/datasets/SAMPOQ/supply-chain-10k.text10K<n<100K1 likes141 downloads2d agoHugging Face29Mxode /Math-Chinese-DeepSeek-R1-10K 中文 DeepSeek-R1-Distil 数学指令微调数据集 💻 Github Repo 基本信息 数据集大小 10K,独立生成指令与回复,并非其他社区数据集的子集。所有数据经过校验,答案正确性可以得到保证。 数据集的组成如下: 问题类型 数据条数 定积分计算 2626 多项式化简 1621 因式分解 2557 多项式展开 2095 多项式方程 1101 总数 10000 数据格式 每条数据的格式如下: { "id": <<12位nanoid>>, "prompt": <<提示词>>, "reasoning": <<模型思考过程>>, "response": <<模型最终回复>> } texttext-generation10K<n<100K3 likes135 downloads1y agoHugging Face30mlx-community /Apertus-v1.5-QAT-10K mlx-community/Apertus-v1.5-QAT-10K This is a 2000 sample subset of the chosen pairs inside swiss-ai/Apertus_v1p5_Preference_Data for MLX-LM-LoRA and MLX-LoRA-Studio and the Quantization Aware Trained Appertus models. texttext-generation10K<n<100K1 likes134 downloads9d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.