CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01artefactory /Argimi-Ardian-Finance-10k-text The ArGiMI Ardian datasets : Text-only version The ArGiMi project is committed to open-source principles and data sharing. Thanks to our generous partners, we are releasing several valuable datasets to the public. Dataset description This text-only dataset comprises 34,000 financial annual reports, written in English, meticulously extracted from their original PDF format to provide a valuable resource for researchers and developers in financial analysis and natural… See the full description on the dataset page: https://huggingface.co/datasets/artefactory/Argimi-Ardian-Finance-10k-text.texttext-retrieval1M<n<10M19 likes5.7k downloads7mo agoHugging Face02Monster-Code /Pytorch-Code-10K Hot Coco Training Dataset A curated collection of 10,625 high-quality PyTorch and Transformers code examples with AI-generated captions. This dataset was specifically built for fine-tuning code-specialized language models like Qimi (Coming soon!) Dataset Description This dataset contains Python code snippets sourced from open-source repositories that utilize PyTorch or Hugging Face Transformers. Each sample includes: code: The raw Python source code (typically… See the full description on the dataset page: https://huggingface.co/datasets/Monster-Code/Pytorch-Code-10K.texttext-generation1K<n<10K1 likes3.9k downloads2mo agoHugging Face03data-is-better-together /10k_prompts_ranked Dataset Card for 10k_prompts_ranked 10k_prompts_ranked is a dataset of prompts with quality rankings created by 314 members of the open-source ML community using Argilla, an open-source tool to label data. The prompts in this dataset include both synthetic and human-generated prompts sourced from a variety of heavily used datasets that include prompts. The dataset contains 10,331 examples and can be used for training and evaluating language models on prompt ranking tasks. The… See the full description on the dataset page: https://huggingface.co/datasets/data-is-better-together/10k_prompts_ranked.tabulartext-classification10K<n<100K170 likes2k downloads3y agoHugging Face04astr010 /sec-10k-markdown-uncompressed 📄 SEC 10-K Full Uncompressed Markdown Filings (12.3k Documents) Dataset Summary This dataset contains 12,361 full-length, uncompressed SEC Form 10-K annual reports converted from EDGAR HTML to clean Markdown format across 1,379 companies (spanning 2004 to 2025, core 2014–2025). The dataset is organized as uncompressed Markdown files structured by company ticker subdirectories (AAPL/10-K_2024.md, NVDA/10-K_2024.md, etc.), complete with company metadata manifests… See the full description on the dataset page: https://huggingface.co/datasets/astr010/sec-10k-markdown-uncompressed.texttext-generation1K<n<10K0 likes1.8k downloads2mo agoHugging Face05PKU-Alignment /PKU-SafeRLHF-10K Paper You can find more information in our paper. Dataset Paper: https://arxiv.org/abs/2307.04657 tabulartext-generation10K<n<100K62 likes1.7k downloads3y agoHugging Face06TeichAI /Ox-Alpha-10k Ox Alpha - 10k 10,005 single-turn prompts for text-response teacher generation Each row carries id, category, subcategory All data was gathered using stealth/ox-alpha via OpenRouter (reasoning effort high) Topic distribution Category Rows Share Coding (incl. Go/Rust, C++/Java/C#, shell/CLI) 944 9.5% Knowledge QA 891 9.0% Logical reasoning & decisions 734 7.4% Web development 720 7.2% Game development 720 7.2% Three.js / browser 3D 620 6.2%… See the full description on the dataset page: https://huggingface.co/datasets/TeichAI/Ox-Alpha-10k.texttext-generation1K<n<10K34 likes800 downloads1mo agoHugging Face07lfaviate /China-K12-STEM-10K-CoT-Reasoning K12-STEM-CoT-Chinese 1.54M Chinese K12 STEM problems with chain-of-thought solutions, 48% with diagrams. The largest structured Chinese math/physics/chemistry reasoning dataset. This is a curated sample (10,000 problems) of the full 1.54M dataset available via API. Full Dataset Access Access the full 1,540,000+ problems via API → This Sample Full API Total problems 10,025 1,540,000+ With CoT solutions 10,025 1,490,000+ With diagrams 6,093 740,000+… See the full description on the dataset page: https://huggingface.co/datasets/lfaviate/China-K12-STEM-10K-CoT-Reasoning.tabularquestion-answering10K<n<100K3 likes598 downloads7mo agoHugging Face08artefactory /Argimi-Ardian-Finance-10k-text-image The ArGiMI Ardian datasets : text and images The ArGiMi project is committed to open-source principles and data sharing. Thanks to our generous partners, we are releasing several valuable datasets to the public. Dataset description This dataset comprises 34,000 financial annual reports, written in English, meticulously extracted from their original PDF format to provide a valuable resource for researchers and developers in financial analysis and natural language… See the full description on the dataset page: https://huggingface.co/datasets/artefactory/Argimi-Ardian-Finance-10k-text-image.imagetext-retrieval1M<n<10M14 likes535 downloads7mo agoHugging Face09bulatSharif /gh-pr-issue-traces-10k GitHub PR Review Traces 10K Dataset Summary This dataset contains 10,000 raw public GitHub pull request traces. Each JSONL row represents one pull request candidate and joins together PR metadata, discussion, code review, changed files, full diff text, optional auxiliary fields, and retrieval provenance. The data is designed for mining software engineering workflows of the form: Pull Request -> Discussion -> Review -> Code Diff -> Merge The uploaded file is:… See the full description on the dataset page: https://huggingface.co/datasets/bulatSharif/gh-pr-issue-traces-10k.text-classification10K<n<100K0 likes533 downloads5mo agoHugging Face10jumplander /JL-Agentic-Behavior-10K JumpLander Agentic Behavior 10K Version 1.0.0 · 10,000 English canonical scenarios · 10 capability datasets JumpLander Agentic Behavior 10K is a process-grounded synthetic suite for training and evaluating agents that plan, use tools, recover from failure, manage memory, verify outcomes, respect authority, coordinate with people and other agents, and control long-horizon tasks. Release status The package is complete as a v1 synthetic candidate release. It… See the full description on the dataset page: https://huggingface.co/datasets/jumplander/JL-Agentic-Behavior-10K.text-generation10K<n<100K6 likes453 downloads2mo agoHugging Face11zai-org /LongReward-10k LongReward-10k 💻 [Github Repo] • 📃 [LongReward Paper] LongReward-10k dataset contains 10,000 long-context QA instances (both English and Chinese, up to 64,000 words). The sft split contains SFT data generated by GLM-4-0520, following the self-instruct method in LongAlign. Using this split, we supervised fine-tune two models: LongReward-glm4-9b-SFT and LongReward-llama3.1-8b-SFT, which are based on GLM-4-9B and Meta-Llama-3.1-8B, respectively. The dpo_glm4_9b and… See the full description on the dataset page: https://huggingface.co/datasets/zai-org/LongReward-10k.texttext-generation10K<n<100K8 likes367 downloads2y agoHugging Face12OysterCoreAI /SFT-General-Simplified.Chinese-10K SFT-General-Simplified.Chinese-10K 概况介绍 Hugging Face仓库名称:OysterCoreAI/SFT-General-Simplified.Chinese-10K SFT-General-Simplified.Chinese-10K 是一个面向简体中文监督微调(SFT)研究的指令—回答数据集,针对数据源以及数据质量做出了严格把控和筛选(通过信息密度、上下文长度等共计12维度筛选),共 10,050 条记录。内容从中文维基百科开放转储的一手条目中提取并进行规则化任务构造;没有使用 Belle、MOSS-SFT、Alpaca-Chinese 或其他既有 SFT/指令数据集作为数据源。 重要边界: 本数据集通过自动清洗、三层次筛选、定向数据收集、双逻辑 18 维评分和独立发布审计。自动分数是筛选指标,并非事实正确率、法律意见或模型效果保证。高风险场景使用前请另行人工复核。 训练集数据概览 项目 结果 记录数 10,050 语言 中文(简体规范化目标)… See the full description on the dataset page: https://huggingface.co/datasets/OysterCoreAI/SFT-General-Simplified.Chinese-10K.text-generation10K<n<100K0 likes341 downloads2d agoHugging Face13astr010 /sec-10k-markdown-compressed 📦 SEC 10-K Full Markdown Filings (Compressed Archives) Dataset Summary This dataset contains 12,361 full-length SEC Form 10-K annual reports converted from EDGAR HTML to clean Markdown format across 1,379 companies (spanning 2004 to 2025, core 2014–2025). The filings are packaged as 1,379 compressed .tar.gz archives (1 archive per company ticker). Each archive contains the un-chunked annual 10-K Markdown reports and a JSON metadata manifest (manifest.json).… See the full description on the dataset page: https://huggingface.co/datasets/astr010/sec-10k-markdown-compressed.text-generation10M<n<100M0 likes311 downloads2mo agoHugging Face14JWei05 /DeepScaleR-Easy-Medium-Hard-Gemma-26B-PT-10k DeepScaleR Easy/Medium/Hard — Gemma 4 26B-A4B PT This dataset contains 9,900 unique, deduplicated DeepScaleR math questions for reinforcement-learning experiments. Difficulty is defined by how often the pretrained google/gemma-4-26B-A4B teacher solved each question across eight temperature-1 samples under the same rule-based grader used by the RL training pipeline. The Hub dataset has three configurations—easy, medium, and hard—and each configuration has a train split with 3,000… See the full description on the dataset page: https://huggingface.co/datasets/JWei05/DeepScaleR-Easy-Medium-Hard-Gemma-26B-PT-10k.texttext-generation1K<n<10K0 likes304 downloads1mo agoHugging Face15responsible-ai-labs /RAIL-HH-10K RAIL-HH-10K: Multi-Dimensional Safety Alignment Dataset The first large-scale safety dataset with 99.5% multi-dimensional annotation coverage across 8 ethical dimensions. 📖 Read Blog • 📖 Paper (Coming Soon) • 🚀 Quick Start • 🔌 RAIL API • 💻 Examples 🌟 What Makes RAIL-HH-10K Special? 🎯 Near-Complete Coverage 99.5% dimension coverage across all 8 ethical dimensions Most existing datasets: 40-70% coverage RAIL-HH-10K: 98-100%… See the full description on the dataset page: https://huggingface.co/datasets/responsible-ai-labs/RAIL-HH-10K.tabulartext-generation10K<n<100K6 likes287 downloads4mo agoHugging Face16openeurollm /nemotron-cc-10K-sample-translated Translated Nemotron-cc-hq samples This dataset contains translated samples from https://huggingface.co/datasets/spyysalo/nemotron-cc-10K-sample Currently, the following are available, we will add other models and languages: Model Languages Gemma-3-4b-it ["Bulgarian", "Czech", "Danish", "German", "Estonian", "Finnish", "French", "Croatian", "Dutch"] EuroLLM-9B-Instruct ["Bulgarian", "Czech", "Danish", "German", "Greek", "Estonia", "Finnish", "French", "Irish"… See the full description on the dataset page: https://huggingface.co/datasets/openeurollm/nemotron-cc-10K-sample-translated.texttext-generation100K<n<1M1 likes174 downloads1y agoHugging Face17dougalldeepmind /2026-08-06-qwen36-table2-80-self-reflection-20-10k-train-mixture Qwen3.6 Table2 80% + SynthDoc self-reflection 20% — 10k-example training bundle field value experiment One-epoch Qwen3.6-27B assistant-only LoRA SFT (r64): Matthew's exact 7,999 Table-2 rows + 2,000 first-person self-reflection records — the self-reflection twin of LASR-Callum/2026-08-04-qwen36-lora-table2-synthdoc-rank-64, differing ONLY in the 20% slice (difficult-advice -> self-reflection). date_generated 2026-08-06 (mixture; Table-2 rows verbatim from the… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-06-qwen36-table2-80-self-reflection-20-10k-train-mixture.text-generation0 likes172 downloads25d agoHugging Face18mishl /StreetVision-10K StreetVision-10K Each sample contains: A system prompt instructing the model to act as an OSINT/geospatial expert A user message with a street-level photo and the instruction to determine coordinates An assistant response with ground-truth coordinates in <direct_lon_lat_output>longitude,latitude</direct_lon_lat_output> format Format Each line is a JSON array of ChatML messages: [ {"role": "system", "content": "..."}, {"role": "user", "content": [… See the full description on the dataset page: https://huggingface.co/datasets/mishl/StreetVision-10K.imagevisual-question-answering10K<n<100K1 likes159 downloads6mo agoHugging Face19wahyurejeki /dapurmu-conversational-commerce-10k 🍳 Dapurmu Conversational AI-Commerce (10.499 Dialogs) Dataset multi-turn conversational e-commerce bergaya pedagang pasar lokal Indonesia ("Kang Dapur") yang dilengkapi dengan fitur: Tawar-menawar produk segar (margin lebar) Penolakan sembako margin tipis dan pengalihan bundling Konsultasi menu resep masakan Nusantara Cek stok & kesegaran Pertahanan anti-jailbreak (floor price protection) Format: ChatML / OpenAI Tool Calling format. text-generation10K<n<100K0 likes152 downloads16d agoHugging Face20Mxode /Math-Chinese-DeepSeek-R1-10K 中文 DeepSeek-R1-Distil 数学指令微调数据集 💻 Github Repo 基本信息 数据集大小 10K,独立生成指令与回复,并非其他社区数据集的子集。所有数据经过校验,答案正确性可以得到保证。 数据集的组成如下: 问题类型 数据条数 定积分计算 2626 多项式化简 1621 因式分解 2557 多项式展开 2095 多项式方程 1101 总数 10000 数据格式 每条数据的格式如下: { "id": <<12位nanoid>>, "prompt": <<提示词>>, "reasoning": <<模型思考过程>>, "response": <<模型最终回复>> } texttext-generation10K<n<100K3 likes144 downloads1y agoHugging Face21UmaiTech /legal-contract-gpt41-redlining-10k legal-contract-gpt41-redlining-10k Dataset Description This dataset contains 9977 synthetic legal contract redlines generated using GPT-4.1 model mix (base, mini, nano) with structured outputs. It is designed for fine-tuning LLMs (including OpenAI GPT-3.5/4, Llama, and other models) to assist with legal document redlining and clause revision. Key Features 🤖 9977 synthetic redlines generated by GPT-4.1 model mix (base, mini, nano) 📋 Multiple training formats:… See the full description on the dataset page: https://huggingface.co/datasets/UmaiTech/legal-contract-gpt41-redlining-10k.texttext-generation10K<n<100K1 likes132 downloads11mo agoHugging Face22mlx-community /Apertus-v1.5-QAT-10K mlx-community/Apertus-v1.5-QAT-10K This is a 2000 sample subset of the chosen pairs inside swiss-ai/Apertus_v1p5_Preference_Data for MLX-LM-LoRA and MLX-LoRA-Studio and the Quantization Aware Trained Appertus models. texttext-generation10K<n<100K1 likes132 downloads9d agoHugging Face23yxx123456 /sft-safe-openai-chat-10k SFT Safe OpenAI Chat 10K This dataset is formatted for chat SFT training. Each JSONL row contains a messages field compatible with OpenAI-style chat fine-tuning data: {"messages":[{"role":"system","content":"..."},{"role":"user","content":"..."},{"role":"assistant","content":"..."}]} Files: train.jsonl: 10,000 training examples validation.jsonl: 200 validation examples eval.jsonl: same content as validation.jsonl, provided as an evaluation alias Example usage: fromdatasets import… See the full description on the dataset page: https://huggingface.co/datasets/yxx123456/sft-safe-openai-chat-10k.texttext-generation10K<n<100K0 likes130 downloads4mo agoHugging Face24Multilingual-Multimodal-NLP /WebWorld-SFT-10k WebWorld-SFT-10k The Browser as a World Model for Self-Improving Web Code — a 10,000-example SFT release from the WebWorld training set. Each example is a single-turn improvement of an HTML artifact that was accepted by the browser-issued certificate: the browser re-executed the candidate, target progress held, and every previously verified capability was preserved. Rejected trajectories are not in the export. Why WebWorld VLM-driven self-improvement of web code… See the full description on the dataset page: https://huggingface.co/datasets/Multilingual-Multimodal-NLP/WebWorld-SFT-10k.text-generation10K<n<100K0 likes129 downloads25d agoHugging Face25Aipresso /10k_rows_cleaned_prompts 10K Rows Cleaned Prompts Dataset Created by Aipresso LIMITED, London, UK ⚠️ IMPORTANT: By using this dataset, you agree to our Terms of Use You must provide attribution when using this data in publications, research, or commercial products. Dataset Overview A chunked collection of 2.7 million cleaned English prompts, organized into 200 files of 10,000 rows each for easy processing and distributed training of language models. 📊 Dataset Statistics Metric… See the full description on the dataset page: https://huggingface.co/datasets/Aipresso/10k_rows_cleaned_prompts.texttext-generation1M<n<10M0 likes128 downloads11mo agoHugging Face263nesdeniz /english-daily-dialogues-10k English Daily Dialogues 10K A general-purpose, open dataset of 10,000 synthetic multi-turn English conversations spanning ten everyday-life domains. Built as a clean NLP resource for dialogue modeling, response generation, intent understanding, and conversational evaluation. This is a general language resource — not a safety or security benchmark. Curated by Enes Deniz (ORCID 0009-0006-9491-3565), Co-Founder at AltaySec. It is the English companion to the Turkish Daily Dialogues… See the full description on the dataset page: https://huggingface.co/datasets/3nesdeniz/english-daily-dialogues-10k.tabulartext-generation10K<n<100K2 likes127 downloads2mo agoHugging Face27twinkle-ai /tw-function-call-reasoning-10k Dataset Card for tw-function-call-reasoning-10k 本資料集為繁體中文版本的函式呼叫(Function Calling)資料集,翻譯自 AymanTarig/function-calling-v0.2-with-r1-cot,而該資料集本身是 Salesforce/xlam-function-calling-60k 的修正版。我們利用語言模型翻譯後,經人工修改,旨在打造高品質的繁體中文工具使用語料。 Dataset Details Dataset Description tw-function-call-reasoning-10k 是一個專為語言模型「工具使用能力(Function Calling)」訓練所設計的繁體中文資料集。其內容源自 AymanTarig/function-calling-v0.2-with-r1-cot,該資料集又為 Salesforce/xlam-function-calling-60k… See the full description on the dataset page: https://huggingface.co/datasets/twinkle-ai/tw-function-call-reasoning-10k.texttext-generation10K<n<100K14 likes113 downloads1y agoHugging Face28ProCreations /grug-think-v3-10k grug-think-v3-10k v2 brain short. v2 brain useful. but some v2 brain wear office shirt. "User wants hello world Python. Provide code." short English, yes. grug, no. v3 tear off office shirt. keep brain meat. old: User wants hello world Python. Simple code snippet, no tools needed. Provide code and brief explanation. new: Need Python hello-world. Tiny snippet. No tool. Give code, brief explain. complex cave different. grug no crush branch into pebble. exact path, error… See the full description on the dataset page: https://huggingface.co/datasets/ProCreations/grug-think-v3-10k.texttext-generation10K<n<100K12 likes113 downloads2mo agoHugging Face29ai4privacy /pii-masking-mini-10k PII Masking Mini: Multilingual Sample A mini-sized stratified sample of pii-masking-openpii-1.5m, the flagship release of the PII-Masking-3M family. Sampled proportionally by (source_dataset, language) so every locale and label gets representation. Asia Pacific rows appear first. 📖 More information: www.ai4privacy.com/datasets/pii-masking-3m-asia-pacific Dataset Details Total Examples Train Validation Labels Languages Regions Annotations Format… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/pii-masking-mini-10k.texttoken-classification1K<n<10K0 likes110 downloads4mo agoHugging Face30tomaarsen /zelo-scores-10kx100-qwen3.6-27b Dataset Card for tomaarsen/zelo-scores-10kx100-qwen3.6-27b Dataset Summary Synthetic data generated by DataForge: Model: Qwen/Qwen3.6-27B (main) Source dataset: tomaarsen/zelo-pairs-10kx100-quantile-anchor (train split). Generation config: temperature=1.0, top_p=0.95, top_k=20, max_tokens=4096, model_max_context=32768 Speculative decoding: disabled System prompt: `You are a relevance scoring system. Given a query and two documents (A and B), your job is to decide which… See the full description on the dataset page: https://huggingface.co/datasets/tomaarsen/zelo-scores-10kx100-qwen3.6-27b.text-generation0 likes106 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.