CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01data-is-better-together /10k_prompts_ranked Dataset Card for 10k_prompts_ranked 10k_prompts_ranked is a dataset of prompts with quality rankings created by 314 members of the open-source ML community using Argilla, an open-source tool to label data. The prompts in this dataset include both synthetic and human-generated prompts sourced from a variety of heavily used datasets that include prompts. The dataset contains 10,331 examples and can be used for training and evaluating language models on prompt ranking tasks. The… See the full description on the dataset page: https://huggingface.co/datasets/data-is-better-together/10k_prompts_ranked.tabulartext-classification10K<n<100K170 likes2k downloads3y agoHugging Face02PKU-Alignment /PKU-SafeRLHF-10K Paper You can find more information in our paper. Dataset Paper: https://arxiv.org/abs/2307.04657 tabulartext-generation10K<n<100K62 likes1.7k downloads3y agoHugging Face03lfaviate /China-K12-STEM-10K-CoT-Reasoning K12-STEM-CoT-Chinese 1.54M Chinese K12 STEM problems with chain-of-thought solutions, 48% with diagrams. The largest structured Chinese math/physics/chemistry reasoning dataset. This is a curated sample (10,000 problems) of the full 1.54M dataset available via API. Full Dataset Access Access the full 1,540,000+ problems via API → This Sample Full API Total problems 10,025 1,540,000+ With CoT solutions 10,025 1,490,000+ With diagrams 6,093 740,000+… See the full description on the dataset page: https://huggingface.co/datasets/lfaviate/China-K12-STEM-10K-CoT-Reasoning.tabularquestion-answering10K<n<100K3 likes598 downloads7mo agoHugging Face04responsible-ai-labs /RAIL-HH-10K RAIL-HH-10K: Multi-Dimensional Safety Alignment Dataset The first large-scale safety dataset with 99.5% multi-dimensional annotation coverage across 8 ethical dimensions. 📖 Read Blog • 📖 Paper (Coming Soon) • 🚀 Quick Start • 🔌 RAIL API • 💻 Examples 🌟 What Makes RAIL-HH-10K Special? 🎯 Near-Complete Coverage 99.5% dimension coverage across all 8 ethical dimensions Most existing datasets: 40-70% coverage RAIL-HH-10K: 98-100%… See the full description on the dataset page: https://huggingface.co/datasets/responsible-ai-labs/RAIL-HH-10K.tabulartext-generation10K<n<100K6 likes287 downloads4mo agoHugging Face053nesdeniz /english-daily-dialogues-10k English Daily Dialogues 10K A general-purpose, open dataset of 10,000 synthetic multi-turn English conversations spanning ten everyday-life domains. Built as a clean NLP resource for dialogue modeling, response generation, intent understanding, and conversational evaluation. This is a general language resource — not a safety or security benchmark. Curated by Enes Deniz (ORCID 0009-0006-9491-3565), Co-Founder at AltaySec. It is the English companion to the Turkish Daily Dialogues… See the full description on the dataset page: https://huggingface.co/datasets/3nesdeniz/english-daily-dialogues-10k.tabulartext-generation10K<n<100K2 likes127 downloads2mo agoHugging Face06nraptisss /telecom-intent-config-sft-10k Telecom Intent→Config SFT Dataset (10K) The first open SFT dataset for training LLMs to translate natural language network intents into structured 5G/6G configurations. This dataset addresses the #1 gap identified in the telecom LLM research landscape: there is no public training dataset for intent-to-policy translation. All existing telecom datasets (TeleQnA, ORANBench-13K, 6G-Bench) are MCQ evaluation benchmarks — not instruction-following format. This dataset fills that gap.… See the full description on the dataset page: https://huggingface.co/datasets/nraptisss/telecom-intent-config-sft-10k.tabulartext-generation10K<n<100K1 likes100 downloads5mo agoHugging Face07alirezaaminzadeh /soc-agent-traces-10k SOC-Agent-Traces-10K Multi-step SOC investigation agent traces in session-trace format. Each record is a complete investigation session: an alert arrives, an analyst agent gathers evidence through nine read-only tools, and closes with a structured JSON triage report. Instead of single-turn alert → answer pairs, every record captures the full reasoning trajectory: alert → get_surrounding_events → get_process_tree → lookup_attack → search_sigma → get_asset_context →… See the full description on the dataset page: https://huggingface.co/datasets/alirezaaminzadeh/soc-agent-traces-10k.tabulartext-generation10K<n<100K0 likes100 downloads2mo agoHugging Face08iknow-lab /JudgeBias-DPO-RefFree-subset-10k JudgeBias-DPO-RefFree-subset-10k A 10K-pair subset of JudgeBias-DPO-RefFree-subset for training LLM judges to evaluate materials science synthesis recipes without bias in a reference-free setting (no ground truth recipe). Sampling Strategy: Stratified Dataset + Top Delta per Sample Equal quota per dataset: 9 datasets × ~1,111 pairs = 10,000 total Within each dataset: for each sample_id, pairs are ranked by score_delta (descending) and selected in round-robin order —… See the full description on the dataset page: https://huggingface.co/datasets/iknow-lab/JudgeBias-DPO-RefFree-subset-10k.tabulartext-generation10K<n<100K0 likes77 downloads7mo agoHugging Face09ozefe /ytc-10ktabularsummarization10K<n<100K0 likes66 downloads4mo agoHugging Face10JWei05 /Nemotron-Math-v2-Medium-10k Nemotron-Math-v2-Medium-10k A lightweight 10,500-problem subset of nvidia/Nemotron-Math-v2 for long-horizon Python-TIR reinforcement learning. It contains 1,500 problems from each metadata.reason_high_with_tool.pass bucket 1 through 7. A deterministic seed-42 shuffle assigns 500 examples to validation and 10,000 to train. This Hugging Face release intentionally contains no teacher traces. The full messages/tools aggregation is retained as a separate local artifact.… See the full description on the dataset page: https://huggingface.co/datasets/JWei05/Nemotron-Math-v2-Medium-10k.tabulartext-generation10K<n<100K0 likes60 downloads2mo agoHugging Face11a13905873166 /China-K12-STEM-10K-CoT-Reasoning K12-STEM-CoT-Chinese 1.54M Chinese K12 STEM problems with chain-of-thought solutions, 48% with diagrams. The largest structured Chinese math/physics/chemistry reasoning dataset. This is a curated sample (10,000 problems) of the full 1.54M dataset available via API. Full Dataset Access Access the full 1,540,000+ problems via API → This Sample Full API Total problems 10,025 1,540,000+ With CoT solutions 10,025 1,490,000+ With diagrams 6,093 740… See the full description on the dataset page: https://huggingface.co/datasets/a13905873166/China-K12-STEM-10K-CoT-Reasoning.tabularquestion-answering10K<n<100K1 likes54 downloads10d agoHugging Face12KavinduHansaka /prompt-gen-10k-flux-sdxl Prompt Generation Dataset (10K Narrative for Flux / SDXL) This dataset (prompt_gen_final_10k.jsonl and prompt_gen_final_10k.csv) was used to train and fine-tune image-prompt models such as KavinduHansaka/Llama-3.2-1B-ImageGen. It contains 10,000 curated narrative prompt samples designed for image generation models like Stable Diffusion XL and Flux.Unlike raw tag-based datasets, the target field provides natural paragraphs (≈80–100 words) that describe cinematic scenes with… See the full description on the dataset page: https://huggingface.co/datasets/KavinduHansaka/prompt-gen-10k-flux-sdxl.tabulartext-generation10K<n<100K0 likes43 downloads1y agoHugging Face13tomaarsen /zelo-scores-10kx100-granite-4.1-30b Dataset Card for tomaarsen/zelo-scores-10kx100-granite-4.1-30b Dataset Summary Synthetic data generated by DataForge: Model: ibm-granite/granite-4.1-30b (main) Source dataset: tomaarsen/zelo-pairs-10kx100-quantile-anchor (train split). Generation config: temperature=None, top_p=None, top_k=None, max_tokens=4096, model_max_context=32768 Speculative decoding: disabled System prompt: `You are a relevance scoring system. Given a query and two documents (A and B), your job… See the full description on the dataset page: https://huggingface.co/datasets/tomaarsen/zelo-scores-10kx100-granite-4.1-30b.tabulartext-generation1M<n<10M0 likes41 downloads5mo agoHugging Face14FierceLLM /ru-instruct-10k 10k Russian chatbot dialogues dataset tabulartext-generation1K<n<10K1 likes38 downloads6mo agoHugging Face15tampakwilll /fineweb2-id-filtered-10k FineWeb2-ID Filtered (min 10.000 karakter) Dataset ini adalah subset turunan dari HuggingFaceFW/fineweb-2, subset bahasa Indonesia (ind_Latn), yang disaring dengan syarat: Panjang teks (setelah strip whitespace) minimal 10.000 karakter per dokumen Diambil secara streaming dari split train, urutan asli (tanpa shuffle) Jumlah sampel: 2,097,152 dokumen, tersebar di 80 file JSONL. Sumber & Lisensi Seluruh isi teks berasal dari FineWeb-2 (lisensi ODC-By), yang pada… See the full description on the dataset page: https://huggingface.co/datasets/tampakwilll/fineweb2-id-filtered-10k.tabulartext-generation1M<n<10M1 likes36 downloads2d agoHugging Face16alirezaaminzadeh /retailopt-10k-verified RetailOpt-10K Verified Solver-verified expansion of Jacoblian/RetailOpt-190. Overview Property Value Base scenarios 190 Variations per base 53 Target total 10,070 Solver Pyomo + HiGHS Verification Optimal objective per scenario Variation Axes Each base scenario is perturbed along: demand_scale — scale demand curves (0.7–1.4×) capacity_scale — scale production and cold storage capacity (0.6–1.3×) cost_multiplier — scale… See the full description on the dataset page: https://huggingface.co/datasets/alirezaaminzadeh/retailopt-10k-verified.tabulartext-generation10K<n<100K0 likes35 downloads2mo agoHugging Face17BEE-spoke-data /sp500-edgar-10k-markdowngated edgar s&p500 Source Datasets The source dataset used for this report is jlohding/sp500-edgar-10k. Dataset Information Configuration: default Feature Data Type cik string sic string company string date timestamp[us] ret float64 mkt_cap float64 report_intro string text string report_returns string word_count int64 Splits: Train: Number of Examples: 6258 Size: 2260000389 bytes Download Size: 974801155 bytesDataset… See the full description on the dataset page: https://huggingface.co/datasets/BEE-spoke-data/sp500-edgar-10k-markdown.tabulartext-generation10K<n<100K6 likes33 downloads9mo agoHugging Face18hybrid-diff-ar /stack-v2-sparse-classes-10k Stack v2 Sparse Python Classes 10k This is a 10,000-sample snapshot for Diffusion + Autoregressive hybrid code generation experiments. Source The data is extracted from bigcode/the-stack-v2-dedup, Python subset. The extraction uses Stack v2 metadata as source of truth, groups candidates by repo_name + revision_id, fetches files with git partial fetch + sparse checkout, then applies AST-level class filters. Splits train.jsonl: 9,000 val.jsonl: 500 test.jsonl:… See the full description on the dataset page: https://huggingface.co/datasets/hybrid-diff-ar/stack-v2-sparse-classes-10k.tabulartext-generation10K<n<100K0 likes31 downloads5mo agoHugging Face19alexshengzhili /Abstract2Appendix_v1_10k Dataset Card: Abstract2Appendix v1 Dataset Description The Abstract2Appendix v1 dataset is a high-quality collection of academic peer reviews and their associated research paper metadata. This dataset combines reviews from four premier machine learning and AI conferences: NeurIPS 2023, EMNLP 2023, TMLR, and ICLR 2023, shuffled into a unified corpus. It is designed to enhance long-context capabilities in Large Language Models (LLMs) and supports tasks such as fine-tuning… See the full description on the dataset page: https://huggingface.co/datasets/alexshengzhili/Abstract2Appendix_v1_10k.tabulartext-generation1K<n<10K3 likes29 downloads2y agoHugging Face20CL-From-Nothing /code_rose_initial_1_7B_SFT_10K_rollouts_Qwen3-4B-Thinking-2507_k12_t0.7_maxtok12288 code_rose_initial_1_7B_SFT_10K — rollouts (Qwen3-4B-Thinking-2507, k=12) Pass@k completions generated with vLLM over the prefixes in CL-From-Nothing/code_rose_initial_1_7B_SFT_10K. Generation config Model Qwen3-4B-Thinking-2507 Samples per question (k) 12 Temperature 0.7 top_p 0.9 max_tokens 12288 max_model_len 32768 Questions 7250 (index 0–7249, full split) Total rows 87000 (7250 × 12) Generated by complete_prefix_vllm.py… See the full description on the dataset page: https://huggingface.co/datasets/CL-From-Nothing/code_rose_initial_1_7B_SFT_10K_rollouts_Qwen3-4B-Thinking-2507_k12_t0.7_maxtok12288.tabulartext-generation10K<n<100K0 likes23 downloads3mo agoHugging Face21JWei05 /Nemotron-Math-v2-10k Nemotron-Math-v2-10k 10,000 machine-verifiable, mid-to-hard math problems subsampled from nvidia/Nemotron-Math-v2, each with all stored gpt-oss-120b high-reasoning, no-tool solution traces. Built for RL training (verifiable final answers) with a teacher-trace column for SFT distillation / analysis. How it was made (seed 42 throughout) Dedupe: scanned all 5 splits (high_part00-02, medium, low), deduped rows by md5(problem) → 324,233 unique problems.… See the full description on the dataset page: https://huggingface.co/datasets/JWei05/Nemotron-Math-v2-10k.tabulartext-generation10K<n<100K0 likes23 downloads2mo agoHugging Face22OpenCoven /fable-forge-10k FableForge — Narrative Reasoning Dataset with Recurrence-Depth Annotations The first narrative dataset designed around recurrence depth requirements. Every example carries a suggested_n_loops field with a theoretically grounded basis — derived from the structural complexity of the task, not a heuristic label or emergent property. Background Standard narrative datasets treat reasoning depth as an emergent property. FableForge is different: it annotates how much… See the full description on the dataset page: https://huggingface.co/datasets/OpenCoven/fable-forge-10k.tabulartext-generation10K<n<100K0 likes22 downloads3mo agoHugging Face23ermiaazarkhalili /alpaca-high-prob-qwen-0.5b-10k High-Probability Sentence Predictions Dataset Dataset Description This dataset contains sentences from tatsu-lab/alpaca where the model Qwen/Qwen2.5-0.5B predicts the token before the final period with ≥90% probability. Source Dataset Attribution This dataset is derived from tatsu-lab/alpaca and inherits its license terms (cc-by-nc-4.0). Please cite the original dataset when using this data. Extraction Parameters Parameter Value Source… See the full description on the dataset page: https://huggingface.co/datasets/ermiaazarkhalili/alpaca-high-prob-qwen-0.5b-10k.tabulartext-generation10K<n<100K0 likes21 downloads9mo agoHugging Face24transhumanist-already-exists /pretraining-high-quality-10k-workshop Lapa HQ 10k Workshop Corpus A small deterministic subset of lapa-llm/pretraining-high-quality for tokenizer-transfer workshop runs. Provenance Source dataset: lapa-llm/pretraining-high-quality Source config: default Source split: train Rows: 10000 Selection: first 10000 rows by dataset-server row order Download window size: 100 Parallel workers: 20 Created at UTC: 2026-06-20T09:22:21.460168+00:00 Added columns: source_row_idx mini_corpus_index tabulartext-generation10K<n<100K0 likes21 downloads3mo agoHugging Face25ckerf /arb-raw-10k Arb-Agent Raw Data This is the main text corpus used to train the QuantOxide Reasoning Agent. It contains clean, semantic text chunks extracted from the 10-K filings of the top 50 S&P 500 companies. The Parsing Logic Parsing SEC filings is unbelievably difficult due to inconsistent HTML, broken table tags, and "incorporation by reference." This dataset was created using a unique parsing technique: Instead of relying on broken regex headers, the parser scores chunks based… See the full description on the dataset page: https://huggingface.co/datasets/ckerf/arb-raw-10k.tabulartext-generation1K<n<10K1 likes18 downloads10mo agoHugging Face26JingweiNi /ocr2_cf1900_k2_qwen35_gpt55_aggregated_10k_seed20260513 Aggregated Qwen3.5 and GPT-5.5 OCR2 Coding Step Labels This dataset follows the same 500-row parquet layout as JingweiNi/ocr2_cf1900_k2_qwen35_fp8_10k_seed20260513 and aggregates the original Qwen3.5 labels with GPT-5.5 medium reannotation of Qwen3.5-positive error steps. Aggregation rule: GPT-5.5 confirmed errors are errors; all other 10,000 manifest-selected steps are correct. Columns qwen35_verified: original Qwen3.5 labels for the 10,000 manifest-selected steps, NaN… See the full description on the dataset page: https://huggingface.co/datasets/JingweiNi/ocr2_cf1900_k2_qwen35_gpt55_aggregated_10k_seed20260513.tabulartext-generationn<1K0 likes16 downloads4mo agoHugging Face27directorybear /open-library-10k Dir Bear Open Library — 10,000 distilled web documents Ten thousand complete, cleaned, English web documents — every one at least 250 words of prose, boilerplate stripped, exact-deduplicated, token-counted and scored by the quality of the site it came from. This is the free, open slice of the Dir Bear corpus: the same records, the same schema and the same pipeline as the datasets we sell, at a size you can read through in an afternoon. Browse it online:… See the full description on the dataset page: https://huggingface.co/datasets/directorybear/open-library-10k.tabulartext-generation10K<n<100K0 likes15 downloads2d agoHugging Face28ermiaazarkhalili /alpaca-cleaned-high-prob-qwen-0.5b-10k High-Probability Sentence Predictions Dataset Dataset Description This dataset contains sentences from yahma/alpaca-cleaned where the model Qwen/Qwen2.5-0.5B predicts the token before the final period with ≥90% probability. Source Dataset Attribution This dataset is derived from yahma/alpaca-cleaned and inherits its license terms (cc-by-4.0). Please cite the original dataset when using this data. Extraction Parameters Parameter Value… See the full description on the dataset page: https://huggingface.co/datasets/ermiaazarkhalili/alpaca-cleaned-high-prob-qwen-0.5b-10k.tabulartext-generation10K<n<100K0 likes11 downloads9mo agoHugging Face29ermiaazarkhalili /alpaca-gpt4-en-high-prob-qwen-0.5b-10k High-Probability Sentence Predictions Dataset Dataset Description This dataset contains sentences from llamafactory/alpaca_gpt4_en where the model Qwen/Qwen2.5-0.5B predicts the token before the final period with ≥90% probability. Source Dataset Attribution This dataset is derived from llamafactory/alpaca_gpt4_en and inherits its license terms (apache-2.0). Please cite the original dataset when using this data. Extraction Parameters Parameter… See the full description on the dataset page: https://huggingface.co/datasets/ermiaazarkhalili/alpaca-gpt4-en-high-prob-qwen-0.5b-10k.tabulartext-generation10K<n<100K0 likes11 downloads9mo agoHugging Face30ermiaazarkhalili /orca-mini-v1-high-prob-qwen-0.5b-10k High-Probability Sentence Predictions Dataset Dataset Description This dataset contains sentences from pankajmathur/orca_mini_v1_dataset where the model Qwen/Qwen2.5-0.5B predicts the token before the final period with ≥90% probability. Source Dataset Attribution This dataset is derived from pankajmathur/orca_mini_v1_dataset and inherits its license terms (apache-2.0). Please cite the original dataset when using this data. Extraction Parameters… See the full description on the dataset page: https://huggingface.co/datasets/ermiaazarkhalili/orca-mini-v1-high-prob-qwen-0.5b-10k.tabulartext-generation10K<n<100K0 likes11 downloads9mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.