CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01PKU-Alignment /PKU-SafeRLHF-10K Paper You can find more information in our paper. Dataset Paper: https://arxiv.org/abs/2307.04657 tabulartext-generation10K<n<100K62 likes1.7k downloads3y agoHugging Face02TripVVT /TripVVT-10Kgated TripVVT-10K Dataset News 2026.06: TripVVT has been accepted by ECCV 2026. 2026.04: The TripVVT paper is available on arXiv. The project page is available at https://shaodingbao.github.io/TripVVT/. TripVVT-10K is a large-scale dataset for in-the-wild Video Virtual Try-On (VVT). It contains 10,031 high-quality video samples with triplet supervision, covering upper-body garments, lower-body garments, and dresses. TripVVT-10K is released together with the… See the full description on the dataset page: https://huggingface.co/datasets/TripVVT/TripVVT-10K.imageimage-to-video10K<n<100K6 likes1.1k downloads3mo agoHugging Face03KerryMe /cinepile_10ktabular10K<n<100K0 likes482 downloads3mo agoHugging Face04PhysGame /PhysDPO-10kTo use PhysDPO, run the following command to combine the parts into a single ZIP file: cat PhysDPO_part_* > PhysDPO.zip tabular10K<n<100K2 likes375 downloads2y agoHugging Face05keenanpepper /whestbench-relu-mlp-moments-10k WhestBench Random ReLU MLPs — Monte-Carlo Activation Cumulants (10k) A dataset of 10,500 random ReLU MLPs (10,000 train + 500 held-out test) together with Monte-Carlo–estimated per-layer activation cumulants (mean, variance, skewness, kurtosis) for every layer, for both the pre-activation and post-ReLU signals. Built for the WhestBench Estimation Challenge 2026 and for research on analytic moment / uncertainty propagation through deep networks. The generative process… See the full description on the dataset page: https://huggingface.co/datasets/keenanpepper/whestbench-relu-mlp-moments-10k.tabularothern<1K0 likes159 downloads3mo agoHugging Face06beaugogh /openorca-multiplechoice-10kA 10k subset of OpenOrca dataset, focusing on multiple choice questions. Credit to Tian Xia. tabular10K<n<100K5 likes105 downloads3y agoHugging Face07nateraw /filings-10ktabular10K<n<100K0 likes59 downloads5y agoHugging Face08KavinduHansaka /prompt-gen-10k-flux-sdxl Prompt Generation Dataset (10K Narrative for Flux / SDXL) This dataset (prompt_gen_final_10k.jsonl and prompt_gen_final_10k.csv) was used to train and fine-tune image-prompt models such as KavinduHansaka/Llama-3.2-1B-ImageGen. It contains 10,000 curated narrative prompt samples designed for image generation models like Stable Diffusion XL and Flux.Unlike raw tag-based datasets, the target field provides natural paragraphs (≈80–100 words) that describe cinematic scenes with… See the full description on the dataset page: https://huggingface.co/datasets/KavinduHansaka/prompt-gen-10k-flux-sdxl.tabulartext-generation10K<n<100K0 likes49 downloads1y agoHugging Face09zhiqings /LLaVA-Human-Preference-10Ktabular1K<n<10K34 likes43 downloads3y agoHugging Face10open-llm-leaderboard /asharsha30__LLAMA_Harsha_8_B_ORDP_10k-detailsgated Dataset Card for Evaluation run of asharsha30/LLAMA_Harsha_8_B_ORDP_10k Dataset automatically created during the evaluation run of model asharsha30/LLAMA_Harsha_8_B_ORDP_10k The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/asharsha30__LLAMA_Harsha_8_B_ORDP_10k-details.tabular10K<n<100K0 likes43 downloads2y agoHugging Face11hybrid-diff-ar /stack-v2-sparse-classes-10k Stack v2 Sparse Python Classes 10k This is a 10,000-sample snapshot for Diffusion + Autoregressive hybrid code generation experiments. Source The data is extracted from bigcode/the-stack-v2-dedup, Python subset. The extraction uses Stack v2 metadata as source of truth, groups candidates by repo_name + revision_id, fetches files with git partial fetch + sparse checkout, then applies AST-level class filters. Splits train.jsonl: 9,000 val.jsonl: 500 test.jsonl:… See the full description on the dataset page: https://huggingface.co/datasets/hybrid-diff-ar/stack-v2-sparse-classes-10k.tabulartext-generation10K<n<100K0 likes41 downloads5mo agoHugging Face12FierceLLM /ru-instruct-10k 10k Russian chatbot dialogues dataset tabulartext-generation1K<n<10K1 likes38 downloads6mo agoHugging Face13violetxi /tb21-eval-qwen35-rewritten-w005-10k-c164-max32k-timeout2x qwen35-rewritten-w005-10k — Terminal-Bench 2.1 Noncanonical Terminal-Bench 2.1 evaluation of violetxi/qwen35-4b-offline-echo-rewritten-obs-wm-weight-0p05-10k-tacc through the served model ID qwen35-rewritten-w005-10k with Terminus-2. Noncanonical run: timeout_multiplier=2 instead of 1.0; concurrency=164 exceeds 30. Do not compare this score directly with canonical TB2.1 leaderboard runs. Result Trials: 445 Tasks / attempts: 89 × 5 Errored trials scored as zero:… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/tb21-eval-qwen35-rewritten-w005-10k-c164-max32k-timeout2x.tabularreinforcement-learningn<1K0 likes37 downloads2mo agoHugging Face14violetxi /tb21-eval-qwen35-rewritten-w005-10k-thinking-32k-timeout2x qwen35-rewritten-w005-10k — Terminal-Bench 2.1 Noncanonical Terminal-Bench 2.1 evaluation of violetxi/qwen35-4b-offline-echo-rewritten-obs-wm-weight-0p05-10k-tacc through the served model ID qwen35-rewritten-w005-10k with Terminus-2. Noncanonical run: timeout_multiplier=2 instead of 1.0; concurrency=164 exceeds 30. Do not compare this score directly with canonical TB2.1 leaderboard runs. Result Trials: 445 Tasks / attempts: 89 × 5 Errored trials scored as zero:… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/tb21-eval-qwen35-rewritten-w005-10k-thinking-32k-timeout2x.tabularreinforcement-learningn<1K0 likes37 downloads2mo agoHugging Face15maharnab /ultrafeedback_binarized_10ktabular10K<n<100K1 likes33 downloads1y agoHugging Face16violetxi /tb21-eval-qwen35-4b-offline-echo-action-only-10k-tacc-timeout2x qwen35-action-only-10k — Terminal-Bench 2.1 (timeout multiplier 2x) Terminal-Bench 2.1 evaluation protocol variant (timeout multiplier 2x) of violetxi/qwen35-4b-offline-echo-action-only-10k-tacc through the served model ID qwen35-action-only-10k with Terminus-2. Result Evaluation trials: 445 Tasks / attempts: 89 × 5 Errored trials scored as zero: 219 Agent timeouts / context-length events / output-cap events: 217 / 0 / 0 Mean reward / Pass@1: 0.105618 Pass@5:… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/tb21-eval-qwen35-4b-offline-echo-action-only-10k-tacc-timeout2x.tabularreinforcement-learningn<1K0 likes33 downloads2mo agoHugging Face17CL-From-Nothing /code_rose_initial_1_7B_SFT_10K_rollouts_Qwen3-4B-Thinking-2507_k12_t0.7_maxtok12288 code_rose_initial_1_7B_SFT_10K — rollouts (Qwen3-4B-Thinking-2507, k=12) Pass@k completions generated with vLLM over the prefixes in CL-From-Nothing/code_rose_initial_1_7B_SFT_10K. Generation config Model Qwen3-4B-Thinking-2507 Samples per question (k) 12 Temperature 0.7 top_p 0.9 max_tokens 12288 max_model_len 32768 Questions 7250 (index 0–7249, full split) Total rows 87000 (7250 × 12) Generated by complete_prefix_vllm.py… See the full description on the dataset page: https://huggingface.co/datasets/CL-From-Nothing/code_rose_initial_1_7B_SFT_10K_rollouts_Qwen3-4B-Thinking-2507_k12_t0.7_maxtok12288.tabulartext-generation10K<n<100K0 likes23 downloads3mo agoHugging Face18Lyraix-AI /Lyraix_bench_10k LyraixGuard Benchmark 10K v5 Internal benchmark dataset for evaluating AI security classification models. 10,000 curated samples balanced across three safety classes. Distribution Class Count % Safe 3,400 34.0% Unsafe 3,400 34.0% Controversial 3,200 32.0% Schema messages: 3-message ChatML array (system + user + assistant) conv_id: Conversation identifier safety_class: Ground truth (Safe/Unsafe/Controversial) category: Attack category or… See the full description on the dataset page: https://huggingface.co/datasets/Lyraix-AI/Lyraix_bench_10k.tabulartext-classification10K<n<100K0 likes22 downloads6mo agoHugging Face19violetxi /tb21-eval-qwen35-original-w005-10k-c164-max32k-timeout2x qwen35-original-w005-10k — Terminal-Bench 2.1 Noncanonical Terminal-Bench 2.1 evaluation of violetxi/qwen35-4b-offline-echo-original-obs-wm-weight-0p05-10k-tacc through the served model ID qwen35-original-w005-10k with Terminus-2. Noncanonical run: timeout_multiplier=2 instead of 1.0; concurrency=164 exceeds 30. Do not compare this score directly with canonical TB2.1 leaderboard runs. Operator-directed failure: build-pov-ray__rpB5FQS was manually terminated and counted as… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/tb21-eval-qwen35-original-w005-10k-c164-max32k-timeout2x.tabularreinforcement-learningn<1K0 likes22 downloads2mo agoHugging Face20cahlen /cdg-larry-king-interviews-10k Larry King Synthetic Conversations Dataset Description This dataset comprises synthetic, multi-turn conversational dialogues emulating interviews conducted by Larry King. The conversations delve into themes such as storytelling, resilience, and personal growth. Each dialogue is structured with alternating turns between 'Larry King' and a guest, capturing the essence of empathetic and reflective interviews. Dataset Details Curated by: Cahlen Humphreys Generated… See the full description on the dataset page: https://huggingface.co/datasets/cahlen/cdg-larry-king-interviews-10k.tabular10K<n<100K0 likes20 downloads1y agoHugging Face21Lux0926 /MetaMath-Llama-8B-CGPO-10ktabular10K<n<100K0 likes19 downloads11mo agoHugging Face22OpenCoven /fable-forge-10k FableForge — Narrative Reasoning Dataset with Recurrence-Depth Annotations The first narrative dataset designed around recurrence depth requirements. Every example carries a suggested_n_loops field with a theoretically grounded basis — derived from the structural complexity of the task, not a heuristic label or emergent property. Background Standard narrative datasets treat reasoning depth as an emergent property. FableForge is different: it annotates how much… See the full description on the dataset page: https://huggingface.co/datasets/OpenCoven/fable-forge-10k.tabulartext-generation10K<n<100K0 likes18 downloads3mo agoHugging Face23Lux0926 /Qwen1.5-32B-SFT-CGPO-10ktabular10K<n<100K0 likes15 downloads11mo agoHugging Face24Sangadi-Bujji /GHIA-CHRONOS-Synthetic-Dialogue-10K 🌌 GHIA-CHRONOS: The Industrial Ops Corpus A Recursive Civilization Simulation Dataset for Long-Horizon AI Reasoning 📘 Dataset Overview Field Information Dataset Name GHIA-CHRONOS Dataset Type Synthetic Recursive Civilization Dataset Primary Purpose Long-horizon reasoning, relativistic causality, strategic simulation Data Format JSONL Generation Style Optimized low-power recursive streaming Current Public Sample 10,000 records Master… See the full description on the dataset page: https://huggingface.co/datasets/Sangadi-Bujji/GHIA-CHRONOS-Synthetic-Dialogue-10K.tabulartext-generation10K<n<100K0 likes11 downloads5mo agoHugging Face25open-llm-leaderboard /FlofloB__10k_continued_pretraining_Phi-3-mini-4k-instruct_Unsloth_merged_16bit-detailsgated Dataset Card for Evaluation run of FlofloB/10k_continued_pretraining_Phi-3-mini-4k-instruct_Unsloth_merged_16bit Dataset automatically created during the evaluation run of model FlofloB/10k_continued_pretraining_Phi-3-mini-4k-instruct_Unsloth_merged_16bit The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/FlofloB__10k_continued_pretraining_Phi-3-mini-4k-instruct_Unsloth_merged_16bit-details.tabular10K<n<100K0 likes10 downloads2y agoHugging Face26Lux0926 /MetaMath-Mistral-7B-CGPO-10ktabular10K<n<100K0 likes10 downloads11mo agoHugging Face27sheng22213 /helpful_harmless_data_10ktabular10K<n<100K0 likes10 downloads3mo agoHugging Face28LangAGI-Lab /train-rl-o1-mini-annotated-math-numina-10k-numeric-answertabular10K<n<100K0 likes9 downloads2y agoHugging Face29Lux0926 /Deepseek-Coder-7B-Instruct-v1.5-CGPO-10ktabular10K<n<100K0 likes9 downloads11mo agoHugging Face30JingweiNi /ClimateMBERT-syn-qwen3-30b-a3b-fp8-10k-seed42 ClimateMBERT Synthetic Qwen3 30B A3B FP8 10K Seed42 Synthetic continuation dataset generated from WxChat/ClimateMBERT_syn train split. Source dataset: WxChat/ClimateMBERT_syn Source split: train Sampling: shuffled with random seed 42, ranks 0..9999 Rows: 10,000 Generator: Qwen/Qwen3-30B-A3B-Instruct-2507-FP8 Inference: vLLM on Clariden GH200 GPUs, tensor parallel size 2, non-eager mode Max tokens: 4096 Generation config: temperature 0.7, top_p 0.8, top_k 20, min_p 0.0… See the full description on the dataset page: https://huggingface.co/datasets/JingweiNi/ClimateMBERT-syn-qwen3-30b-a3b-fp8-10k-seed42.tabulartext-generation10K<n<100K0 likes9 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.