datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
PKU-SafeRLHF-10K
Paper
You can find more information in our paper.
Dataset Paper: https://arxiv.org/abs/2307.04657
TripVVT-10K
TripVVT-10K Dataset
News
2026.06: TripVVT has been accepted by ECCV 2026.
2026.04: The TripVVT paper is available on arXiv.
The project page is available at https://shaodingbao.github.io/TripVVT/.
TripVVT-10K is a large-scale dataset for in-the-wild Video Virtual Try-On (VVT). It contains 10,031 high-quality video samples with triplet supervision, covering upper-body garments, lower-body garments, and dresses.
TripVVT-10K is released together with the… See the full description on the dataset page: https://huggingface.co/datasets/TripVVT/TripVVT-10K.cinepile_10kPhysDPO-10kTo use PhysDPO, run the following command to combine the parts into a single ZIP file:
cat PhysDPO_part_* > PhysDPO.zip
whestbench-relu-mlp-moments-10k
WhestBench Random ReLU MLPs — Monte-Carlo Activation Cumulants (10k)
A dataset of 10,500 random ReLU MLPs (10,000 train + 500 held-out test) together with
Monte-Carlo–estimated per-layer activation cumulants (mean, variance, skewness, kurtosis) for
every layer, for both the pre-activation and post-ReLU signals. Built for the WhestBench
Estimation Challenge 2026 and for research on analytic moment / uncertainty propagation
through deep networks.
The generative process… See the full description on the dataset page: https://huggingface.co/datasets/keenanpepper/whestbench-relu-mlp-moments-10k.openorca-multiplechoice-10kA 10k subset of OpenOrca dataset, focusing on multiple choice questions.
Credit to Tian Xia.
filings-10kprompt-gen-10k-flux-sdxl
Prompt Generation Dataset (10K Narrative for Flux / SDXL)
This dataset (prompt_gen_final_10k.jsonl and prompt_gen_final_10k.csv) was used to train and fine-tune image-prompt models such as KavinduHansaka/Llama-3.2-1B-ImageGen.
It contains 10,000 curated narrative prompt samples designed for image generation models like Stable Diffusion XL and Flux.Unlike raw tag-based datasets, the target field provides natural paragraphs (≈80–100 words) that describe cinematic scenes with… See the full description on the dataset page: https://huggingface.co/datasets/KavinduHansaka/prompt-gen-10k-flux-sdxl.LLaVA-Human-Preference-10Kasharsha30__LLAMA_Harsha_8_B_ORDP_10k-details
Dataset Card for Evaluation run of asharsha30/LLAMA_Harsha_8_B_ORDP_10k
Dataset automatically created during the evaluation run of model asharsha30/LLAMA_Harsha_8_B_ORDP_10k
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/asharsha30__LLAMA_Harsha_8_B_ORDP_10k-details.stack-v2-sparse-classes-10k
Stack v2 Sparse Python Classes 10k
This is a 10,000-sample snapshot for Diffusion + Autoregressive hybrid code generation experiments.
Source
The data is extracted from bigcode/the-stack-v2-dedup, Python subset. The extraction uses Stack v2 metadata as source of truth, groups candidates by repo_name + revision_id, fetches files with git partial fetch + sparse checkout, then applies AST-level class filters.
Splits
train.jsonl: 9,000
val.jsonl: 500
test.jsonl:… See the full description on the dataset page: https://huggingface.co/datasets/hybrid-diff-ar/stack-v2-sparse-classes-10k.ru-instruct-10k
10k Russian chatbot dialogues dataset
tb21-eval-qwen35-rewritten-w005-10k-c164-max32k-timeout2x
qwen35-rewritten-w005-10k — Terminal-Bench 2.1
Noncanonical Terminal-Bench 2.1 evaluation of violetxi/qwen35-4b-offline-echo-rewritten-obs-wm-weight-0p05-10k-tacc through the served
model ID qwen35-rewritten-w005-10k with Terminus-2.
Noncanonical run: timeout_multiplier=2 instead of 1.0; concurrency=164 exceeds 30. Do not compare this score directly with canonical TB2.1 leaderboard runs.
Result
Trials: 445
Tasks / attempts: 89 × 5
Errored trials scored as zero:… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/tb21-eval-qwen35-rewritten-w005-10k-c164-max32k-timeout2x.tb21-eval-qwen35-rewritten-w005-10k-thinking-32k-timeout2x
qwen35-rewritten-w005-10k — Terminal-Bench 2.1
Noncanonical Terminal-Bench 2.1 evaluation of violetxi/qwen35-4b-offline-echo-rewritten-obs-wm-weight-0p05-10k-tacc through the served
model ID qwen35-rewritten-w005-10k with Terminus-2.
Noncanonical run: timeout_multiplier=2 instead of 1.0; concurrency=164 exceeds 30. Do not compare this score directly with canonical TB2.1 leaderboard runs.
Result
Trials: 445
Tasks / attempts: 89 × 5
Errored trials scored as zero:… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/tb21-eval-qwen35-rewritten-w005-10k-thinking-32k-timeout2x.ultrafeedback_binarized_10ktb21-eval-qwen35-4b-offline-echo-action-only-10k-tacc-timeout2x
qwen35-action-only-10k — Terminal-Bench 2.1 (timeout multiplier 2x)
Terminal-Bench 2.1 evaluation protocol variant (timeout multiplier 2x) of violetxi/qwen35-4b-offline-echo-action-only-10k-tacc through the served
model ID qwen35-action-only-10k with Terminus-2.
Result
Evaluation trials: 445
Tasks / attempts: 89 × 5
Errored trials scored as zero: 219
Agent timeouts / context-length events / output-cap events:
217 / 0 /
0
Mean reward / Pass@1: 0.105618
Pass@5:… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/tb21-eval-qwen35-4b-offline-echo-action-only-10k-tacc-timeout2x.code_rose_initial_1_7B_SFT_10K_rollouts_Qwen3-4B-Thinking-2507_k12_t0.7_maxtok12288
code_rose_initial_1_7B_SFT_10K — rollouts (Qwen3-4B-Thinking-2507, k=12)
Pass@k completions generated with vLLM over the prefixes in
CL-From-Nothing/code_rose_initial_1_7B_SFT_10K.
Generation config
Model
Qwen3-4B-Thinking-2507
Samples per question (k)
12
Temperature
0.7
top_p
0.9
max_tokens
12288
max_model_len
32768
Questions
7250 (index 0–7249, full split)
Total rows
87000 (7250 × 12)
Generated by complete_prefix_vllm.py… See the full description on the dataset page: https://huggingface.co/datasets/CL-From-Nothing/code_rose_initial_1_7B_SFT_10K_rollouts_Qwen3-4B-Thinking-2507_k12_t0.7_maxtok12288.Lyraix_bench_10k
LyraixGuard Benchmark 10K v5
Internal benchmark dataset for evaluating AI security classification models. 10,000 curated samples balanced across three safety classes.
Distribution
Class
Count
%
Safe
3,400
34.0%
Unsafe
3,400
34.0%
Controversial
3,200
32.0%
Schema
messages: 3-message ChatML array (system + user + assistant)
conv_id: Conversation identifier
safety_class: Ground truth (Safe/Unsafe/Controversial)
category: Attack category or… See the full description on the dataset page: https://huggingface.co/datasets/Lyraix-AI/Lyraix_bench_10k.tb21-eval-qwen35-original-w005-10k-c164-max32k-timeout2x
qwen35-original-w005-10k — Terminal-Bench 2.1
Noncanonical Terminal-Bench 2.1 evaluation of violetxi/qwen35-4b-offline-echo-original-obs-wm-weight-0p05-10k-tacc through the served
model ID qwen35-original-w005-10k with Terminus-2.
Noncanonical run: timeout_multiplier=2 instead of 1.0; concurrency=164 exceeds 30. Do not compare this score directly with canonical TB2.1 leaderboard runs.
Operator-directed failure: build-pov-ray__rpB5FQS was manually terminated and counted as… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/tb21-eval-qwen35-original-w005-10k-c164-max32k-timeout2x.cdg-larry-king-interviews-10k
Larry King Synthetic Conversations
Dataset Description
This dataset comprises synthetic, multi-turn conversational dialogues emulating interviews conducted by Larry King. The conversations delve into themes such as storytelling, resilience, and personal growth. Each dialogue is structured with alternating turns between 'Larry King' and a guest, capturing the essence of empathetic and reflective interviews.
Dataset Details
Curated by: Cahlen Humphreys
Generated… See the full description on the dataset page: https://huggingface.co/datasets/cahlen/cdg-larry-king-interviews-10k.MetaMath-Llama-8B-CGPO-10kfable-forge-10k
FableForge — Narrative Reasoning Dataset with Recurrence-Depth Annotations
The first narrative dataset designed around recurrence depth requirements.
Every example carries a suggested_n_loops field with a theoretically grounded basis —
derived from the structural complexity of the task, not a heuristic label or emergent property.
Background
Standard narrative datasets treat reasoning depth as an emergent property. FableForge is
different: it annotates how much… See the full description on the dataset page: https://huggingface.co/datasets/OpenCoven/fable-forge-10k.Qwen1.5-32B-SFT-CGPO-10kGHIA-CHRONOS-Synthetic-Dialogue-10K
🌌 GHIA-CHRONOS: The Industrial Ops Corpus
A Recursive Civilization Simulation Dataset for Long-Horizon AI Reasoning
📘 Dataset Overview
Field
Information
Dataset Name
GHIA-CHRONOS
Dataset Type
Synthetic Recursive Civilization Dataset
Primary Purpose
Long-horizon reasoning, relativistic causality, strategic simulation
Data Format
JSONL
Generation Style
Optimized low-power recursive streaming
Current Public Sample
10,000 records
Master… See the full description on the dataset page: https://huggingface.co/datasets/Sangadi-Bujji/GHIA-CHRONOS-Synthetic-Dialogue-10K.FlofloB__10k_continued_pretraining_Phi-3-mini-4k-instruct_Unsloth_merged_16bit-details
Dataset Card for Evaluation run of FlofloB/10k_continued_pretraining_Phi-3-mini-4k-instruct_Unsloth_merged_16bit
Dataset automatically created during the evaluation run of model FlofloB/10k_continued_pretraining_Phi-3-mini-4k-instruct_Unsloth_merged_16bit
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/FlofloB__10k_continued_pretraining_Phi-3-mini-4k-instruct_Unsloth_merged_16bit-details.MetaMath-Mistral-7B-CGPO-10khelpful_harmless_data_10ktrain-rl-o1-mini-annotated-math-numina-10k-numeric-answerDeepseek-Coder-7B-Instruct-v1.5-CGPO-10kClimateMBERT-syn-qwen3-30b-a3b-fp8-10k-seed42
ClimateMBERT Synthetic Qwen3 30B A3B FP8 10K Seed42
Synthetic continuation dataset generated from WxChat/ClimateMBERT_syn train split.
Source dataset: WxChat/ClimateMBERT_syn
Source split: train
Sampling: shuffled with random seed 42, ranks 0..9999
Rows: 10,000
Generator: Qwen/Qwen3-30B-A3B-Instruct-2507-FP8
Inference: vLLM on Clariden GH200 GPUs, tensor parallel size 2, non-eager mode
Max tokens: 4096
Generation config: temperature 0.7, top_p 0.8, top_k 20, min_p 0.0… See the full description on the dataset page: https://huggingface.co/datasets/JingweiNi/ClimateMBERT-syn-qwen3-30b-a3b-fp8-10k-seed42.
