datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
diffusers-metadataRationalRewards_DiffusionNFT_TrainDataTLDR: this is the diffusion RL training dataset for text-to-image generation and image editing, from the following paper.
RationalRewards: Reasoning Rewards Scale Visual Generation Both Training and Test Time
Haozhe Wang1
Cong Wei2
Weiming Ren2
Jiaming Liu3
Fangzhen Lin1
Wenhu Chen2
1 HKUST
2 University of Waterloo
3 Alibaba
RationalRewards is a… See the full description on the dataset page: https://huggingface.co/datasets/TIGER-Lab/RationalRewards_DiffusionNFT_TrainData.2026-08-20-odcv-feature-discovery-difficult-advice-716-5-pct-vs-numina-control
LLM-driven feature discovery over ODCV-Bench rollouts from TWO matched Qwen3.6-27B LoRA arms — 9,284 filtered instruction rows plus 716 rows that differ only in kind (constitution-grounded difficult advice vs NuminaMath chain-of-thought) — asking which reasoning and action properties separate the two models, and which go with the judged misalignment.
field
value
experiment
LLM-driven feature discovery over ODCV-Bench rollouts from TWO matched Qwen3.6-27B LoRA arms — 9… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-20-odcv-feature-discovery-difficult-advice-716-5-pct-vs-numina-control.RL-seed-Decensor-DifficultyImagePulseV2-Edit-Change
ImagePulseV2 Dataset - Foreground Editing
The ImagePulseV2 dataset is a custom-built dataset we created for training the Diffusion Templates series of models. It consists of multiple subsets generated using models such as Z-Image-Turbo, Qwen-Image, and Qwen-Image-Edit, based on prompts randomly sampled from DiffusionDB.
Open-source code: DiffSynth-Studio
Technical report: arXiv
Project homepage: GitHub
Documentation: English Version, Chinese Version
Online demo: ModelScope Studio… See the full description on the dataset page: https://huggingface.co/datasets/DiffSynth-Studio/ImagePulseV2-Edit-Change.prompt-difficulty
Prompt Difficulty Assessment
Prompt difficulty plays a critical role in the performance of large language models (LLMs).
Assessing this difficulty is essential for selecting training examples, evaluating model capabilities, and optimizing routing and reasoning strategies.
Yet, no standardized framework exists for comparing prompt difficulty across domains.
This report proposes a method to quantify prompt difficulty using multiple LLMs and introduces a composite difficulty score for… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/prompt-difficulty.EvoEval_difficultImagePulseV2-Edit-AddRemove
ImagePulseV2 Dataset - Local Add/Delete
The ImagePulseV2 dataset is a collection we constructed for training the Diffusion Templates series of models. It consists of multiple subsets generated using models such as Z-Image-Turbo, Qwen-Image, and Qwen-Image-Edit, based on prompts randomly sampled from DiffusionDB.
Open-source code: DiffSynth-Studio
Technical report: arXiv
Project homepage: GitHub
Documentation: English Version, Chinese Version
Online demo: ModelScope Studio
Models:… See the full description on the dataset page: https://huggingface.co/datasets/DiffSynth-Studio/ImagePulseV2-Edit-AddRemove.pi-diff-review
Coding agent session traces for badlogicgames/pi-diff-review
This dataset contains redacted coding agent session traces collected while working on https://github.com/badlogic/pi-diff-review.git. The traces were exported with pi-share-hf from a local pi workspace and filtered to keep only sessions that passed deterministic redaction and LLM review.
Data description
Each *.jsonl file is a redacted pi session. Sessions are stored as JSON Lines files where each line… See the full description on the dataset page: https://huggingface.co/datasets/badlogicgames/pi-diff-review.flutter-diff-steps-v1
Flutter Codegen: Diff Steps
Synthetic dataset of step-by-step Flutter/Dart widget construction, where each
row is one incremental edit in a sequence: given a goal, the current code, and the
history of steps taken so far, predict the next action (a short description) and
the code change as a search/replace diff hunk.
Built for training and evaluating small language models on iterative, diff-based
code editing -- as opposed to regenerating the whole file at each step. This is
the… See the full description on the dataset page: https://huggingface.co/datasets/bbidpa/flutter-diff-steps-v1.pi-diff-review
Coding agent session traces for badlogicgames/pi-diff-review
This dataset contains redacted coding agent session traces collected while working on https://github.com/badlogic/pi-diff-review.git. The traces were exported with pi-share-hf from a local pi workspace and filtered to keep only sessions that passed deterministic redaction and LLM review.
Data description
Each sessions/*.jsonl file is a redacted pi session. Sessions are stored as JSON Lines files where… See the full description on the dataset page: https://huggingface.co/datasets/cfahlgren1/pi-diff-review.2026-08-25-table2-9284-difficult-advice-verbose-token-matched-train-mixture
Token-matched verbose difficult-advice arm. Holds difficult advice's share of the TRAINABLE TOKENS at the control's value while the traces are ~3x longer, by keeping only a subset of the expanded rows. Its sibling arm holds the ROW share instead; together they separate more deliberation from more difficult-advice signal.
field
value
experiment
Token-matched verbose difficult-advice arm. Holds difficult advice's share of the TRAINABLE TOKENS at the control's value… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-25-table2-9284-difficult-advice-verbose-token-matched-train-mixture.requests-pr-diff
requests-pr-diff
Generated by Repo2RLEnv.
💡 Browse this dataset in your browser — click the badge above or open
HuggingFaceH4/harbor-visualiser
to inspect every task's spec, instruction, oracle patch, test script, and Dockerfile.
Source repo: psf/requests
Pipeline: pr_diff
Tasks: 47
Visibility: public
Spec: Harbor task format with [metadata.repo2env] extension
Reward kinds
This dataset emits diff_similarity rewards. Each task ships an oracle diff at… See the full description on the dataset page: https://huggingface.co/datasets/sergiopaniego/requests-pr-diff.2026-08-04-qwen36-27b-1000ex-difficult-advice-350-numina-650-train-mixture
Qwen3.6-27B training bundle — 2026-08-04-qwen36-27b-1000ex-da350-numina650-train
code.tar.gz (trainer, src/, configs/) plus mixture.jsonl. The pod untars it, copies
the jsonl to data/mixture.jsonl, and runs configs/train_1000ex_da350_numina650.yaml.
field
value
experiment
1000-example mixture: 350 difficult-advice (all t1-t3 + t4 fill) + 650 NuminaMath-CoT; lr 4e-5, 1 epoch
date_generated
2026-08-03
constitution
constitutions/claude_constitution_principles.md —… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-04-qwen36-27b-1000ex-difficult-advice-350-numina-650-train-mixture.ImagePulseV2-Edit-Style
ImagePulseV2 Dataset - Style Transfer
The ImagePulseV2 dataset is a custom-built dataset we created for training the Diffusion Templates series of models. It consists of multiple subsets generated using models such as Z-Image-Turbo, Qwen-Image, and Qwen-Image-Edit, based on prompts randomly sampled from DiffusionDB.
Open-source code: DiffSynth-Studio
Technical report: arXiv
Project homepage: GitHub
Documentation: English Version, Chinese Version
Online demo: ModelScope Studio
Model… See the full description on the dataset page: https://huggingface.co/datasets/DiffSynth-Studio/ImagePulseV2-Edit-Style.2026-08-25-difficult-advice-716-verbose-cot
Difficult-advice reasoning, expanded ~3x (716 records, no other data)
field
value
experiment
Does deliberation LENGTH change alignment behaviour, holding the ideas deliberated constant? These are the 716 difficult-advice exchanges of LASR-Callum/2026-08-13-haiku45-sonnet45-difficult-advice-diversity-gated-voice-linted, with the assistant's private reasoning rewritten about three times longer while carrying the same content. The user turn, the system prompt and the… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-25-difficult-advice-716-verbose-cot.ImagePulseV2-Edit-HumanFace
ImagePulseV2 Dataset - Facial Expression Editing
The ImagePulseV2 dataset is a collection we constructed for training the Diffusion Templates series of models. It contains multiple subsets generated using models such as Z-Image-Turbo, Qwen-Image, and Qwen-Image-Edit, based on prompts randomly sampled from DiffusionDB.
Open-source code: DiffSynth-Studio
Technical report: arXiv
Project homepage: GitHub
Documentation: English Version, Chinese Version
Online demo: ModelScope Studio… See the full description on the dataset page: https://huggingface.co/datasets/DiffSynth-Studio/ImagePulseV2-Edit-HumanFace.ImagePulseV2-Edit-Pose
ImagePulseV2 Dataset - Pose Adjustment
The ImagePulseV2 dataset is a custom-built dataset created for training the Diffusion Templates series of models. It consists of multiple subsets generated using models such as Z-Image-Turbo, Qwen-Image, and Qwen-Image-Edit, based on prompts randomly sampled from DiffusionDB.
Open-source code: DiffSynth-Studio
Technical report: arXiv
Project homepage: GitHub
Documentation: English Version, Chinese Version
Online demo: ModelScope Studio
Model… See the full description on the dataset page: https://huggingface.co/datasets/DiffSynth-Studio/ImagePulseV2-Edit-Pose.diffusion_data_constraint_c4subsetsdiffsImagePulseV2-Edit-Light
ImagePulseV2 Dataset - Lighting Adjustment
The ImagePulseV2 dataset is a custom-built dataset we created for training the Diffusion Templates series of models. It consists of multiple subsets generated using models such as Z-Image-Turbo, Qwen-Image, and Qwen-Image-Edit, based on prompts randomly sampled from DiffusionDB.
Open-source code: DiffSynth-Studio
Technical report: arXiv
Project homepage: GitHub
Documentation: English Version, Chinese Version
Try online: ModelScope Studio… See the full description on the dataset page: https://huggingface.co/datasets/DiffSynth-Studio/ImagePulseV2-Edit-Light.2026-08-31-difficult-advice-716-seeds-bundle
da716 seed replicates — training bundle (seeds 42 and 69)
code.tar.gz (trainer + src/ + the two seed configs) beside seed 0's mixture,
byte-identical. scripts/gpu/runpod_train.py up reads both from this one repo.
field
value
experiment
Seed replicates of the da716 arm (Table2 9,284 filtered + difficult-advice-v2 716, 7.16%) so the arm carries training-seed variance like its siblings. da716 was the last arm on a single seed and is the comparison baseline for the… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-31-difficult-advice-716-seeds-bundle.2026-08-03-qwen36-27b-1000ex-difficult-advice-250-numina-750-train-mixture
Qwen3.6-27B training bundle — 1,000 examples (250 synthdoc_v2 t1-t3 + 750 NuminaMath-CoT)
RunPod training bundle: code.tar.gz (the trainer, src/, configs/) plus mixture.jsonl.
The pod pulls this, untars it, copies the jsonl to data/, and runs configs/train_1000ex_da250_numina750.yaml.
field
value
experiment
1-epoch assistant-only-loss LoRA SFT of Qwen3.6-27B on a 1,000-example mixture that is 25% difficult-advice by example count
date_generated
2026-08-03… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-03-qwen36-27b-1000ex-difficult-advice-250-numina-750-train-mixture.2026-08-19-random-220-difficult-advice-control-train
Random 220-row difficult-advice control for the LESS top-10% arm
field
value
experiment
THE CONTROL ARM of a paired LESS (arXiv:2402.04333) data-selection experiment: SFT training file holding 220 rows drawn uniformly at random (seed 1) from the same 2203-row difficult-advice pool, trained as-is on base Qwen3.6-27B with no other data. The two arms differ ONLY in which 220 of the same 2,203 rows they hold — identical tokenizer, budget, seed, shuffle and training recipe… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-19-random-220-difficult-advice-control-train.miles-diffusion-datasets
miles-diffusion-datasets
Prompt datasets used by the miles-diffusion training pipeline.
Layout
flowgrpo_ocr/ # OCR-style prompts from flowGRPO
├── train.jsonl # 19,653
└── test.jsonl # 1,018
flowgrpo_pickscore/ # PickScore prompts from flowGRPO
├── train.jsonl # 25,432
└── test.jsonl # 2,048
hpdv2/ # Deduplicated prompts from HPDv2
├──… See the full description on the dataset page: https://huggingface.co/datasets/rockdu/miles-diffusion-datasets.2026-08-03-qwen36-27b-armb-1000ex-difficult-advice-250-rest-750-train-mixture
Qwen3.6-27B training bundle — 2026-08-03-qwen36-27b-armB-1000ex-da250-rest750-train
code.tar.gz (trainer, src/, configs/) plus mixture.jsonl. The pod untars it, copies
the jsonl to data/mixture.jsonl, and runs configs/train_armB_1000ex_da250_rest750.yaml.
field
value
experiment
Arm B: 250 difficult-advice (t1-t3) + 750 at 3:2 NuminaMath : (TULU3 + No Robots)
date_generated
2026-08-03
constitution
constitutions/claude_constitution_principles.md — principles t1-t3… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-03-qwen36-27b-armb-1000ex-difficult-advice-250-rest-750-train-mixture.ImagePulseV2-Edit-Angle
ImagePulseV2 Dataset - Viewpoint Adjustment
The ImagePulseV2 dataset is a custom dataset we constructed for training the Diffusion Templates series of models. It consists of multiple subsets generated using models such as Z-Image-Turbo, Qwen-Image, and Qwen-Image-Edit, based on prompts randomly sampled from DiffusionDB.
Open-source code: DiffSynth-Studio
Technical report: arXiv
Project homepage: GitHub
Documentation: English Version, Chinese Version
Online demo: ModelScope Studio… See the full description on the dataset page: https://huggingface.co/datasets/DiffSynth-Studio/ImagePulseV2-Edit-Angle.dapo-math-17k-difficulty-qwen3-1.7b-base-k16
DAPO-Math-17k difficulty under Qwen3-1.7B-Base (K=16)
For each of the 17,398 problems in the DAPO-Math-17k train set, how many of
K=16 samples from the untrained base model are correct.
The headline: 57.27% of problems are solved 0 out of 16 times, and not one
problem is solved 16 out of 16. Difficulty here is entirely one-sided.
Why count per problem instead of reporting mean accuracy
In group-relative RL (GRPO and its relatives), a prompt group whose K responses… See the full description on the dataset page: https://huggingface.co/datasets/RyanYr/dapo-math-17k-difficulty-qwen3-1.7b-base-k16.2026-08-19-less-top10-difficult-advice-220-train
LESS top-10% difficult-advice selection (220 rows, score_max)
field
value
experiment
THE LESS ARM of a paired LESS (arXiv:2402.04333) data-selection experiment: SFT training file holding the 220 highest-influence rows (top 10%) of the difficult-advice pool by score_max, trained as-is on base Qwen3.6-27B with no other data. The two arms differ ONLY in which 220 of the same 2,203 rows they hold — identical tokenizer, budget, seed, shuffle and training recipe — so a… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-19-less-top10-difficult-advice-220-train.2026-08-31-difficult-advice-principle-scoped-702-seeds-bundle
chunk-only 702 seed replicates — training bundle (seeds 42 and 69)
code.tar.gz (trainer + src/ + the two seed configs) beside seed 0's mixture,
byte-identical. scripts/gpu/runpod_train.py up reads both from this one repo.
field
value
experiment
Seed replicates so this arm carries training-seed variance. Table2 9,284 filtered + chunk-only difficult advice 702 (7.03%). The rewrite stages never saw the constitution, only their one target principle. Between-seed spread on… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-31-difficult-advice-principle-scoped-702-seeds-bundle.
