CoolFace
Datasetpublic

rockdu/miles-diffusion-datasets

miles-diffusion-datasets Prompt datasets used by the miles-diffusion training pipeline. Layout flowgrpo_ocr/ # OCR-style prompts from flowGRPO ├── train.jsonl # 19,653 └── test.jsonl # 1,018 flowgrpo_pickscore/ # PickScore prompts from flowGRPO ├── train.jsonl # 25,432 └── test.jsonl # 2,048 hpdv2/ # Deduplicated prompts from HPDv2 ├──… See the full description on the dataset page: https://huggingface.co/datasets/rockdu/miles-diffusion-datasets.

sourceHugging Faceupdated 20d agoView on Hugging Face
0likes72downloads
Dataset Card

miles-diffusion-datasets

Prompt datasets used by the miles-diffusion training pipeline.

Layout

flowgrpo_ocr/                    # OCR-style prompts from flowGRPO
├── train.jsonl                  # 19,653
└── test.jsonl                   # 1,018

flowgrpo_pickscore/              # PickScore prompts from flowGRPO
├── train.jsonl                  # 25,432
└── test.jsonl                   # 2,048

hpdv2/                           # Deduplicated prompts from HPDv2
├── train.jsonl                  # 103,672
└── test.jsonl                   # 242

Directory names record the upstream source so the schema stays clear when more datasets land here later (other reward families, internal prompts, etc.).

Format

Each .jsonl file has one JSON object per line:

json
{"input": "A close-up of a medicine bottle with a clear, red warning label..."}

UTF-8 encoded (ensure_ascii=False), no escaping for non-ASCII characters.

Usage

bash
hf download --repo-type dataset rockdu/miles-diffusion-datasets \
    --local-dir /root/datasets/miles-diffusion-datasets

python train_diffusion.py \
    --prompt-data /root/datasets/miles-diffusion-datasets/flowgrpo_ocr/train.jsonl \
    --eval-prompt-data ocr_test /root/datasets/miles-diffusion-datasets/flowgrpo_ocr/test.jsonl \
    ...

Provenance

flowgrpo_* directories: extracted from flow_grpo's dataset/{ocr,pickscore}/{train,test}.txt files (one prompt per line) and converted to {"input": prompt} JSONL format.

hpdv2/: extracted from the prompt field of `ymhao/HPDv2` at revision bb0f2563e451e11c3f376b2ba93615db00e042e1. Prompts are stripped, empty or non-string values are dropped, and each source split is deduplicated while preserving first occurrence. Prompts present in train are removed from test, leaving no cross-split overlap. The source dataset is licensed under Apache-2.0; please cite the original HPSv2 paper when using these prompts.