datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
GPT-Image-Edit-1.5M
GPT-Image-Edit-1.5M A Million-Scale, GPT-Generated Image Dataset
📃Arxiv | 🌐 Project Page | 💻Github
GPT-Image-Edit-1.5M is a comprehensive image editing dataset that is built upon HQ-Edit, UltraEdit, OmniEdit and Complex-Edit, with all output images regenerated with GPT-Image-1.
📣 News
[2025.08.20] 🚀 We provide a script for multi-process downloading. See Multi-process Download.
[2025.07.27] 🤗 We release GPT-Image-Edit, a state-of-the-art image editing model with… See the full description on the dataset page: https://huggingface.co/datasets/UCSC-VLAA/GPT-Image-Edit-1.5M.gptsovits_dataset
bhyuan/gptsovits_dataset
GPT-SoVITS speech dataset, packed as WebDataset tar shards.
Layout
data/
train/
metadata.csv
audio/
train-000.tar
train-001.tar
...
validation/
metadata.csv
audio/
validation-000.tar
...
test/
metadata.csv
audio/
test-000.tar
...
Shard counts:
youshengshu_v5_test: 6536 tar shard(s)
Inside each tar, every sample is a pair sharing a unique key:
<key>.wav # raw… See the full description on the dataset page: https://huggingface.co/datasets/bhyuan/gptsovits_dataset.aime24_R1-Dist-Qwen-1.5B_32K_tok_VER-Qwen2.5-1.5B_dist_r1_qwen_1p5B_gpt_4o_verify_all_train_1e-5math128_qwen2p57B_VER-Llama-3.1-8B-Inst_data-qwen_25_7b_gpt_4o_verify_train_e3_LR-5e-7_7KlenMATH128_Qwen2.57BInst_ver_model-qwen_2.5_7b_verifier_train_gpt_4o_lr5e-7math128_qwen2.5-7b_ver-llama3.1-8b_train_gpt_4o_verifications_e3_lr5e-7-31389-merged_4Ktokensaime24_qwen2.5-7b_ver-llama3.1-8b_train_gpt_4o_verifications_e3_lr5e-7-31389-merged_4Ktokensaime24_qwen2.5-7b_ver_Llama-3.1-8B-Instruct_data-qwen_25_7b_gpt_4o_verify_train_e3_LR-5e-7_7Klenaime24_qwen2.5_7b_ver_qwen_2.5_7b_verifier_train_gpt_4o_lr5e-7_4Ktokensgpt4o_captions_1k5samples
PACO
WebDataset export for PACO-style localized caption data.
Summary
Samples: 1500
Shards: 2
Payload format inside each shard: pickle
Split: train
Config: PACO
Layout
Media files are stored in WebDataset tar shards.
Each sample key is stable and becomes __key__ in the dataset viewer.
Hugging Face will infer columns such as jpg, pickle, json, __key__, and __url__ from the shard contents.
Manifest file: PACO/annotations.json
math128_R1_1.5B_VER-Qwen2.5-1.5B_data-distill_r1_qwen_1p5B_gpt_4o_verify_proc_all_train_1e-5
