CoolFace
11 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01UCSC-VLAA /GPT-Image-Edit-1.5M GPT-Image-Edit-1.5M A Million-Scale, GPT-Generated Image Dataset 📃Arxiv | 🌐 Project Page | 💻Github GPT-Image-Edit-1.5M is a comprehensive image editing dataset that is built upon HQ-Edit, UltraEdit, OmniEdit and Complex-Edit, with all output images regenerated with GPT-Image-1. 📣 News [2025.08.20] 🚀 We provide a script for multi-process downloading. See Multi-process Download. [2025.07.27] 🤗 We release GPT-Image-Edit, a state-of-the-art image editing model with… See the full description on the dataset page: https://huggingface.co/datasets/UCSC-VLAA/GPT-Image-Edit-1.5M.imageimage-to-image1M<n<10M90 likes6.8k downloads1y agoHugging Face02bhyuan /gptsovits_datasetgated bhyuan/gptsovits_dataset GPT-SoVITS speech dataset, packed as WebDataset tar shards. Layout data/ train/ metadata.csv audio/ train-000.tar train-001.tar ... validation/ metadata.csv audio/ validation-000.tar ... test/ metadata.csv audio/ test-000.tar ... Shard counts: youshengshu_v5_test: 6536 tar shard(s) Inside each tar, every sample is a pair sharing a unique key: <key>.wav # raw… See the full description on the dataset page: https://huggingface.co/datasets/bhyuan/gptsovits_dataset.audiotext-to-speech10M<n<100M1 likes1.6k downloads4mo agoHugging Face03nishadsinghi /aime24_R1-Dist-Qwen-1.5B_32K_tok_VER-Qwen2.5-1.5B_dist_r1_qwen_1p5B_gpt_4o_verify_all_train_1e-5text1K<n<10K0 likes8 downloads2y agoHugging Face04nishadsinghi /math128_qwen2p57B_VER-Llama-3.1-8B-Inst_data-qwen_25_7b_gpt_4o_verify_train_e3_LR-5e-7_7Klentext10K<n<100K0 likes7 downloads2y agoHugging Face05nishadsinghi /MATH128_Qwen2.57BInst_ver_model-qwen_2.5_7b_verifier_train_gpt_4o_lr5e-7text10K<n<100K0 likes5 downloads2y agoHugging Face06nishadsinghi /math128_qwen2.5-7b_ver-llama3.1-8b_train_gpt_4o_verifications_e3_lr5e-7-31389-merged_4Ktokenstext10K<n<100K0 likes5 downloads2y agoHugging Face07nishadsinghi /aime24_qwen2.5-7b_ver-llama3.1-8b_train_gpt_4o_verifications_e3_lr5e-7-31389-merged_4Ktokenstext1K<n<10K0 likes5 downloads2y agoHugging Face08nishadsinghi /aime24_qwen2.5-7b_ver_Llama-3.1-8B-Instruct_data-qwen_25_7b_gpt_4o_verify_train_e3_LR-5e-7_7Klentext1K<n<10K0 likes5 downloads2y agoHugging Face09nishadsinghi /aime24_qwen2.5_7b_ver_qwen_2.5_7b_verifier_train_gpt_4o_lr5e-7_4Ktokenstext1K<n<10K0 likes5 downloads2y agoHugging Face10Khoale11hcmut /gpt4o_captions_1k5samples PACO WebDataset export for PACO-style localized caption data. Summary Samples: 1500 Shards: 2 Payload format inside each shard: pickle Split: train Config: PACO Layout Media files are stored in WebDataset tar shards. Each sample key is stable and becomes __key__ in the dataset viewer. Hugging Face will infer columns such as jpg, pickle, json, __key__, and __url__ from the shard contents. Manifest file: PACO/annotations.json image1K<n<10K1 likes5 downloads5mo agoHugging Face11nishadsinghi /math128_R1_1.5B_VER-Qwen2.5-1.5B_data-distill_r1_qwen_1p5B_gpt_4o_verify_proc_all_train_1e-5text10K<n<100K0 likes4 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.