CoolFace
20 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01APRIL-AIGC /UltraVideo UltraVideo: High-Quality UHD 4K Video Dataset 🤓 Project    | 📑 Paper    | 🤗 Hugging Face (UltraVideo Dataset))   | 🤗 Hugging Face (UltraVideo-Long Dataset))   | 🤗 Hugging Face (UltraWan-1K/4K Weights)   UltraVideo: High-Quality UHD Video Dataset with Comprehensive Captions 🎋 Click below image to watch the 4K demo video. 🤓 First open-sourced UHD-4K/8K video datasets with comprehensive structured (10 types) captions.🤓 Native 1K/4K videos generation by UltraWan.… See the full description on the dataset page: https://huggingface.co/datasets/APRIL-AIGC/UltraVideo.tabularimage-to-video10K<n<100K66 likes11k downloads1y agoHugging Face02APRIL-AIGC /UltraVideo-Long UltraVideo: High-Quality UHD 4K Video Dataset 🤓 Project    | 📑 Paper    | 🤗 Hugging Face (UltraVideo Dataset))   | 🤗 Hugging Face (UltraVideo-Long Dataset))   | 🤗 Hugging Face (UltraWan-1K/4K Weights)   UltraVideo: High-Quality UHD Video Dataset with Comprehensive Captions 🎋 Click below image to watch the 4K demo video. 🤓 First open-sourced UHD-4K/8K video datasets with comprehensive structured (10 types) captions.🤓 Native 1K/4K videos generation by UltraWan.… See the full description on the dataset page: https://huggingface.co/datasets/APRIL-AIGC/UltraVideo-Long.tabularimage-to-video10K<n<100K7 likes3.9k downloads1y agoHugging Face03MAPLE-WestLake-AIGC /OpenstoryPlusPlus Openstory++: A Large-scale Dataset and Benchmark for Instance-aware Open-domain Visual Storytelling We introduce OpenStory++, a large-scale open-domain dataset contains focusing on enabling MLLMs to perform storytelling generation tasks. related resorcce paper: https://arxiv.org/abs/2408.03695 code: https://github.com/YeLuoSuiYou/openstorypp News 2024/7/31 We have reorganized and distributed the high-quality subset and released most of the story data collected… See the full description on the dataset page: https://huggingface.co/datasets/MAPLE-WestLake-AIGC/OpenstoryPlusPlus.image100K<n<1M5 likes1.3k downloads2y agoHugging Face04APRIL-AIGC /Soul-Bench Soul 🤓 Project    | 📑 Paper    | 🤖 Online Experience    | 🤖 API Documentation    | 🤗 Soul Model | Eval Suite   | 🤗 Soul-Bench | Results   Soul: Breathe Life into Digital Human for High-fidelity Long-term Multimodal Animation 🎋 Click ↓ to watch brief introduction for Soul, Soul-1M, and Soul-Bench TODO Release evaluation tool for Soul-Bench. Release inference code. Release training code. Inference (Soul Model) It will be released soon.… See the full description on the dataset page: https://huggingface.co/datasets/APRIL-AIGC/Soul-Bench.audioimage-to-videon<1K55 likes691 downloads9mo agoHugging Face05techjam-aigc /wildfake-eval-subset WildFake Eval Subset Reference benchmark for the AIGC-detection track, repackaged from WildFake as parquet so it loads in one line. Four configs: the spec-faithful set, plus three that remove artifacts which make the spec-faithful set trivially gameable. [!WARNING] Demonstration purposes only. Do not train on any config here. These exist so you can sanity-check a model and track iterative improvements. They do not contribute to the final score, and the final test set is drawn… See the full description on the dataset page: https://huggingface.co/datasets/techjam-aigc/wildfake-eval-subset.image100K<n<1M0 likes323 downloads23d agoHugging Face06basakdemirok /AIGCodeSet LLM vs Human Code Dataset A Benchmark Dataset for AI-generated and Human-written Code Classification Description This dataset contains code samples generated by various Large Language Models (LLMs), including CodeStral (Mistral AI), Gemini (Google DeepMind), and CodeLLaMA (Meta), along with human-written codes from CodeNet. The dataset is designed to support research on distinguishing LLM-generated code from human-written code. Dataset Structure 1.… See the full description on the dataset page: https://huggingface.co/datasets/basakdemirok/AIGCodeSet.tabulartext-classification10K<n<100K7 likes277 downloads1y agoHugging Face07Bluemaki /Flux_AIGC_Datasetimage10K<n<100K0 likes162 downloads25d agoHugging Face08ecnu-aigc /EMID Dataset Summary Emotionally paired Music and Image Dataset (EMID) is a novel dataset designed for the emotional matching of music and images. The EMID dataset contains 10,738 unique music clips, each of which is paired with 3 images in the same emotional category,as well as rich annotations. These musical clips are categorized into the 13 emotional categories proposed by What music makes us feel: At least 13 dimensions organize subjective experiences associated with music across… See the full description on the dataset page: https://huggingface.co/datasets/ecnu-aigc/EMID.audio10K<n<100K6 likes145 downloads3y agoHugging Face09strawhat /aigciqa-20kDataset from paper: `[CVPR2024] Aigiqa-20k: A large database for ai-generated image quality assessment Code: https://www.modelscope.cn/datasets/lcysyzxdxc/AIGCQA-30K-Image @inproceedings{li2024aigiqa, title={Aigiqa-20k: A large database for ai-generated image quality assessment}, author={Li, Chunyi and Kou, Tengchuan and Gao, Yixuan and Cao, Yuqin and Sun, Wei and Zhang, Zicheng and Zhou, Yingjie and Zhang, Zhichao and Zhang, Weixia and Wu, Haoning and others}, booktitle={Proceedings of… See the full description on the dataset page: https://huggingface.co/datasets/strawhat/aigciqa-20k.texttext-to-image10K<n<100K0 likes76 downloads2y agoHugging Face10aigc-x /Pronunciation-boldvoice Pronunciation Assessment Dataset (BoldVoice + speechocean762) Dataset for fine-tuning multimodal models on English pronunciation assessment. Overview Source Samples Audio Duration Description BoldVoice 38,182 10-20s Non-native English learners, BoldVoice API annotations speechocean762 5,000 1.6-20s Public dataset, 5-expert scored, Mandarin speakers Total 43,182 Schema Column Type Description audio Audio (16kHz mono) Speech… See the full description on the dataset page: https://huggingface.co/datasets/aigc-x/Pronunciation-boldvoice.audioaudio-classification10K<n<100K0 likes67 downloads6mo agoHugging Face11jizhongpeng /AIGCQA-30KAIGCQA-30K dataset ready for Q-Align training text10K<n<100K3 likes59 downloads2y agoHugging Face12stevenfan /AIGCBench_v1.0 AIGCBench v1.0 AIGCBench is a novel and comprehensive benchmark designed for evaluating the capabilities of state-of-the-art video generation algorithms. Official dataset for the paper:AIGCBench: Comprehensive Evaluation of Image-to-Video Content Generated by AI, BenchCouncil Transactions on Benchmarks, Standards and Evaluations (TBench). Description This dataset is intended for the evaluation of video generation tasks. Our dataset includes image-text pairs and… See the full description on the dataset page: https://huggingface.co/datasets/stevenfan/AIGCBench_v1.0.image1K<n<10K4 likes58 downloads3y agoHugging Face13strawhat /aigciqa2023 Dataset from paper [CICAI2023] AIGCIQA2023: A Large-scale Image Quality Assessment Database for AI Generated Images: from the Perspectives of Quality, Authenticity and Correspondence Seems that this dataset does not have a specified license. Please refer to the original source and paper for more information on its usage and redistribution policies. https://github.com/wangjiarui153/AIGCIQA2023 @misc{wang2023aigciqa2023, title={AIGCIQA2023: A Large-scale Image Quality Assessment Database… See the full description on the dataset page: https://huggingface.co/datasets/strawhat/aigciqa2023.tabular1K<n<10K0 likes43 downloads6mo agoHugging Face14mehdidc /compositionality_aigciqa2023tabular1K<n<10K0 likes30 downloads3y agoHugging Face15maze /aigcAsian photography dataset win3000: about 18k asian celebrity photo. jiepaigou: streetsnap and celebrity cybesx: about 13k street photography audio10K<n<100K2 likes28 downloads2y agoHugging Face16feeday /aigc-image Dataset Schema & Labels Column Type Description file_name image / string The primary key for the Hugging Face dataset builder prompt string The original text prompt provided to the generative model. model string AI model (NanoBanana2, Z-Image-Turbo, SRPO). generation_method string •T2I: Text-to-Image• I2I: Image-to-Image label string The semantic classification of the image subject (man, ppt). watermark string • visible: Visible watermarks (e.g., logos, text).•… See the full description on the dataset page: https://huggingface.co/datasets/feeday/aigc-image.imagen<1K0 likes14 downloads6mo agoHugging Face17Littlesalt33 /zxy-aigctextn<1K0 likes13 downloads2y agoHugging Face18AigcBtree /dsltexttext-generationn<1K0 likes9 downloads1y agoHugging Face19yingzhitao /AIGCBench_v1.0 AIGCBench v1.0 AIGCBench is a novel and comprehensive benchmark designed for evaluating the capabilities of state-of-the-art video generation algorithms. Official dataset for the paper:AIGCBench: Comprehensive Evaluation of Image-to-Video Content Generated by AI, BenchCouncil Transactions on Benchmarks, Standards and Evaluations (TBench). Description This dataset is intended for the evaluation of video generation tasks. Our dataset includes image-text pairs and… See the full description on the dataset page: https://huggingface.co/datasets/yingzhitao/AIGCBench_v1.0.image1K<n<10K0 likes9 downloads8mo agoHugging Face20Littlesalt33 /surgical-aigctext10K<n<100K1 likes6 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.