datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
UltraVideo
UltraVideo: High-Quality UHD 4K Video Dataset
🤓 Project | 📑 Paper | 🤗 Hugging Face (UltraVideo Dataset)) | 🤗 Hugging Face (UltraVideo-Long Dataset)) | 🤗 Hugging Face (UltraWan-1K/4K Weights)
UltraVideo: High-Quality UHD Video Dataset with Comprehensive Captions
🎋 Click below image to watch the 4K demo video.
🤓 First open-sourced UHD-4K/8K video datasets with comprehensive structured (10 types) captions.🤓 Native 1K/4K videos generation by UltraWan.… See the full description on the dataset page: https://huggingface.co/datasets/APRIL-AIGC/UltraVideo.UltraVideo-Long
UltraVideo: High-Quality UHD 4K Video Dataset
🤓 Project | 📑 Paper | 🤗 Hugging Face (UltraVideo Dataset)) | 🤗 Hugging Face (UltraVideo-Long Dataset)) | 🤗 Hugging Face (UltraWan-1K/4K Weights)
UltraVideo: High-Quality UHD Video Dataset with Comprehensive Captions
🎋 Click below image to watch the 4K demo video.
🤓 First open-sourced UHD-4K/8K video datasets with comprehensive structured (10 types) captions.🤓 Native 1K/4K videos generation by UltraWan.… See the full description on the dataset page: https://huggingface.co/datasets/APRIL-AIGC/UltraVideo-Long.OpenstoryPlusPlus
Openstory++: A Large-scale Dataset and Benchmark for Instance-aware Open-domain Visual Storytelling
We introduce OpenStory++, a large-scale open-domain dataset contains focusing on enabling MLLMs to perform storytelling generation tasks.
related resorcce
paper: https://arxiv.org/abs/2408.03695
code: https://github.com/YeLuoSuiYou/openstorypp
News
2024/7/31 We have reorganized and distributed the high-quality subset and released most of the story data collected… See the full description on the dataset page: https://huggingface.co/datasets/MAPLE-WestLake-AIGC/OpenstoryPlusPlus.AIGC-Detection-Benchmark
AIGC Detection Benchmark Dataset
📝 Dataset Description
Dataset Summary
The AIGC Detection Benchmark Dataset is a high-quality collection of images and associated metadata designed to benchmark models for detecting and identifying the source of artificially generated content. The dataset contains a mix of real-world images and images generated by a wide array of prominent AI models, including diffusion models (like Stable Diffusion, DALL-E 2, Midjourney, ADM) and GANs… See the full description on the dataset page: https://huggingface.co/datasets/TheKernel01/AIGC-Detection-Benchmark.Soul-Bench
Soul
🤓 Project | 📑 Paper | 🤖 Online Experience | 🤖 API Documentation | 🤗 Soul Model | Eval Suite | 🤗 Soul-Bench | Results
Soul: Breathe Life into Digital Human for High-fidelity Long-term Multimodal Animation
🎋 Click ↓ to watch brief introduction for Soul, Soul-1M, and Soul-Bench
TODO
Release evaluation tool for Soul-Bench.
Release inference code.
Release training code.
Inference (Soul Model)
It will be released soon.… See the full description on the dataset page: https://huggingface.co/datasets/APRIL-AIGC/Soul-Bench.AIGCViViDAIGC_LipSync_Benchmark
AIGC-LipSync Benchmark
📋 Overview
AIGC-LipSync Benchmark is a comprehensive evaluation benchmark specifically designed for lip synchronization in AI-Generated Content (AIGC). This benchmark consists of 615 high-quality videos covering a wide spectrum of visual representations, from realistic humans to stylized characters, enabling thorough assessment of lip synchronization methods across diverse AI-generated video scenarios.
🎯 Key Features
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/ZiqiaoPeng/AIGC_LipSync_Benchmark.AIGC-text-bank
AIGC-text-bank
AIGC-text-bank is a large-scale, multi-domain dataset designed for real-world AI-generated content (AIGC) detection. It is introduced in the paper: Reasoning-Aware AIGC Detection via Alignment and Reinforcement.
🔗 Project Homepage & Code: https://aka.ms/reveal
🌟 Dataset Overview
As Large Language Models (LLMs) rapidly advance, traditional AIGC detectors struggle to generalize, especially when facing human-AI collaborative writing (e.g., AI-polished… See the full description on the dataset page: https://huggingface.co/datasets/bmbgsj/AIGC-text-bank.wildfake-eval-subset
WildFake Eval Subset
Reference benchmark for the AIGC-detection track, repackaged from
WildFake as parquet so it loads in one
line. Four configs: the spec-faithful set, plus three that remove artifacts which make the
spec-faithful set trivially gameable.
[!WARNING]
Demonstration purposes only. Do not train on any config here.
These exist so you can sanity-check a model and track iterative improvements. They do not
contribute to the final score, and the final test set is drawn… See the full description on the dataset page: https://huggingface.co/datasets/techjam-aigc/wildfake-eval-subset.AIGCodeSet
LLM vs Human Code Dataset
A Benchmark Dataset for AI-generated and Human-written Code Classification
Description
This dataset contains code samples generated by various Large Language Models (LLMs), including CodeStral (Mistral AI), Gemini (Google DeepMind), and CodeLLaMA (Meta), along with human-written codes from CodeNet. The dataset is designed to support research on distinguishing LLM-generated code from human-written code.
Dataset Structure
1.… See the full description on the dataset page: https://huggingface.co/datasets/basakdemirok/AIGCodeSet.AIGC_Image_Steganography_Dataset
AIGC Image Steganography Dataset
📖 Dataset Description
This dataset is specifically designed for research in Artificial Intelligence Generated Content (AIGC) image steganography, steganalysis, and image forensics.
To construct a highly diverse and standardized dataset, we selected 10 prominent domestic and international text-to-image (T2I) large models and batch-generated the images via their official APIs.
During the generation process, we carefully defined 10 typical… See the full description on the dataset page: https://huggingface.co/datasets/Asketla/AIGC_Image_Steganography_Dataset.AIGC-Blur-BenchmarkAIGCDetect_testsetM3CoTBenchAIGCIQA2023EMID
Dataset Summary
Emotionally paired Music and Image Dataset (EMID) is a novel dataset designed for the emotional matching of music and images. The EMID dataset contains 10,738 unique music clips, each of which is paired with 3 images in the same emotional category,as well as rich annotations. These musical clips are categorized into the 13 emotional categories proposed by What music makes us feel: At least 13 dimensions organize subjective experiences associated with music across… See the full description on the dataset page: https://huggingface.co/datasets/ecnu-aigc/EMID.Flux_AIGC_DatasetAIGC-Blur-BenchmarkAIGCDetect_testsetaigciqa-20kDataset from paper: `[CVPR2024] Aigiqa-20k: A large database for ai-generated image quality assessment
Code: https://www.modelscope.cn/datasets/lcysyzxdxc/AIGCQA-30K-Image
@inproceedings{li2024aigiqa,
title={Aigiqa-20k: A large database for ai-generated image quality assessment},
author={Li, Chunyi and Kou, Tengchuan and Gao, Yixuan and Cao, Yuqin and Sun, Wei and Zhang, Zicheng and Zhou, Yingjie and Zhang, Zhichao and Zhang, Weixia and Wu, Haoning and others},
booktitle={Proceedings of… See the full description on the dataset page: https://huggingface.co/datasets/strawhat/aigciqa-20k.Pronunciation-boldvoice
Pronunciation Assessment Dataset (BoldVoice + speechocean762)
Dataset for fine-tuning multimodal models on English pronunciation assessment.
Overview
Source
Samples
Audio Duration
Description
BoldVoice
38,182
10-20s
Non-native English learners, BoldVoice API annotations
speechocean762
5,000
1.6-20s
Public dataset, 5-expert scored, Mandarin speakers
Total
43,182
Schema
Column
Type
Description
audio
Audio (16kHz mono)
Speech… See the full description on the dataset page: https://huggingface.co/datasets/aigc-x/Pronunciation-boldvoice.aigc-security-iddm-anime-experiment
AIGC Security Experiment Materials
This repository contains public report materials for a course experiment on AIGC synthetic image detection and black-box evasion attacks.
The experiment uses an IDDM diffusion model trained on anime face images to generate synthetic images, evaluates a real-vs-synthetic detector, and compares traditional post-processing attacks with black-box pixel-level attacks.
Public Release Scope
Included:
1000 IDDM-generated anime face… See the full description on the dataset page: https://huggingface.co/datasets/hhunugryy/aigc-security-iddm-anime-experiment.AIGCQA-30KAIGCQA-30K dataset ready for Q-Align training
AIGCBench_v1.0
AIGCBench v1.0
AIGCBench is a novel and comprehensive benchmark designed for evaluating the capabilities of state-of-the-art video generation algorithms. Official dataset for the paper:AIGCBench: Comprehensive Evaluation of Image-to-Video Content Generated by AI, BenchCouncil Transactions on Benchmarks, Standards and Evaluations (TBench).
Description
This dataset is intended for the evaluation of video generation tasks. Our dataset includes image-text pairs and… See the full description on the dataset page: https://huggingface.co/datasets/stevenfan/AIGCBench_v1.0.Sampled_AIGCBench_text2image_ar_0.625
Description
This dataset is intended for the implementation of image-to-video generation evaluations in the paper of AdaptiveDiffusion, which is composed of the original text-image pairs collected from AIGCBench v1.0 and a text file listing the randomly selected samples.
Data Organization
The dataset is organized into the following files:
AIGCBench_t2i_aspect_ratio_625.zip: 2002 images named by the index and the text description, adjusted to an aspect ratio of 0.625.… See the full description on the dataset page: https://huggingface.co/datasets/HankYe/Sampled_AIGCBench_text2image_ar_0.625.AIGC-Loc-Testsetsaigciqa2023
Dataset from paper [CICAI2023] AIGCIQA2023: A Large-scale Image Quality Assessment Database for AI Generated Images: from the Perspectives of Quality, Authenticity and Correspondence
Seems that this dataset does not have a specified license. Please refer to the original source and paper for more information on its usage and redistribution policies.
https://github.com/wangjiarui153/AIGCIQA2023
@misc{wang2023aigciqa2023,
title={AIGCIQA2023: A Large-scale Image Quality Assessment Database… See the full description on the dataset page: https://huggingface.co/datasets/strawhat/aigciqa2023.compositionality_aigciqa2023AIGC-Director-Masterclass
