datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
OpenstoryPlusPlus
Openstory++: A Large-scale Dataset and Benchmark for Instance-aware Open-domain Visual Storytelling
We introduce OpenStory++, a large-scale open-domain dataset contains focusing on enabling MLLMs to perform storytelling generation tasks.
related resorcce
paper: https://arxiv.org/abs/2408.03695
code: https://github.com/YeLuoSuiYou/openstorypp
News
2024/7/31 We have reorganized and distributed the high-quality subset and released most of the story data collected… See the full description on the dataset page: https://huggingface.co/datasets/MAPLE-WestLake-AIGC/OpenstoryPlusPlus.AIGC-Detection-Benchmark
AIGC Detection Benchmark Dataset
📝 Dataset Description
Dataset Summary
The AIGC Detection Benchmark Dataset is a high-quality collection of images and associated metadata designed to benchmark models for detecting and identifying the source of artificially generated content. The dataset contains a mix of real-world images and images generated by a wide array of prominent AI models, including diffusion models (like Stable Diffusion, DALL-E 2, Midjourney, ADM) and GANs… See the full description on the dataset page: https://huggingface.co/datasets/TheKernel01/AIGC-Detection-Benchmark.Soul-Bench
Soul
🤓 Project | 📑 Paper | 🤖 Online Experience | 🤖 API Documentation | 🤗 Soul Model | Eval Suite | 🤗 Soul-Bench | Results
Soul: Breathe Life into Digital Human for High-fidelity Long-term Multimodal Animation
🎋 Click ↓ to watch brief introduction for Soul, Soul-1M, and Soul-Bench
TODO
Release evaluation tool for Soul-Bench.
Release inference code.
Release training code.
Inference (Soul Model)
It will be released soon.… See the full description on the dataset page: https://huggingface.co/datasets/APRIL-AIGC/Soul-Bench.ViViDwildfake-eval-subset
WildFake Eval Subset
Reference benchmark for the AIGC-detection track, repackaged from
WildFake as parquet so it loads in one
line. Four configs: the spec-faithful set, plus three that remove artifacts which make the
spec-faithful set trivially gameable.
[!WARNING]
Demonstration purposes only. Do not train on any config here.
These exist so you can sanity-check a model and track iterative improvements. They do not
contribute to the final score, and the final test set is drawn… See the full description on the dataset page: https://huggingface.co/datasets/techjam-aigc/wildfake-eval-subset.AIGC_Image_Steganography_Dataset
AIGC Image Steganography Dataset
📖 Dataset Description
This dataset is specifically designed for research in Artificial Intelligence Generated Content (AIGC) image steganography, steganalysis, and image forensics.
To construct a highly diverse and standardized dataset, we selected 10 prominent domestic and international text-to-image (T2I) large models and batch-generated the images via their official APIs.
During the generation process, we carefully defined 10 typical… See the full description on the dataset page: https://huggingface.co/datasets/Asketla/AIGC_Image_Steganography_Dataset.AIGCDetect_testsetAIGCIQA2023M3CoTBenchFlux_AIGC_DatasetEMID
Dataset Summary
Emotionally paired Music and Image Dataset (EMID) is a novel dataset designed for the emotional matching of music and images. The EMID dataset contains 10,738 unique music clips, each of which is paired with 3 images in the same emotional category,as well as rich annotations. These musical clips are categorized into the 13 emotional categories proposed by What music makes us feel: At least 13 dimensions organize subjective experiences associated with music across… See the full description on the dataset page: https://huggingface.co/datasets/ecnu-aigc/EMID.AIGCDetect_testsetAIGCBench_v1.0
AIGCBench v1.0
AIGCBench is a novel and comprehensive benchmark designed for evaluating the capabilities of state-of-the-art video generation algorithms. Official dataset for the paper:AIGCBench: Comprehensive Evaluation of Image-to-Video Content Generated by AI, BenchCouncil Transactions on Benchmarks, Standards and Evaluations (TBench).
Description
This dataset is intended for the evaluation of video generation tasks. Our dataset includes image-text pairs and… See the full description on the dataset page: https://huggingface.co/datasets/stevenfan/AIGCBench_v1.0.Sampled_AIGCBench_text2image_ar_0.625
Description
This dataset is intended for the implementation of image-to-video generation evaluations in the paper of AdaptiveDiffusion, which is composed of the original text-image pairs collected from AIGCBench v1.0 and a text file listing the randomly selected samples.
Data Organization
The dataset is organized into the following files:
AIGCBench_t2i_aspect_ratio_625.zip: 2002 images named by the index and the text description, adjusted to an aspect ratio of 0.625.… See the full description on the dataset page: https://huggingface.co/datasets/HankYe/Sampled_AIGCBench_text2image_ar_0.625.aigcAsian photography dataset
win3000: about 18k asian celebrity photo.
jiepaigou: streetsnap and celebrity
cybesx: about 13k street photography
AIGCQA-30Kaigcdetect_testEditDataThis repository hosts the dataset resources for training and benchmarking image editing models.All paths are relative to the repository root when cloned via Git.
📂 Dataset Structure
After cloning, you should have the following directories:
project_name/
├── dataset/
│ └── image_edit/
│ └── OmniEdit/ # Training data
└── edit_benchmark/
└── data/ # Benchmark data
Training Data
The training set is located under:… See the full description on the dataset page: https://huggingface.co/datasets/AIGCers/EditData.aigc-image
Dataset Schema & Labels
Column
Type
Description
file_name
image / string
The primary key for the Hugging Face dataset builder
prompt
string
The original text prompt provided to the generative model.
model
string
AI model (NanoBanana2, Z-Image-Turbo, SRPO).
generation_method
string
•T2I: Text-to-Image• I2I: Image-to-Image
label
string
The semantic classification of the image subject (man, ppt).
watermark
string
• visible: Visible watermarks (e.g., logos, text).•… See the full description on the dataset page: https://huggingface.co/datasets/feeday/aigc-image.AIGCBench_v1.0
AIGCBench v1.0
AIGCBench is a novel and comprehensive benchmark designed for evaluating the capabilities of state-of-the-art video generation algorithms. Official dataset for the paper:AIGCBench: Comprehensive Evaluation of Image-to-Video Content Generated by AI, BenchCouncil Transactions on Benchmarks, Standards and Evaluations (TBench).
Description
This dataset is intended for the evaluation of video generation tasks. Our dataset includes image-text pairs and… See the full description on the dataset page: https://huggingface.co/datasets/yingzhitao/AIGCBench_v1.0.AdRectification_AIGC
