datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
generative-sound-masking-generated-energy-v7
Generative Sound Masking — fixed-background audio masking
Incrementally generated unfiltered candidates. This is not a final selected dataset.
Each background has 15 separately generated prompt–seed outputs using gain-compensated reconstruction residuals. run_config.json pins models, source pools, parameters and implementation hashes.
For multiple workers read workers/worker-NN/progress.json; each worker reports only its assigned IDs.
Global completion requires all worker… See the full description on the dataset page: https://huggingface.co/datasets/AE-W/generative-sound-masking-generated-energy-v7.flux_generatedCT-RATE_Generated_Scans
Dataset Card for Synthetic Text-to-CT Scans - VLM3D Challenge
Dataset Details
Dataset Description
This dataset contains 1,000 synthetic 3D chest CT scans generated using the model introduced in
From Alignment to Synthesis: Contrastive Volumetric Grounding for Text-to-CT Generation (Molino et al., BMVC 2026).
The model was trained on the CT-RATE dataset, the largest publicly available collection of paired CT volumes and radiology reports.
It… See the full description on the dataset page: https://huggingface.co/datasets/dmolino/CT-RATE_Generated_Scans.Nameonly_generated
Just Say the Name: Online Continual Learning with Category Names Only via Data Generation
We provide the dataset used for Name-only continual learning, generated using Stable Diffusion XL, DALL.E-2, CogView2, and DeepFloyd IF models.
Disclaimer
This dataset is created solely for academic purposes. We minimized human intervention to ensure a fair comparison with the baseline methods discussed in our paper. Despite our efforts, the extensive size of the dataset prevented us… See the full description on the dataset page: https://huggingface.co/datasets/seongwon980/Nameonly_generated.robocasa-v02-generated300-success-env-depth-rgb-256
RoboCasa v0.2 Generated-300 Success — Env Depth + RGB (256)
Per-step environment depth and RGB sidecars rendered from the original
RoboCasa v0.2 mg_im Generated-300 HDF5 demos, restricted to demos
whose final state matches the task-success predicate. The selection
covers 300 success-audited demos per task across 24 RoboCasa Kitchen
atomic tasks (~7,200 demos total) at native 256-pixel render
resolution. Geometry and RGB are emitted as per-camera contiguous
.npy memmaps designed for… See the full description on the dataset page: https://huggingface.co/datasets/SeonghuJeon/robocasa-v02-generated300-success-env-depth-rgb-256.generated-csvsMLLM-Generated-Image-Detection-Dataset
MLLM-Generated Image Dataset
This dataset contains real and AI-generated image samples organized for binary MLLM-generated image detection.
Paper | Code
Dataset Summary
We construct an MLLM-generated image detection benchmark from GPT Image2 and Nano Banana2. This benchmark covers texture-dominated, structure-dominated, and hybrid-dominated. It is designed to evaluate detector performance under the new challenges introduced by large-scale image generation models.… See the full description on the dataset page: https://huggingface.co/datasets/zr-zhang/MLLM-Generated-Image-Detection-Dataset.arc-voicesamples-generatedgenerative-sound-masking-generated-energy-v3-smokegenerated-passport-faces-aditya-second-halfgenerated-imagesgenerated_chat_0.4M
Generated_Chat_0.4M
内容
包含约40万条由BELLE项目生成的个性化角色对话数据,包含角色介绍。
注意:此数据集是由ChatGPT产生的,未经过严格校验,题目或解题过程可能包含错误。使用过程中请注意这一点。
样例
{
"instruction": "于以下角色信息完成一段对话\nAlice是一个年轻的律师,她正在为一个涉及濒危物种保护的案件做准备。Bob是一位当地的居民,他所在的社区经常受到该濒危物种的干扰。\n",
"input": "",
"output": "\nAlice: 你好Bob,我是濒危物种保护案的律师,很高兴见到你。我们前几天发给你的文件你看到了吗?\nBob: 看到了,你们说要禁止我们在该物种的常住地采集蘑菇,这很不公平。\nAlice: 保护濒危物种对整个生态系统都很重要,如果我们的举动损害了这个物种的生存,那么整个生态系统都将遭受损失。\nBob: 我理解您的立场,但是我们一直以来都依靠这个物种来维持我们的经济生活,现在我们要怎么办?\nAlice:… See the full description on the dataset page: https://huggingface.co/datasets/BelleGroup/generated_chat_0.4M.SDXL-Generated-Stanford-Dogs
Dataset Card for Generated Dogs
10+ images for each class in the Stanford Dogs dataset, but all generated with SDXL. Images were filtered for CLIP score, and cartoonish images were removed.
This is a FiftyOne dataset with 1305 samples.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import fiftyone as fo
import fiftyone.utils.huggingface as fouh
# Load the dataset
# Note: other available arguments include… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/SDXL-Generated-Stanford-Dogs.AI-Generated-vs-Real-Images-Datasets
Dataset Card for "AI-Generated-vs-Real-Images-Datasets"
More Information needed
lvdec-flf-wave-20260815-generated
LVDEC FLF Wave 2026-08-15 — Generated Returns
This public dataset contains the completed generated return package for the
Yueshengli/lvdec-flf-wave-20260815
execution wave.
Contents
The dataset contains 8,193 completed jobs with no final failures:
Generator / batch
Jobs
ltx23_flf2v
1,200
ltx23_flf2va
1,200
ltx25_flf2v
470
ltx25_flf2va
370
wan22_fun_5b_control
990
minimax_h3_flf2v
1,166
minimax_h3_flf2va
2,797
Each job directory may… See the full description on the dataset page: https://huggingface.co/datasets/Hallucinatie/lvdec-flf-wave-20260815-generated.FIGNEWS_generated_queriesThis repository contains the FIGNEWS dataset with predicted queries, a core component used in the paper QAEncoder: Towards Aligned Representation Learning in Question Answering Systems.
The official implementation and related code are available on GitHub: https://github.com/IAAR-Shanghai/QAEncoder
Introduction
Modern QA systems entail retrieval-augmented generation (RAG) for accurate and trustworthy responses. However, the inherent gap between user queries and relevant documents… See the full description on the dataset page: https://huggingface.co/datasets/zr-wang/FIGNEWS_generated_queries.real-fake-ai-generated-art-images
🎨 Real and Fake (AI-Generated) Art Images Dataset
21,642 balanced images — 10,821 real artworks and 10,821 AI-generated
images — for training models to distinguish authentic art from GAN-generated fakes.
🧭 Overview
This dataset is part of the FauxFinder project, designed to build
advanced models capable of distinguishing between authentic artworks
and AI-generated images. Ideal for binary classification, GAN research,
and computer vision benchmarking.… See the full description on the dataset page: https://huggingface.co/datasets/hmnshudhmn24/real-fake-ai-generated-art-images.ai-generated-images-classifierGenerated_summariesgenerated-videosreachy-mini-generated-moves
Reachy Mini - Generated Moves
Moves generated by the reachy-mini-text-to-motion
Space (natural language -> motion clip -> 50 Hz recorded-move JSON).
Each generation produces three files under moves/:
File
Content
<stamp>_<slug>.json
Standard Reachy Mini recorded-move JSON (50 Hz), playable by the SDK/daemon
<stamp>_<slug>.clip.json
The LLM-authored source: per-channel curves with easing + oscillators
<stamp>_<slug>.meta.json
Prompt, model, timestamp, generated… See the full description on the dataset page: https://huggingface.co/datasets/tfrere/reachy-mini-generated-moves.generated-datagenerative-sound-masking-generated-energy-v4
Generative Sound Masking — fixed-background audio masking
Incrementally generated unfiltered candidates. This is not a final selected dataset.
Each background has 15 separately generated prompt–seed outputs using gain-compensated reconstruction residuals. run_config.json pins models, source pools, parameters and implementation hashes.
For multiple workers read workers/worker-NN/progress.json; each worker reports only its assigned IDs.
Global completion requires all worker… See the full description on the dataset page: https://huggingface.co/datasets/AE-W/generative-sound-masking-generated-energy-v4.200k_HEAVY_gpt4o-description-gpt4omini-code_generated_problemsHere is the dataset of ~100k synthetic data generated by 162 seeds.
We generate the dataset with the following steps and two approaches:
Generate ~110k descriptions by GPT4o.
Approach 1: Generate ~110k codes follow each description by GPT4o-mini.
Approach 2: Generate ~110k codes follow each description by GPT4o-mini and suggest it to use specific library functions.
Run the ~220k codes and do auto-filtering.
Get the final ~200k legitimate ARC-like tasks with examples.
human_ai_generated_text
Human or AI-Generated Text
The data can be valuable for educators, policymakers, and researchers interested in the evolving education landscape, particularly in detecting or identifying texts written by Humans or Artificial Intelligence systems.
File Name
model_training_dataset.csv
File Structure
id: Unique identifier for each record.
human_text: Human-written content.
ai_text: AI-generated texts.
instructions: Description of the task given to both Humans and… See the full description on the dataset page: https://huggingface.co/datasets/dmitva/human_ai_generated_text.generated-novelsGenerated-LoRA-Input-Images-for-Mitigating-Biasgpt-oss120b-generated-perfectblendmsmarco_generated_question_runsanny-render-corpus-generated-train
anny-render-corpus-generated
Images generated by OmniGen2 from the constructed renders in
chibifire/anny-render-corpus.
Code: weftspun/anny-render-corpus, on the 6-datasource side of the hexagon.
Why this is a separate repository
These are generated synthetic, not constructed. They were sampled from a model rather than
rendered deterministically from a rig, so their labels are inferred and not true by
construction. Our working agreement requires generated data to… See the full description on the dataset page: https://huggingface.co/datasets/chibifire/anny-render-corpus-generated-train.
